Istio 1.31 Adds Agentgateway as a Waypoint, Giving AI Agent API Traffic Its Own Proxy
Istio 1.31 is out, and one line in the release notes says a lot about where API infrastructure is heading. You can now run agentgateway, a proxy built for AI agent and MCP traffic, as a waypoint inside the mesh.
If “waypoint” means nothing to you yet, that’s fine. Here’s the short version, then why it matters if you build APIs.
What changed in 1.31
Istio’s ambient mode splits the mesh in two. A small per-node component handles encryption and basic connectivity. Anything that needs to understand HTTP, like routing, auth policy or rate limits, goes through a separate proxy called a waypoint. You opt services into a waypoint with a label.
Istio 1.30 added experimental support for agentgateway at the edge only. Version 1.31 adds a new istio-agentgateway-waypoint GatewayClass, so the same proxy can sit between services inside the cluster. The release also fixes several bugs around ListenerSet handling and mTLS to agentgateway backends.
Two other changes are worth knowing about. Istio stopped publishing release artifacts to its Google Cloud registries, so if your install scripts pull from gcr.io/istio-release, they need updating before you upgrade. And there’s new support for canary waypoints, where you send a set share of traffic to a new waypoint before cutting over.
Why a separate proxy for agents?
A normal API gateway was built for human-driven apps. Requests are short, the caller is a known client, and the rate limit is counted in requests per minute.
Agent traffic breaks some of those assumptions. A single agent session might call ten tools in a row. The expensive part of a call to an LLM backend is tokens, not requests. And MCP has its own message structure that a generic HTTP proxy can’t look inside.
That’s the gap agentgateway is aimed at. Solo.io’s docs for its enterprise version list rate limiting by tokens and limits specific to MCP, which a classic gateway doesn’t offer.
Lesson one: count what actually costs money
Most beginner API courses teach rate limiting as “100 requests per minute per key.” That works when every request costs you about the same.
It stops working when one request can cost a thousand times more than another. If you’re putting an LLM behind your API, or exposing tools to agents, think about limiting by the real cost unit:
- Tokens in and out, for model calls
- Rows scanned or returned, for data endpoints
- Compute seconds, for jobs
- Tool calls per session, for agent workflows
Return the remaining budget in response headers so a well-behaved client can slow down on its own. That’s the same idea as X-RateLimit-Remaining, just with a better unit.
Lesson two: put policy outside the service
The waypoint model is a nice illustration of something you’ll keep running into. Auth checks, rate limits and routing rules don’t have to live in your application code. They can live in a proxy in front of it, written once, applied to many services.
That doesn’t mean your service skips validation. It means cross-cutting rules that are the same for every endpoint get enforced in one place, and your handlers focus on business logic. When the rule changes, you change one config, not twelve repos.
With Istio, attaching a service to a waypoint is just a label:
kubectl label service orders istio.io/use-waypoint=waypoint
Lesson three: roll out proxies like code
The canary waypoint feature is worth copying as a habit even if you never touch Istio. A gateway change can break every API behind it at once. So treat it like a code deploy. Send a small slice of traffic through the new version first, watch error rates and latency, then move the rest.
Should you learn this now?
If you’re early in learning APIs, no need to stand up a service mesh this week. But pay attention to the direction. Agents are becoming a major class of API caller, and the infrastructure is starting to treat them that way. Learn rate limiting and gateway policy with that in mind, and you’ll be ahead of most courses still teaching the 2015 version.