AI Gateway Rate Limiting: Requests, Tokens, Models, Providers, and Budgets
API7.ai
September 28, 2026
AI gateway rate limiting controls more than requests per second. LLM and agent workloads consume variable numbers of input and output tokens, hold streaming connections open, trigger retries, call tools, and use providers with different quotas and costs. A request-only limit can protect connection volume while still allowing one workload to consume most of a team's token budget.
A production design therefore combines several controls: request rate, token rate, concurrency, provider quota, model access, and spending guardrails. The correct combination depends on the workload and provider contract.
For the broader architecture, start with the AI Gateway Guide.
Why Traditional Request Limits Are Not Enough
A traditional API limit often counts each request as one unit. That works when requests have relatively predictable cost. AI requests are different:
- A short classification request may use a few hundred tokens.
- A long-context analysis may use tens of thousands of input tokens.
- Output length is not fully known when the request begins.
- Streaming responses occupy connections while tokens are generated.
- Agent workflows may create multiple model and tool calls from one user action.
- Retries and fallback can multiply provider consumption.
- Providers impose different request and token quotas.
Two users sending the same number of requests can create very different load and spend. Request limits remain useful, but they should be paired with usage-aware controls.
Rate-Limit Dimensions for AI Traffic
| Dimension | What it controls | Typical purpose |
|---|---|---|
| Requests | Calls over a time window | Protect endpoints and provider request quotas |
| Input tokens | Prompt and context volume | Limit large-context consumption |
| Output tokens | Generated response volume | Control generation cost and duration |
| Total tokens | Combined input and output usage | Apply a common usage budget |
| Concurrency | Simultaneous in-flight calls | Protect connections and upstream capacity |
| User or consumer | Usage by an authenticated identity | Fairness and accountability |
| Team or tenant | Shared organizational usage | Department or customer budgets |
| Model | Calls or tokens for a selected model | Protect scarce or expensive models |
| Provider | Consumption sent to one vendor | Respect provider quotas and diversify traffic |
| Route or application | Workload-specific usage | Separate chat, batch, agent, and embedding workloads |
| Tool or MCP server | Agent tool calls | Prevent loops or excessive tool execution |
The gateway needs a trustworthy key for the selected dimension. A client-supplied header should not be treated as user or tenant identity unless an upstream authentication process validates and controls it.
Request-Based Rate Limiting
Request limits answer questions such as:
- How many calls may this consumer send per minute?
- How much burst traffic may an application create?
- How many concurrent streaming requests may a tenant hold?
- How quickly may an agent call one tool?
They are useful for abuse protection and fairness, especially before token usage is known. However, they do not represent actual model consumption. A practical design keeps a request or concurrency limit as the first boundary and adds token-aware controls for cost and provider quotas.
Token-Based Rate Limiting
Token limits treat model usage as the constrained resource. They can be applied to input tokens, output tokens, or the total.
Input Tokens
Input-token controls can reject or route requests whose prompt and context exceed a policy. This protects context-window usage and limits unexpectedly large retrieval results. The gateway or upstream integration must use a tokenizer compatible with the target model or rely on provider-reported usage; token counts are not interchangeable across every model.
Output Tokens
Output limits are more difficult because the final number is unknown before generation. Teams can set a requested maximum, meter usage while streaming, or account for provider-reported usage after completion. These mechanisms have different enforcement timing. Post-response accounting can inform future decisions but cannot retroactively stop tokens already delivered.
Total Tokens
Total-token limits support quotas such as tokens per user per minute or tokens per team per month. They are useful for budgets and fair allocation but still need a policy for unknown output. A design may reserve an estimated amount before the call and reconcile after completion, or allow the call and update the balance from provider usage data.
Rate, Quota, and Budget Are Different
- A rate limit controls usage over a short window, such as tokens per minute.
- A quota controls a longer allowance, such as tokens per day or month.
- A budget converts usage into a spending boundary, often using model-specific prices.
Do not use the terms interchangeably. A team can stay under a per-minute rate while exceeding its monthly budget. It can also have budget remaining while hitting a provider's short-window token quota.
Cost estimates depend on current provider pricing, model, input/output classification, cached-token rules, and other contract details. A gateway can meter usage, but billing reconciliation should acknowledge that provider invoices remain authoritative unless the integration contract states otherwise.
Limit Scope and Fairness
A global limit is easy to configure but can create noisy-neighbor problems. Scope limits according to ownership and risk:
flowchart TD org[Organization Budget] --> teamA[Team A Quota] org --> teamB[Team B Quota] teamA --> appA[Application Limit] teamB --> appB[Application Limit] appA --> userA[Consumer Limit] appB --> userB[Consumer Limit] appA --> modelA[Model and Provider Limit] appB --> modelB[Model and Provider Limit]
Hierarchical policies let a platform enforce an organizational ceiling while giving teams and applications predictable shares. The policy needs clear precedence: when multiple limits apply, document which one denies the request and how the caller can identify the exceeded boundary.
Provider and Model Quotas
Model providers may enforce both requests-per-minute and tokens-per-minute quotas. An AI gateway can track internal usage and route traffic, but provider-side accounting may differ because of timing, tokenizer behavior, retries, or requests sent outside the gateway.
Platform teams should:
- Track quotas by provider account and region where applicable.
- Reserve capacity for critical workloads.
- Separate expensive or capacity-constrained models from default models.
- Decide whether a limit failure should stop, queue, or route a request.
- Monitor provider
429responses and quota headers when available. - Reconcile gateway estimates with provider usage reports.
Multi-provider routing can reduce dependency on one quota, but fallback is not free. The alternate model may have different behavior, price, latency, context limits, or data-handling terms. Routing changes should be explicit and observable.
Streaming and Concurrency
Streaming changes resource use even when request volume is low. A connection may remain open while the provider generates tokens, and the gateway may not know final usage until the stream ends.
Combine:
- A maximum number of concurrent streams per consumer or application.
- Request timeouts appropriate to the workload.
- Input-size and requested-output limits.
- Incremental or post-completion token accounting.
- Cleanup for disconnected clients and interrupted streams.
If the client disconnects, verify whether the provider call is cancelled or continues generating. Continuing upstream work can consume quota even though the user no longer receives the response.
Retries, Fallback, and Hidden Multiplication
Retries and fallback improve resilience but can increase usage. One client request may create several provider attempts.
The rate-limit design should specify:
- Whether each provider attempt counts against internal request limits.
- How token usage from failed or partial attempts is recorded.
- Maximum retry count and retryable failure classes.
- Whether fallback changes the model or data-processing boundary.
- How duplicate or non-idempotent tool calls are prevented.
Retries should be bounded and observable. A retry loop that is invisible to the consumer can exhaust both quota and budget.
AI Agent and MCP Traffic
An agent may call a model, select a tool, call an MCP server, process the result, and call the model again. A single human action can produce an unpredictable number of downstream operations.
Useful controls include:
- Maximum model turns per task.
- Maximum tool calls per task or time window.
- Per-tool and per-MCP-server limits.
- Total token or spending budget per agent run.
- Concurrency limits for autonomous workflows.
- Circuit breakers for repeated failure or looping behavior.
The gateway can enforce traffic-level limits, while the agent runtime may need to enforce task semantics such as maximum steps. See What Is an MCP Gateway? and Building an AI Agent Traffic Management Platform.
Responses When a Limit Is Reached
A limit policy should define both enforcement and developer experience:
- Return a consistent status and machine-readable error.
- Identify the applicable limit without exposing other tenants' data.
- Include retry timing only when it is accurate.
- Distinguish internal policy denial from provider quota failure.
- Record the dimension and scope that triggered the decision.
- Avoid automatic fallback when it changes an approved model or data boundary.
Queueing can smooth batch workloads, but it may be unsuitable for interactive chat. Routing to a less expensive model can preserve service, but it can also change quality. The response strategy should match the workload rather than using one global behavior.
Metrics and Alerts
Track at least:
- Allowed and denied requests by policy and scope.
- Input, output, and total tokens.
- Concurrent requests and streams.
- Provider
429responses. - Retry and fallback attempts.
- Usage by consumer, team, application, model, and provider.
- Estimated cost and remaining quota where available.
- Disconnects and incomplete streams.
Avoid high-cardinality labels that make the monitoring system expensive or unstable. User-level reporting may belong in logs or analytics storage rather than every metrics label.
For the broader signal model, continue with AI Gateway Observability.
Production Checklist
- Authenticate the identity used as the limit key.
- Use request and concurrency limits even when token limits exist.
- Define how input and output tokens are counted for each supported model.
- Document pre-request enforcement versus post-response accounting.
- Scope limits by caller API key, team, member, model, and provider as needed.
- Bound retries and include all attempts in usage records.
- Define fallback behavior and model-approval boundaries.
- Test streaming completion, cancellation, timeout, and provider failure.
- Alert before provider quotas or organizational budgets are exhausted.
- Reconcile estimated usage with provider reporting.
- Load-test shared counters if limits must work across gateway nodes.
AISIX AI Gateway and API7 Gateway AI Plugins
AISIX AI Gateway documents request, token, and concurrency controls for caller keys and models, plus conditional policies that can match teams, members, keys, models, model names, and providers. In AISIX Cloud, budgets can target an organization, environment, caller API key, provider key, team, member, or each member in a team. An application can receive its own budget when its traffic is represented by a dedicated caller API key. Token usage is settled from provider responses, so teams should test enforcement timing for streaming, failed, and retried requests.
API7 Gateway and Apache APISIX expose a separate ai-rate-limiting plugin for APISIX-based deployments. That plugin uses local or Redis-backed counters and can work with AI proxy plugins. Route, Service, Consumer, Consumer Group, and plugin-instance configuration belong to this API7 Gateway/APISIX model, not to the AISIX caller-key and policy model.
For either product, verify supported models, token accounting behavior, policy precedence, failure responses, and deployment requirements against the documentation for the exact version you plan to run.
Next Steps
- Review the broader AI Gateway Guide.
- Design the signal model with AI Gateway Observability.
- Compare gateway types in AI Gateway, MCP Gateway, and API Gateway.
- Evaluate AISIX AI Gateway for AI-native traffic controls.
AISIX AI Gateway
Manage LLM and agent traffic with routing, token controls, security, and observability.
Explore AISIX AI Gateway