API Gateway Architecture for Large-Scale SaaS Platforms

API7.ai

September 15, 2026

API Gateway Guide

A large-scale SaaS gateway should turn verified tenant identity into bounded routing, quota, and observability context while keeping domain authorization and data isolation in the services that own them. It should also limit blast radius: one tenant, region, configuration change, or dependency failure must not consume the entire platform.

The gateway is one isolation layer, not the whole multi-tenant design. It can authenticate, route, rate-limit, and label requests, but it cannot by itself prove row-level isolation, make every backend tenant-aware, or create capacity that downstream systems do not have.

Key Takeaways

  • Derive tenant context from verified credentials and server-side mapping, never directly from a public tenant header.
  • Define separate isolation budgets for requests, concurrency, expensive operations, and shared dependencies.
  • Partition the data plane into cells or regions when a single global failure domain becomes unacceptable.
  • Keep policy configuration centralized enough to govern but scoped enough to prevent a bad change from affecting every tenant.
  • Choose fail-open or fail-closed behavior per dependency and document both paths; availability is not a universal reason to bypass isolation.

Start With the Tenant Trust Boundary

“Tenant ID” is useful only when its source is clear. A client may send a hostname, path segment, header, API key, or token claim that suggests a tenant. The gateway must first validate the credential and then map it to an internal tenant identifier. RFC 8725 warns implementations not to trust received JWT claims without validation, including issuer, audience, signature, and application-specific rules.

Use this sequence:

sequenceDiagram
    participant C as Client
    participant G as API gateway
    participant I as Identity provider or trusted key source
    participant T as Authoritative tenant directory
    participant S as SaaS service
    participant D as Tenant-aware data layer
    C->>G: Request plus credential and claimed context
    G->>I: Obtain identity response or trusted verification keys
    I-->>G: Identity response or keys
    G->>G: Validate signature, issuer, audience, and credential state
    G->>T: Map verified principal to allowed tenant
    T-->>G: Authoritative tenant and cell mapping
    G->>G: Replace public tenant hint with trusted context
    G->>S: Request plus gateway-created tenant context
    S->>S: Enforce object and workflow authorization
    S->>D: Query within authoritative tenant boundary
    D-->>S: Tenant-scoped result
    S-->>G: Response
    G-->>C: Response

If a public X-Tenant-ID header is useful for routing, compare it with the verified mapping and overwrite or reject it. Do not forward two competing versions. The service must trust the gateway-created context only over an authenticated and bypass-resistant internal path.

Build Isolation in Layers

LayerGateway responsibilityResponsibility that remains elsewhere
IdentityValidate supported credentials and create minimal trusted contextAccount lifecycle, tenant membership, revocation policy
AdmissionApply request, concurrency, payload, and cost boundsProtect workers, queues, database pools, and third-party spend
RoutingSelect region, cell, API version, and serviceKeep tenant placement authoritative and consistent
DataPrevent obvious cross-tenant routing and strip untrusted contextEnforce row, schema, database, or account isolation
ObservabilityEmit tenant-safe route and policy outcomesControl sensitive fields, retention, audit, and incident access
ConfigurationApply reviewed policy artifactsApprove entitlements, exceptions, and business policy

Rate limits are only one budget. A tenant can stay below an RPS limit while exhausting long-lived connections, database locks, AI inference spend, or a slow third-party API. Define at least request rate, concurrency, payload, and expensive-operation budgets, then align them with downstream capacity.

Choose a Data-Plane Topology

Shared Regional Fleet

A shared fleet is simple and efficient. It fits smaller tenant counts and workloads with similar policy. Its risk is a large failure domain: a bad plugin, route, configuration, or traffic spike can affect many tenants.

Cell-Based Fleet

Cells assign subsets of tenants to independent gateway and service capacity. A global directory maps a verified tenant to a cell; the cell owns routing, quotas, and backends for that subset. Cells reduce blast radius and make capacity easier to reason about, but placement, migration, and cross-cell operations become platform responsibilities.

flowchart TB
    C[Clients] --> E[Global edge and tenant routing]
    E --> A[Cell A gateway]
    E --> B[Cell B gateway]
    A --> AS[Cell A services and data]
    B --> BS[Cell B services and data]
    P[Reviewed policy source] --> PA[Cell A control path]
    P --> PB[Cell B control path]

The global edge should route from verified placement data, not a caller-selected cell header. Control paths distribute configuration; they do not carry user requests.

Dedicated Tenant Fleet

Dedicated gateways may be justified for regulatory, network, performance, or contractual isolation. They cost more to operate and can create configuration drift. Use the same policy source, validation, and observability contract as shared cells unless a documented tenant requirement differs.

Apply Tenant Quotas Carefully in APISIX

Apache APISIX 3.18.0 provides Consumer Groups for sharing plugin configuration across consumers and the limit-count plugin for local or Redis-backed counters. A consumer group can represent an operational policy cohort or a tenant, but it is not a substitute for application data isolation.

This JSON is a configuration excerpt for a tenant-scoped group. It assumes authentication has already selected consumers that are legitimately bound to tenant-acme; callers cannot choose group_id directly.

{ "plugins": { "limit-count": { "count": 200, "time_window": 60, "policy": "redis", "redis_host": "redis.internal", "redis_port": 6379, "key_type": "constant", "key": "tenant-acme", "group": "quota-tenant-acme", "allow_degradation": false, "rejected_code": 429 } } }

key_type: constant and key: tenant-acme make all authenticated consumers bound to this tenant-specific Consumer Group use one quota key. The separate group value lets matching plugin configurations share that counter across routes; it does not replace the counting key. If key_type and key are omitted, APISIX 3.18 defaults to var and remote_addr, which creates a separate quota per source IP rather than one tenant-wide quota. The constant is trustworthy only because an operator sets it server-side in the Consumer Group configuration, and authentication selects a Consumer already bound to that group; never derive it from a public tenant header.

Redis credentials and transport security are intentionally omitted from this excerpt. For production, configure Redis authentication and TLS according to the version-pinned plugin documentation.

allow_degradation changes the isolation contract. Its default is false: in APISIX 3.18.0, a limiter or Redis error on this request path returns HTTP 500 rather than bypassing the limiter. With true, APISIX continues handling requests without that limiter, so a Redis failure can remove the quota protection. The configured rejected_code: 429 applies when the quota is exhausted; it is not the Redis-failure response. These branches are implemented in the version-pinned limit-count source and should still be exercised in the target runtime.

Choose deliberately:

  • false returns request-time HTTP 500 for the documented APISIX 3.18.0 limiter/dependency error path rather than bypassing enforcement;
  • true favors request availability but can expose shared capacity to an unbounded tenant;
  • neither path proves tenant identity—the authentication and consumer binding do that;
  • local counters apply per gateway instance, while a Redis policy coordinates counters across instances with its own availability and consistency trade-offs.

Separate Entitlements From Runtime Counters

Store commercial plans and entitlements in an authoritative control system. Compile the allowed quota policy into gateway configuration, identify the revision, and audit changes. Do not call a billing database on every request if a cached entitlement with a defined staleness window is sufficient.

Metering and enforcement are different. A quota counter decides whether a request proceeds. Usage events support analytics or billing and may be delivered asynchronously. Losing a metering event should trigger reconciliation; it should not silently change the runtime enforcement decision unless that behavior is explicitly designed.

Plan Regional and Dependency Failure

For each dependency, record the absent, slow, failed, and recovered paths:

  • identity provider: cache duration, revocation exposure, and behavior after cache expiry;
  • quota store: fail-open or fail-closed decision and capacity guardrail;
  • tenant directory: stale placement handling and migration window;
  • control plane: last-known-good configuration and recovery process;
  • regional cell: failover eligibility, data residency, and state availability;
  • telemetry pipeline: buffering, sampling, and sensitive-data limits.

Do not promise automatic regional failover when tenant data, keys, queues, or residency rules cannot move with the request. The gateway can redirect traffic only to a backend that can safely serve it.

Verification Checklist

  • Can a public tenant header override the verified tenant mapping?
  • Can a credential for tenant A reach a route, object, log stream, or counter for tenant B?
  • Do authenticated consumers using different source IPs in the same tenant consume one shared quota window?
  • Does one tenant's concurrency or expensive work starve others?
  • What happens when the shared counter store is absent, slow, or recovered?
  • Do local and distributed counters have acceptable accuracy for the business rule?
  • Can a policy revision be rolled out to one cell and rolled back without global impact?
  • Are logs useful without placing personal or secret data in labels?
  • Can the platform reconstruct which identity, policy revision, route, and cell handled an incident?

FAQ

Is a tenant header safe after the gateway adds it?

Only if the gateway first removes or rejects the caller's value, derives the replacement from verified identity, and sends it over a path the caller cannot bypass or impersonate.

Should every tenant have a dedicated gateway?

No. Shared fleets or cells are usually more efficient. Dedicated fleets are justified by concrete isolation, network, capacity, or contractual requirements.

Does a distributed rate limit guarantee exact fairness?

No. Counter algorithm, synchronization, failure behavior, retries, and request cost all affect the result. Treat it as one admission control in a layered design.

Next Steps

Define distributed rate-limit accuracy and failure trade-offs, add concurrency budgets and backpressure, and establish auditable access evidence before onboarding high-impact tenants.

Share article link