AI Gateway High Availability: Survive Control-Plane Outages
September 29, 2026
An AI gateway should not stop serving accepted inference policy merely because its management plane is temporarily unreachable. The core design is an active-active data plane that handles traffic directly, keeps an accepted configuration locally, exposes traffic readiness separately from configuration freshness, and degrades shared features deliberately.
That is only one layer of availability. A production design must also address gateway instance loss, restarts during a control-plane outage, Redis degradation, telemetry backpressure, provider failure, and mixed-version upgrades. This article turns those cases into an architecture and test plan, using AISIX AI Gateway 1.4.0 as the concrete implementation reference.
Key Takeaways
- Keep live inference traffic off the control-plane request path.
- Run gateway instances across the failure domains you intend to survive.
- Distinguish in-memory continuity from restart continuity through an on-disk snapshot.
- Monitor traffic readiness and configuration freshness as separate signals.
- Define degraded behavior for budgets, rate limits, cache, telemetry, and certificates.
- Test control-plane loss, container restart, Pod replacement, mixed versions, and recovery independently.
High Availability Has Multiple Failure Domains
"Run three replicas" addresses instance loss, not the entire AI request path. At minimum, review four domains:
- Application traffic: DNS, ingress, client connections, and load balancers.
- Gateway data plane: processes or Pods that authenticate, route, enforce policy, and proxy inference.
- Management path: configuration distribution, administrative APIs, heartbeats, and certificate rotation.
- Upstream and shared services: model providers, Redis, identity systems, telemetry destinations, MCP servers, and A2A agents.
flowchart LR
A[Applications] --> L[Regional load balancer]
L --> G1[AISIX gateway A]
L --> G2[AISIX gateway B]
C[Control plane] -. config and heartbeat .-> G1
C -. config and heartbeat .-> G2
G1 --> R[(Redis)]
G2 --> R
G1 --> P[Model providers]
G2 --> P
The gateways, not the control plane, receive application traffic. A management outage can therefore coexist with a working but increasingly stale data plane. Availability claims must name which domain failed and what behavior remains.
Build for Instance Failure and Graceful Drain
Run at least two gateway instances across the hosts, zones, or regions in scope. Size the survivors for the loss you claim to tolerate; replicas do not help if losing one zone saturates the rest.
AISIX documents /livez for process liveness and /readyz for traffic readiness. On SIGTERM or SIGINT, /readyz returns 503 immediately while /livez remains healthy. The gateway continues accepting connections for at least shutdown.min_drain_secs, which defaults to 30 seconds in 1.4.0, then waits without its own fixed deadline for in-flight work. Align the platform termination deadline with the longest accepted inference stream plus operational margin.
In Kubernetes, use a startup probe so slow initialization is not mistaken for a crash loop. The AISIX Helm chart's default startup-probe settings provide a 300-second retry budget, not a measured startup guarantee. The Kubernetes probe documentation likewise distinguishes startup, liveness, and readiness responsibilities.
What a Control-Plane Outage Changes
The behavior below is pinned to AISIX AI Gateway 1.4.0 and API7 documentation commit cdc36eb4bd04ead937fef000fa7774b75a4c8395. Check the release notes for another version.
After a gateway applies a valid projected configuration, it retains that accepted snapshot in memory. The offline resilience documentation states that a running gateway continues serving it when the configuration connection is unavailable. New resources, revocations, and policy changes do not arrive until connectivity returns.
This creates two simultaneous states:
- Traffic continuity: accepted routes, callers, and policy can keep working.
- Management staleness: the running configuration may no longer reflect current intent.
AISIX 1.4.0 keeps /readyz healthy after a valid configuration is applied and separates source connectivity, revision, and configuration hash into /status/config. That prevents a shared control-plane outage from making every otherwise serviceable replica withdraw itself. It also means operators must alert on configuration age and disconnected state. Stale policy is not automatically safe simply because requests still succeed.
Running Continuity Is Not Restart Continuity
In-memory state disappears on restart. AISIX can persist its accepted configuration, but the snapshot cache defaults to disabled for both AISIX Cloud and etcd-backed deployments. Enable it explicitly when restart-during-disconnect is a requirement:
managed: snapshot_cache_enabled: true snapshot_cache_path: "/var/lib/aisix/config_cache.json"
A restarted instance can serve while reconnecting only if the enabled cache contains a usable snapshot. Without one, AISIX does not bind the proxy listener until valid configuration applies; /readyz is therefore not available on that listener during the wait.
Storage lifetime changes the guarantee. A Kubernetes emptyDir survives a container restart in the same Pod but not Pod replacement. Surviving replacement requires storage that outlives the Pod, with a separate cache location per replica. Do not describe the container-level path as Pod-level recovery.
The snapshot includes provider credentials and other sensitive configuration without AISIX application-level encryption. Restrict file permissions, storage access, backups, and diagnostic collection. Availability should not create a new credential-exposure path.
Prove That Configuration Reached Callers
An administrative write succeeding does not prove that every gateway serves the intended model. Distribution also differs by deployment: standalone open-source file configuration, etcd watches, and AISIX Cloud projection have different source paths.
A release pipeline should apply dependencies in order, inspect each gateway's revision or hash, query the caller-visible model interface, and send a real request through the production route. During a partition, preserve the accepted configuration as the explicit operating state rather than repeatedly assuming new writes reached isolated gateways.
Track at least:
- source connectivity and last successful update;
- published versus applied revision;
- configuration hash across replicas;
- last load failure and last-known-good state;
- caller-visible model and policy behavior.
Define Shared-Service Degradation
High availability requires documented failure modes for stateful dependencies.
| Feature | AISIX 1.4.0 boundary | Operational consequence |
|---|---|---|
| Cloud budget | Decisions are cached for 5 seconds; no cached decision denies. After that, the last decision may be reused within AISIX_DP_BUDGET_STALE_MAX_SECONDS, default 600 | Choose stale window and failure mode for financial risk |
| Budget stale ceiling exceeded | Sticky and fail-closed modes deny; fail-open allows | Test cold, stale, and first-request paths |
| Rate-limit Redis unavailable | timeout_secs defaults to 5 seconds (minimum 1) for each Redis connection attempt or round trip. At startup, an unavailable or refusing Redis serves with per-replica counters while a background task reconnects. During runtime, a connectivity failure or timeout triggers automatic local fallback, a 30-second short circuit, and background PING recovery | The first affected request may pay one Redis timeout; later requests use local quotas. Cluster-wide counting stops, and outage counts are not reconciled |
| Cache Redis unavailable | Redis errors and timeouts are treated as misses; after a failure the circuit short-circuits for 30 seconds while background PING probes recovery | Protect provider capacity and cost |
| External exporter | Sends directly when its destination is reachable | Control-plane continuity is not exporter delivery proof |
| Certificate rotation | Requires control-plane connectivity | Current certificate works only through its validity window |
The 600-second budget value is a product default, not a universal recommendation. A strict financial boundary may need fail-closed behavior, while a low-risk user flow may accept a bounded stale decision. The 5-second decision cache is separate from this stale-reuse ceiling.
Per-replica fallback weakens enforcement: three local counters can exceed a shared ceiling. Protect upstreams; alert on fallback duration.
Telemetry Has Retry and Compatibility Limits
AISIX 1.4.0 usage telemetry is queued in memory. When the connected control plane advertises batch deduplication, the gateway can resend the same batch ID with exponential backoff. It gives up when the oldest event reaches 30 minutes or after eight consecutive failures that returned a response; connection errors and timeouts are bounded by the age limit instead.
Deduplication state also has a mixed-version boundary. A response without the signal clears remembered support, while a no-response failure preserves the last signal. Therefore, after a newer replica advertises deduplication, a no-response attempt to an older replica can be retried with the same batch ID even though that replica cannot deduplicate it. Duplicate-free retry is safe only when every control-plane replica is on 1.4.0; otherwise a batch may be single-attempt or, in that mixed-fleet sequence, duplicated. A full in-memory queue also drops newly emitted events. Consequently, 1.4.0 does not guarantee lossless telemetry. Monitor queue drops, retry exhaustion, batch age, destination connectivity, and the versions on both sides.
Provider Availability Is a Separate Contract
A healthy gateway cannot make an unavailable model provider respond. Use bounded timeouts and retries, and consider fallback or load balancing only when the workload permits it. The AI gateway load-balancing guide and multi-cloud AI gateway architecture describe those patterns.
A fallback model may have different tool support, privacy terms, regional processing, price, latency, or output quality. Validate it as a separate capability contract. Avoid retry amplification during overload; Google's SRE guidance on handling overload and cascading failures explains that feedback risk.
The same boundary applies to MCP servers and A2A agents. Gateway availability does not imply downstream tool availability or preserved side effects.
Exercise the Failure Matrix
| Scenario | Expected evidence |
|---|---|
| Stop one gateway | readiness transition, completed streams, survivor saturation |
| Disconnect control plane | successful accepted traffic plus freshness alert and stable revision |
| Restart container while disconnected | snapshot load result and listener-open timing |
| Replace Pod while disconnected | storage survival or intentional wait for control plane |
| Remove Redis | local-limit or cache-miss behavior and aggregate traffic |
| Exhaust provider | bounded attempts, timeout stage, total latency, fallback identity |
| Mix 1.3.x and 1.4.0 | dedup capability, retry or drop metrics, component versions |
| Restore connectivity | revision convergence, policy validation, caller-visible success |
Measure steady-state and degraded capacity before the exercise. A recovery drill should also test a revocation made during the partition, a cold budget decision, a long stream during drain, and telemetry that exceeds its retry window.
Product Boundaries and Non-Claims
AISIX is API7.ai's AI gateway product. Apache APISIX and API7 Enterprise are separate products or projects. The snapshot, budget, and managed telemetry behavior described here applies specifically to AISIX AI Gateway 1.4.0; standalone file mode has no AISIX Cloud control-plane dependency.
High availability does not mean configuration remains fresh, snapshots are encrypted by AISIX, rate limits stay cluster-wide during Redis degradation, telemetry cannot be lost, certificates rotate offline, provider fallback preserves semantics or residency, or spare replicas guarantee enough capacity.
Keep Serving Without Hiding Staleness
The strongest AI gateway design does two things at once: it keeps accepted traffic flowing when management systems fail, and it makes every stale or degraded condition visible to operators.
Explore AISIX AI Gateway and turn control-plane loss, restart, shared-service failure, and recovery into tested parts of an AI infrastructure runbook.



