AISIX 1.5.0: Run Non-Chat AI Workloads Through Model Groups

September 29, 2026

Products

Chat is often the first AI workload a team sends through a gateway. It is rarely the last. Production systems also submit embeddings, reranking, image work, audio, and video jobs—and each workload can otherwise grow its own provider-selection and recovery logic.

AISIX 1.5.0, released September 29, 2026, makes Model Groups available on supported single-request endpoints. The practical change is not that every AI endpoint now behaves alike. It is that a team can apply a Model Group's target selection, retry and failover path to more than chat, while retaining endpoint-specific support checks and per-attempt operational evidence.

AISIX processes gateway traffic in your environment. AISIX Cloud provides the control plane, dashboard, organizations, and centralized usage views in its supported deployment options. The routing behavior described here belongs to the AISIX gateway; it does not turn every provider or endpoint into an interchangeable service.

Key Takeaways

  • In 1.5.0, a Model Group can serve supported single-request endpoints for completions, embeddings, rerank, image generation and edits, audio, and video submission; those endpoints previously returned 400 for a Model Group.
  • The group applies its existing strategy order, health filter, retry budget, and failover path. least_latency and least_busy can balance targets on these endpoints.
  • Every target must support the endpoint being addressed. /v1/realtime and the jobs surface still reject Model Groups by design.
  • The serving target is reflected in the usage record and routing telemetry. A failed attempt creates a zero-token, non-billed usage row; it is not proof that an upstream provider did not charge independently.

One Routing Policy for the Workloads Around a Product

Consider an AI product that indexes submitted documents, reranks search candidates, and creates an image variation for a finished brief. Before this release, its chat path could use a Model Group, while a single-request path had to address an individual model or receive 400 when it named the group.

With AISIX 1.5.0, the application can address a Model Group on the supported single-request endpoint. AISIX filters eligible targets, applies the group's strategy and health information, then records the target that served the request. A failed eligible attempt can enter the group's retry or failover path according to its configured budget.

flowchart LR
  A[Embedding or rerank request] --> G[Model Group]
  G --> H[Health and endpoint support]
  H -->|eligible target| P[Provider target]
  P -->|success| U[Serving target and usage row]
  P -->|retryable failure| R[Retry or failover]
  R --> H

This is useful when routing is an operating policy, rather than application glue. The same group can use its strategy order and health filter for chat and for the supported single-request traffic that belongs to the same service. For least_latency and least_busy, the group can balance eligible targets on those endpoints as well.

The model group does not create endpoint capability. Each target must support the endpoint the caller addressed. Treat that as a deployment check: a group that is appropriate for embeddings may not be appropriate for image edits, and an eligible chat target is not automatically eligible for a different workload.

Keep Failure Evidence Separate From Spend

The routing extension becomes more useful when operators can distinguish a request's terminal outcome from the attempts that preceded it. AISIX 1.5.0 writes a usage row for each failed upstream attempt on single-request endpoints, including completions, embeddings, rerank, image generation and edits, audio endpoints, video submission, and /v1/messages/count_tokens.

Those failed-attempt rows carry the attempt kind and error, report zero tokens, and are not billed by AISIX. When all attempts fail, the final attempt is the request's terminal row. A direct model whose retry budget replays a 5xx follows the same attempt-recording behavior.

That is gateway accounting and observability, not a statement about a provider invoice. Do not infer an upstream provider's charge from a zero-token AISIX row. Use the per-attempt rows to diagnose whether the group retried or failed over, then reconcile any provider billing through the provider's own records.

The release also makes the destination visible in normal operation: the usage row is attributed to the target that served the request, and deployment/fallback metrics plus the access-log routing summary are populated. This supplies the evidence needed to validate a rollout without guessing which target received a successful request.

Do Not Generalize the Endpoint Boundary

The new scope is broad, but not universal. AISIX lists the supported paths as /v1/completions, /v1/embeddings, /v1/rerank, /v1/images/generations, /v1/images/edits, /v1/audio/transcriptions, /v1/audio/translations, /v1/audio/speech, and video submission. /v1/realtime and the jobs surface still refuse Model Groups by design.

This distinction matters especially for asynchronous-looking workloads. Video submission is in the supported set, but polling a video job and downloading its content remain calls that do not write a usage row. Those later calls are not evidence that model-group routing or usage accounting applies to the entire video lifecycle.

Keep fallback behavior equally specific. A group can walk its targets under the configured retry and failover policy; it does not promise that a target which lacks the requested capability will serve the request, or that every failure can be retried safely. Validate the actual target inventory and the configured retry budget before moving production traffic.

Upgrade the Routing and Observability Contract Together

Use a controlled rollout to verify both routing and the surrounding 1.5.0 changes:

  1. List every Model Group target against every single-request endpoint your application calls. Remove or segregate targets that do not support the endpoint; do not use chat support as a proxy.
  2. Send a representative request to each supported endpoint through the group. Verify the serving target in the usage record, deployment/fallback metrics, and access-log routing summary.
  3. Test a retryable upstream failure and a non-retryable one. Confirm the per-attempt zero-token rows, attempt kind, terminal row, and your alerting behavior. Reconcile provider charges separately.
  4. Keep /v1/realtime and jobs paths on their supported routing configuration. Do not treat the release as enabling Model Groups there.
  5. Review the other 1.5.0 behavior changes before upgrading: metrics that add a side label can change unfiltered query totals; usage-event occurred_at now has millisecond precision; and streamed failures after headers now use the failure status in AISIX records while cost, budgets, and rate limits remain unchanged.

For security-sensitive deployments, review the release notes separately before adopting its new shared-secret OIDC option. It selects HMAC verification when hmac_secret is configured, has its own version-compatibility gate for registered gateways, and is not part of Model Group routing.

Expand the Routing Policy, Preserve Its Boundaries

AISIX 1.5.0 gives teams a clearer way to run supported non-chat AI work through the routing policy they already use for chat. That can reduce one-off selection logic for embeddings, reranking, media, and related single-request paths while making the serving target and failed attempts visible.

The value depends on the boundaries: targets must support the endpoint, /v1/realtime and jobs remain excluded, and a zero-token gateway attempt row is not a provider-billing conclusion. Start with the AISIX 1.5.0 release notes, then test the endpoint and target combinations you operate before expanding traffic.

Tags:
Share article link