Designing APIs for Agentic CLIs: Safe Discovery and Change
September 29, 2026
The next important API consumer may not be a person reading documentation or an SDK written once by a developer. It may be an AI agent discovering an operation, assembling parameters, executing it, and interpreting the result inside one task.
That shift was visible in this week's Hacker News discussion about Cloudflare's new cf CLI. Cloudflare says agent usage reached 48% of Wrangler use during the preceding week, and its new CLI generates commands from the same OpenAPI descriptions used for documentation and SDKs. It also makes JSON the default output and adds natural-language command search across more than 3,000 API operations.
The lesson is broader than one CLI: an agentic CLI exposes the quality of the API beneath it. A schema can help an agent discover syntax, but safe automation also needs explicit semantics for identity, side effects, concurrency, retries, approval, and recovery.
Key Takeaways
- Use one versioned API description as the source for documentation, SDKs, and CLI discovery.
- Return stable, bounded, machine-readable results; keep human presentation as a separate view.
- Describe side effects and reversibility, not only request and response fields.
- Give agents short-lived, least-privilege credentials and constrain them at the gateway.
- Make mutations observable and retry-safe with idempotency, preconditions, and operation status.
- Separate planning from applying when an action is expensive, destructive, or difficult to reverse.
Why an Agentic CLI Changes the API Contract
A traditional CLI assumes that a human already knows the command hierarchy. Help text can be long, output can be a table, and an error can suggest that the operator "try again". An agent approaches the interface differently. It searches for a capability, selects an operation from metadata, fills parameters from context, then feeds structured output into its next decision.
This makes ambiguity expensive. If two commands use different verbs for the same action, discovery becomes less reliable. If a success message is prose, the agent may parse the wrong identifier. If an error does not distinguish a conflict from a transient failure, an automatic retry may repeat a mutation.
The OpenAPI Specification is designed so humans and computers can understand an HTTP API without inspecting its implementation. That is a strong foundation for an agentic interface, but it is not a complete operational contract. OpenAPI can describe an operation, parameters, schemas, and security schemes; the API team still has to define what the operation changes and how a caller can recover.
Start with Discoverable, Versioned Operations
An agent should not have to load thousands of endpoints into its context. Give it a compact index that supports search by purpose, resource, risk, and required permission. Each result should provide a stable operation identifier, a concise description, the target environment, and a link or command for retrieving the full schema.
Good operation descriptions answer four questions:
- What resource does this operation read or change?
- Which identity and permission does it require?
- Is the action read-only, reversible, or destructive?
- What proves that it completed?
Generate the CLI, SDK, and documentation from the same versioned contract where practical. This reduces drift, but generation does not make the source correct. Treat schema reviews like code reviews: check descriptions, examples, defaults, deprecations, and security requirements. API7.ai's OpenAPI and Swagger guide explains how the description becomes a shared interface for design and governance.
Keep operation identifiers stable even if the visible command changes. An agent can migrate from a deprecated operation when the replacement is explicit; it cannot safely infer that two similarly named actions have equivalent effects.
Make JSON Predictable and Bounded
Machine-readable output should be the primary contract for automation. A terminal table can remain useful for humans, but it should be rendered from the same structured result rather than becoming a second semantic interface.
A successful response should expose stable field names, resource identifiers, version or revision, and an unambiguous status. List operations need deterministic pagination and documented ordering. Long-running operations need an operation ID and a status endpoint rather than a connection that remains open until an unknown deadline.
Errors need structure too. Return a stable error code, a human-readable message, whether retry may be appropriate, and field-level details when input is invalid. Do not tell an agent to retry every 5xx response blindly: a server may have committed the change before the connection failed.
Bound response size and pagination. An agent that requests "all logs" or "all resources" can consume excessive tokens, memory, or upstream capacity even when the request is authorized. Limits, continuation tokens, and time ranges are part of the interface, not implementation trivia.
Treat Every Mutation as a State Transition
For a write operation, define the starting state, requested transition, and success evidence. A useful pattern is:
discover current revision -> plan change -> validate inputs -> apply with precondition -> observe operation -> verify resource
Use conditional requests or explicit revision fields so an agent does not overwrite a change made after it read the resource. HTTP defines PUT and DELETE as idempotent methods, but method semantics alone do not prove that a specific application handles retries safely. For POST operations such as purchases, deployments, or job creation, support an idempotency key when duplicate execution would be harmful.
Return the same logical result for a repeated request with the same key, and document the key's scope and retention window. When an outcome is uncertain, let the caller query by operation ID or idempotency key before retrying.
For high-impact actions, add a plan/apply split. The plan should identify affected resources, permissions, estimated scope, and irreversible consequences. Apply should bind to the reviewed plan or resource revision so a stale approval cannot authorize a different change.
An API gateway cannot invent these application semantics. It can enforce the traffic contract around them, while the control-plane API remains responsible for transaction and recovery behavior.
Put a Policy Boundary in Front of Agent Traffic
An agent should receive the narrowest identity that can complete the task. Prefer short-lived credentials tied to a workload, user delegation, environment, and allowed operation set. Avoid sharing a broad personal token across agents or repositories.
At the gateway, authenticate the caller, authorize the route and method, validate request shape, apply quotas, and attach a correlation identifier. Apache APISIX 3.18.0, for example, documents plugins for OpenID Connect, request validation, and request-count limits. These controls are independent: a valid token does not imply permission for every route, and a valid schema does not make a requested action safe.
Use server-validated identity and policy attributes. A caller-supplied header such as X-Agent-Name is useful as a label only after a trusted component sets or verifies it; it is not proof of identity by itself. Similarly, a request ID supports correlation but should not be used as an authorization principal.
Separate read and write credentials where possible. A discovery process may need broad read access while the apply phase needs a narrow, time-limited write grant. Rate limits should reflect operation cost, not merely raw request count: a list call and a region-wide deployment are not equivalent units of risk.
Design an Approval Boundary, Not an Approval Illusion
Some changes should require a person or an independent policy decision. The API must make that boundary explicit. Mark operations by risk class, return a reviewable plan, and require approval for actions such as deleting durable data, changing identity policy, purchasing resources, or exposing a public endpoint.
Approval must bind to concrete inputs. "Allow the agent to fix DNS" is too broad; approving a hash of a plan that changes one record from one value to another is auditable. Expire approvals and invalidate them when the underlying resource revision changes.
Do not rely on a CLI prompt as the sole control. An agent may run non-interactively, suppress prompts, or call the API directly. The authoritative approval check belongs on the server side or at a policy enforcement point that every path crosses.
Observe Intent, Decision, and Effect
An audit trail should let an operator reconstruct:
- who or what initiated the request;
- which delegated user, workload, and environment were involved;
- which operation and resource revision were selected;
- what policy decision allowed or denied it;
- what changed, when it completed, and whether it was rolled back.
Log enough to explain the action without copying secrets or sensitive payloads. Keep the caller-provided intent separate from gateway-generated identity and from the backend's final result. Monitor repeated denials, large discovery scans, retry storms, unusual environments, and changes outside an approved window.
This is where an API-first approach and an API gateway reinforce each other: the description makes the surface discoverable, while the gateway applies consistent runtime policy and telemetry across consumers.
A Practical Readiness Checklist
- Every operation has a stable ID, purpose, risk class, and permission requirement.
- JSON responses and errors have versioned, bounded schemas.
- Pagination, ordering, limits, and long-running operation behavior are documented.
- Mutations support preconditions, idempotency, or explicit recovery where needed.
- Destructive actions expose plan/apply and server-side approval boundaries.
- Agent credentials are short-lived, scoped, revocable, and separated by environment.
- Gateway authentication, authorization, validation, and quotas are independently configured.
- Audit records distinguish caller claims, trusted identity, policy decisions, and effects.
- Contract tests exercise absent, duplicate, stale-revision, timeout, and partial-failure paths.
- Deprecated operations name a tested replacement and a removal date.
Build for Automation Without Surrendering Control
Agentic CLIs are making APIs the primary user interface for infrastructure automation. The winning design is not the interface that lets an agent do everything with one broad credential. It is the interface that makes the right operation easy to discover, the result easy to interpret, and a risky change difficult to execute accidentally.
Explore API7 Enterprise to place authentication, authorization, validation, traffic controls, and observability at a consistent gateway boundary while your control-plane APIs define safe state transitions.



