Make Your Product Callable by Agents: An Agent-Friendly API Design Guide
In 2026, APIs are increasingly called by agents, not humans. This guide dissects the five pillars of agent-friendly API design — contracts, error codes, idempotency keys, pagination, and machine credentials — through real cases from Stripe and GitHub, plus a ready-to-use checklist, OpenAPI quality scorecard, and self-test prompt template.

No matter how beautiful your API docs are, the first real "reader" may no longer be human. In 2026, a growing share of product calls happen between agents: a coding agent creates a Stripe invoice for a user at 3 AM, a support agent calls your returns endpoint directly, a data agent pulls your reporting API every hour. They don't read your marketing page or watch your screenshot tutorials. They recognize exactly three things: machine-readable contracts, predictable errors, and retry-safe semantics.
The uncomfortable truth: most APIs are "readable by humans, uncallable by agents." A human developer who hits a 400 opens the docs, guesses at parameters, retries by hand. An agent that hits the same 400 stuffs the vague error message into its context and starts banging its head against the wall with its "intelligence" — retrying three times, reformatting parameters, finally hallucinating a field name. Those "inexplicable calls" in your support queue? Probably not hackers. Probably an agent feeling around in the dark.
This guide gives you a complete playbook: why Stripe and GitHub are the honor students, two real-world cautionary tales, the five pillars (contracts, error codes, idempotency, pagination, machine credentials), and three things you can take away today — an agent-callability checklist, an OpenAPI description quality scorecard, and a self-test prompt template that lets Claude or GPT actually call your API and grade it.
1. The API's "Second User" Has Arrived
For twenty years, API design assumed one thing: the caller is a human engineer who reads docs, debugs, and files tickets. Every "best practice" orbited that assumption: write clear docs, build a pretty dashboard, make error messages read like prose.
An agent caller is a different species, and its traits change what good design means:
- It doesn't "understand," it matches. A human sees
user_nameanduserNameand guesses they're the same thing. An agent leans on literal contracts. Inconsistent field naming doesn't get "intuited" — it gets mis-sent. - It has no patience, only a token budget. A human will spend 20 minutes debugging one failed call. Every agent retry burns money. A vague error message forces more rounds of guessing — your API is helping it burn cash.
- It runs concurrently, retries, and reconnects after drops. A human clicks "submit" once. An agent's retry logic can fire the same POST five times in a minute. Without idempotency, your "create order" endpoint becomes a "create five orders" endpoint.
- It reads your OpenAPI, not your website. Agent tool descriptions are often generated straight from your OpenAPI JSON. A lazy
descriptionfield is a torn map handed directly to the agent.
The test is simple: if your API in 2026 is still optimized only for humans, you're locking out the fastest-growing class of callers. Agent traffic won't knock and announce itself. It will quietly pick the competitor whose API actually works.
2. Good Example #1: Stripe Turned "Retry-Safety" into Infrastructure
Stripe is the gold standard of the agent era not because its docs are pretty (that's a side effect), but because it answers, at the protocol level, the three questions agents fear most: may I retry? will I know why it failed? will an upgrade break me?
Permission to retry: Idempotency-Key
Stripe supports the Idempotency-Key header on all POST requests. Resend the same key within 24 hours and Stripe returns the first request's result instead of executing again. Translated for agents: it can finally write "retry on failure" logic without fear of double-charging.
POST /v1/charges
Idempotency-Key: order-8f3a2b-create-001
Authorization: Bearer sk_test_...
# Network timeout — retry with the same key
POST /v1/charges
Idempotency-Key: order-8f3a2b-create-001
# → returns the first result; no second charge created
Note the design subtlety: the key is generated by the caller (a UUID or a business-meaningful key); the server only has to "recognize the key, not the caller." That's smarter than server-side dedup: an agent that crashes and restarts can safely resume as long as it remembers the key.
Knowing why it failed: the structured error object
Stripe's error body is textbook material:
{
"error": {
"type": "card_error",
"code": "card_declined",
"decline_code": "insufficient_funds",
"param": "exp_month",
"message": "Your card's expiration month is invalid.",
"doc_url": "https://stripe.com/docs/error-codes/card-declined"
}
}
Four layers, each with a job: type tells the agent what kind of problem this is (retry or give up), code is a stable, programmable identifier (goes in a switch branch), param pinpoints the offending field (no guessing which parameter was wrong), message is for humans. The agent can make decisions without NLP — that is what "machine-readable" actually means.
Upgrades that don't explode: Stripe-Version
Stripe pins API versions with the Stripe-Version: 2024-06-20 header and never ships silent breaking changes. For an agent, that means: the call sequence that worked last week still works this week. Compare that with APIs that rename a field on Wednesday night and break every integration by Thursday morning.
The verdict: none of Stripe's three moves is "good documentation." They're all protocol design. Agent-friendliness isn't a copywriting problem — it's a contract problem.
3. Good Example #2: GitHub's REST API Makes "Quota" a Programmable Fact
GitHub's REST API is the other model student. Its strength: turning runtime state into part of the response, so agents can "drive while watching the dashboard."
Transparent rate limiting
Every response carries three headers:
X-RateLimit-Limit: 5000
X-RateLimit-Remaining: 4999
X-RateLimit-Reset: 1712345678 # Unix timestamp
Secondary rate limits return 403/429 with Retry-After. The agent never has to guess "am I being throttled?" or blindly exponential-backoff — the headers say exactly how much is left and when it recovers. Operational knowledge encoded into the protocol, consumable by the agent's scheduler directly.
Conditional requests save tokens
GitHub supports ETag / If-None-Match: unchanged resources return 304 Not Modified with an empty body. Agents polling for sync don't re-download everything each time — real savings in tokens and bandwidth. When designing an API, ask "will this GET be polled at high frequency?" If yes, conditional requests are mandatory.
Link-header pagination
GitHub expresses pagination relationships ( rel="next", rel="last" ) in the RFC 5988 Link response header instead of burying page URLs in some corner of the body. Parsing response headers is standard behavior for agents — no per-vendor body-pagination parser required.
The verdict: GitHub's philosophy is "never make the caller maintain a mental model." Quotas, caching, pagination — everything the caller would have to "remember" arrives inside each response. Agents have no reliable long-term memory; self-describing responses are their external brain.
4. Bad Example #1: GitHub Search API's 1,000-Result Ceiling
Same company, different story. GitHub's Search API hides an agent-killer: only the first 1,000 search results are available — the docs state it plainly ("only the first 1,000 search results are available").
For humans, this doesn't matter — nobody reads past page 10. For agents, it's a disaster: it writes a dutiful "walk all results" loop, pages honestly to page 34 (30 per page), then the API starts returning empty and the agent concludes "that's everything." It has no idea it's looking at a truncated world — worse, the response carries no machine-readable signal saying "this isn't all of it."
This is "silent truncation" — the nastiest class of problem in agent-friendly design. A returned error can at least be handled; data that looks complete but isn't leads the agent to make wrong decisions on a wrong universe, with zero exceptions along the way.
Three actionable lessons:
- Any hard cap must be declared explicitly in the response (
truncated: true, withtotal_hitsandreturned_hitsreported separately); - Provide tools to shrink the result set (finer filters, sort options) instead of making callers "page and pray";
- If truncation is a business necessity, say so in bold in the OpenAPI
description— that text is what agents read when generating tool descriptions.
The diagram above traces API paradigm shifts, each answering "who is the caller?" RPC-era callers were internal services, REST-era callers were human developers, GraphQL-era callers were frontend apps — and in the agent era, the caller is a program that can't read hints, only contracts. Design has to evolve with the caller.
5. Bad Example #2: "200 for Everything" — the Cost of Lying Status Codes
Another real-world anti-pattern: countless internal APIs and early SaaS APIs "always return 200, with the error in the body":
HTTP/1.1 200 OK
{ "success": false, "errorMsg": "Insufficient balance" }
A human glances at the body and gets it. But the agent's HTTP client — and the retry middleware, gateways, and observability stack behind it — only reads the status code: 200 means success, so it caches, records a success metric, and moves on. A "successful" 200 carrying a failed body sails through the agent's toolchain with all green lights until it explodes three steps later, with logs full of "success." Worse, field names like errorMsg are spelled differently at every company (error_msg, errMsg, message) — agents can't write generic parsers.
GraphQL takes this to the extreme: per the spec, a GraphQL service always returns 200 OK, with errors in a top-level errors array. That's spec design, not a bug — but it means agents need a separate "find errors inside 200s" code path, and HTTP-layer retries, circuit breakers, and alerts all go blind. If you're exposing a GraphQL endpoint to agents, at minimum use stable machine-readable codes in errors[].extensions.code (e.g. UNAUTHENTICATED, BAD_USER_INPUT) and state explicitly in the docs: "this endpoint always returns 200; errors are reported via the errors array."
The verdict: the status code is the cheapest machine signal in HTTP, and lying with it makes the entire toolchain misjudge. Better an ugly 422 than a lying 200.
6. Pillar One: Contract First — the OpenAPI Description Quality Scorecard
An agent's "documentation" is your OpenAPI JSON. Plenty of teams ship "good enough to generate docs" OpenAPI: operations without descriptions, schema fields unexplained, examples all empty. Feed that contract to an agent and you're asking it to drive blindfolded.
This scorecard works in code review or release gates. Each dimension scores 0–2, 20 points total:
| Dimension | 0 points | 1 point | 2 points |
|---|---|---|---|
| operationId uniqueness | Missing or duplicated | Unique but meaningless (e.g. op1) | Unique and semantic (e.g. createOrder) |
| Operation description | Missing | One sentence restating the title | Documents preconditions, side effects, idempotency |
| Parameter description coverage | <30% | 30%–80% | >80%, with formats and value ranges |
| Request/response examples | None | Success example only | Success + typical error examples |
| Schema field docs | Bare types | Some fields documented | All core fields documented with examples |
| Error code docs | None | Lists HTTP statuses only | Each business code explained with handling advice |
| Auth definition | No securityScheme | Defined but scopes unexplained | Scheme + scopes + how to obtain, complete |
| Versioning & change marks | None | Has a version number | Deprecated marks + replacements + sunset dates |
| Webhook/callback docs | None (when webhooks exist) | Event names only | Event payload schemas + signature verification + retry policy |
| Rate limit docs | None | Mentioned somewhere in docs | Quota numbers + headers + over-limit behavior, all three |
Grading: 16–20, agents can integrate directly; 10–15, workable but needs human escort; below 10, don't expect any agent to succeed — fix the contract before promoting it. Put this in CI: OpenAPI-change PRs must include a score, and anything under 16 doesn't merge. Contract quality can be managed with numbers.
One oft-missed detail: description should be written from the caller's perspective. "Create order" is a failing description; "Creates a pending-payment order for the given user; repeated calls with the same Idempotency-Key within 24h return the first result; amounts are in cents" is one an agent can use. Those two extra sentences save the agent ten rounds of trial and error.
7. Pillar Two: Error Code Conventions — Let Agents Make Decisions
Error responses have exactly one design goal: let the caller know "what happened and what to do" without reading the docs. The spec:
{
"error": {
"code": "INSUFFICIENT_BALANCE", // stable, unique, programmable; never renamed
"message": "Insufficient balance, 1200 cents available", // human-readable, may change
"param": "amount", // the offending field (required for param errors)
"request_id": "req_8f3a2b1c", // end-to-end trace ID
"doc_url": "https://docs.example.com/errors/INSUFFICIENT_BALANCE",
"retryable": false, // worth retrying? consumed directly by agents
"details": { "available": 1200, "required": 5000 }
}
}
Key decisions:
codein SCREAMING_SNAKE_CASE, never renamed. A rename is a breaking change. The code is the agent's switch-branch condition and matters more than the HTTP status — there are only a few dozen statuses and they lack expressive power ("insufficient balance" and "bad parameter" are both 400, but the agent's handling strategy differs completely).- The
retryableboolean is an "action instruction" for agents. A human seeing "service busy" decides on their own to try later; an agent needs an explicit signal. Moving the "should I retry?" judgment from caller to server is the highest-ROI single field in agent-friendly design. request_idon every request (successes too, via theX-Request-Idheader). When an agent reports a failure it will paste the request_id, and your on-call engineer locates the issue with it. Without it, an agent's bug report is just "it says no."- 4xx is the caller's fault, 5xx is yours — keep them strictly separate. Wrapping a "downstream timeout" as a 400 forces the agent to "fix" a parameter it cannot fix. Semantically scrambled error codes are worse than none.
8. Pillar Three: Idempotency Keys — the Prerequisite for Agent Retries
Agents retry as a way of life: network jitters, timeouts, killed processes, failed steps re-run from scratch. A POST endpoint without idempotency keys is "Schrödinger's create" for an agent — it can never know whether that timed-out request actually executed.
The standard semantics (copied from Stripe, battle-tested)
- Clients send an
Idempotency-Keyheader on all non-idempotent writes (POST/PATCH), with a UUID or business-unique key as the value; - The server stores the first request's result (status code + body) keyed by it, TTL 24 hours;
- A repeat arrival with the same key: if parameters match, return the stored result (status
200, optionally with anIdempotent-Replayed: trueheader marking it a replay); if parameters differ, return422withcode: IDEMPOTENCY_KEY_IN_USE— preventing "one key, different business"; - Keys are only for "create-like" semantics: naturally idempotent PUT (full replace) and DELETE don't need them. Don't add keys to everything for "consistency" — that's noise.
Minimal implementation (Python/Flask-style pseudocode)
def idempotent(ttl=86400):
def deco(fn):
def wrapper():
key = request.headers.get("Idempotency-Key")
if not key:
return fn() # read request or keyless caller: execute as-is
saved = store.get(f"idem:{key}")
if saved:
if saved["fingerprint"] != fingerprint(request.json):
return error(422, "IDEMPOTENCY_KEY_IN_USE"), 422
resp = make_response(saved["body"], saved["status"])
resp.headers["Idempotent-Replayed"] = "true"
return resp
status, body = fn() # first execution
store.setex(f"idem:{key}", ttl,
{"status": status, "body": body,
"fingerprint": fingerprint(request.json)})
return make_response(body, status)
return wrapper
return deco
The verdict: idempotency keys are the seatbelt of agent-era APIs. You can drive without one, but you wouldn't let the car drive itself. Schedule it as P0 — it determines whether agents dare put your write endpoints into automated flows.
9. Pillar Four: Pagination Contracts — Don't Let Agents Get Lost Flipping Pages
Pagination is where agents stumble most, because they need to "walk the entire set" — humans only read page one. One unified contract:
{
"data": [ ... ],
"pagination": {
"next_cursor": "eyJpZCI6MTAwMH0=", // null means the end
"has_more": true,
"limit": 100
}
}
- Cursor, not offset. Offset skips or duplicates rows when data changes mid-walk; a cursor anchors to "where the last page ended," keeping traversal stable. For agents doing full syncs, offset pagination is a data-quality incident waiting to happen.
- Cursors must be opaque. Agents will try to parse your cursor ("looks like base64 JSON, let me tweak it"). State in the docs: "do not parse, do not construct" — and return 400, not 500, for malformed cursors.
- Sorting must be stable and declared.
sort=created_at:desc,id:desc— the second key is the tie-breaker so rows created in the same second don't reshuffle between pages. Put the default sort in the OpenAPI description. has_moreplusnext_cursor: nullas double confirmation. The simpler the agent's "am I done?" logic, the better — don't make it infer completion from "this page wasn't full."- Limits must be sane and honest. Whether the cap is 100 or 1,000 depends on your per-row cost, but say so in the docs and in the 400 error. GitHub Search's lesson: truncation is fine, silence is not.
10. Pillar Five: Machine Credentials and Quota Transparency
The most-forgotten pillar: if the agent can't get through the door, nothing inside matters.
- Offer machine-to-machine credentials. The OAuth authorization-code flow requires "a user clicking allow in a browser" — a headless agent can't click. Provide at least one human-free path: long-lived API keys (rotatable, fine-grained scopes), OAuth client-credentials mode, or GitHub-style fine-grained personal access tokens. Make "getting credentials" a three-step-or-fewer pure API/CLI flow; every extra manual step tanks agent adoption.
- Scopes should be fine-grained, defaults minimal. Split
orders:writefromorders:read— the smaller the permission an agent requests, the more willing users are to grant it. Stripe's restricted keys are the model: two-dimensional grants by resource and action. - Quotas belong in response headers, not doc paragraphs. Copy GitHub's
X-RateLimit-*trio. An agent's scheduler is a program; it reads headers, not prose. - A sandbox is the agent's test drive. Ship a test mode / sandbox so agents can get it working before touching production. Stripe's
sk_test_prefix is elegant: the environment is visible in the key itself, so an agent can't accidentally spend real money with a test key.
Four more iron rules for the semantic layer, listed as a cheat sheet for space: timestamps always RFC 3339 UTC (2026-10-10T08:00:00Z) — no bare epoch numbers, no local timezones; money as integers in the smallest currency unit (cents) — floats in JSON are precision bombs; enum values in UPPER_SNAKE, additive-only; webhooks must be signed (X-Signature + timestamp anti-replay) — the first thing an agent does on receiving a callback is verify, not parse.
11. Deliverable #1: The Agent-Callability Checklist
Three tiers by priority. P0 is "agents can't call it" breakage, P1 is "works but expensive," P2 is "nice to have." Check each before shipping:
P0 — without these, don't bother
- □ An OpenAPI 3.x contract exists and matches production (CI checks drift automatically)
- □ All write endpoints (POST/PATCH) support Idempotency-Key; replays return the first result
- □ Error bodies carry a stable code (never renamed) + request_id + doc_url
- □ Errors declare retryability (4xx param errors non-retryable; 429/5xx retryable)
- □ Honest status codes: no 200-wrapped errors; GraphQL endpoints declare their errors semantics
- □ Machine credentials with no human interaction required (API key / client credentials)
- □ Cursor pagination with has_more; hard caps declared explicitly in responses
- □ Globally consistent time/money/enum formats (RFC 3339 / smallest-unit integers / additive-only)
P1 — halve the cost of calling
- □ OpenAPI scorecard ≥ 16 points (success + error examples included)
- □ Rate-limit trio headers (limit/remaining/reset) + Retry-After
- □ Conditional requests (ETag/If-None-Match) on high-frequency-polled GETs
- □ Fine-grained scopes, least privilege by default
- □ Sandbox/test-mode environment, key prefixes distinguishing environments
- □ Signed webhooks with replay protection, event payload schemas
- □ Sensitive operations (charges/deletes) require a confirmation parameter or dry-run mode
- □ Breaking changes announced 30 days ahead in the changelog +
Deprecationresponse header
P2 — bonus points
- □ An official MCP server or OpenAPI-to-tool config snippets
- □ X-Request-Id on responses, traceable on success and failure alike
- □ Batch endpoints to cut agents' N+1 calls
- □ Structured filters on search endpoints, avoiding "pull everything then filter" waste
- □ Transition validation on state-machine resources (illegal transitions return 422 + hint at legal next steps)
- □ An official "Agent quickstart" page: from key to first write call in ≤ 10 minutes
Note the P1 item on sensitive-operation confirmation: agents make mistakes. Requiring confirm: true or offering dry_run before a charge is a "confirm lane change" button for self-driving. It's not distrust of agents — it's acknowledging that every caller errs.
12. Deliverable #2: The "Machine-Readability" Self-Test Prompt Template
Don't ship right after writing the contract. Copy this prompt into Claude or GPT (with a read-only/sandbox key) and let it play "an agent seeing your API for the first time," calling and grading it. This is the highest-ROI action in the whole guide: half an hour of test-driving exposes every "obvious to us" assumption in your docs.
You are an API integration engineer seeing this API for the first time.
You may only use the OpenAPI contract (attached) and the sandbox
credentials I give you. Do not guess at undocumented behavior.
Tasks:
1. Complete one full business loop with the sandbox key
(create → read → update → delete/cancel). For each step,
record: which endpoint you called, what parameters you sent,
what came back.
2. Deliberately cause 3 kinds of errors: invalid parameters,
a nonexistent resource ID, and a duplicate create with the
same Idempotency-Key. Record each HTTP status, error.code,
and whether you could tell "should I retry?"
3. Page through an entire list endpoint. Record whether the
pagination mechanism is reliable and whether you hit any
silent truncation.
Deliver a report containing:
- A callability score (0–100) with a deduction breakdown
- Every place you were forced to "guess" (field meanings,
formats, defaults, error meanings)
- A P0/P1/P2-tiered fix list
- What stands between you and 100 ("zero-guess onboarding")
Rules: when something is ambiguous, log it first, then proceed
the most conservative way. Never invent fields. Never assume
"it's probably like this."
Usage notes: the key must be sandbox/read-only — never feed a production key to a model; paste the full OpenAPI JSON into the prompt (the contract itself is the exam paper); cross-validate with two models — Claude and GPT stumble in different places, and the intersection is where the real problems are. Run it before every release and put the score in your release checklist.
13. Rollout SOP: A Two-Week Retrofit Plan
The checklist is long, but rollout has an order. At this pace, an existing API goes from "human-usable" to "agent-usable" in two weeks:
Days 1–2: measure. Run the self-test prompt above for a baseline score and a P0 list. Also check OpenAPI-vs-implementation drift (run schemathesis or openapi-diff once — a drifted contract hurts agents more than no contract).
Days 3–5: fix P0. In order: error-body conventions (half a day) → honest status codes (one day) → Idempotency-Key (two days, storage + replay logic) → cursor pagination (one day). If machine credentials are missing, add them this week — nothing else is testable without them.
Days 6–8: fill the contract. Bring OpenAPI descriptions/examples to 16 points per the scorecard. The trick: make the endpoint's author write the description. If they can't articulate "preconditions and side effects," the endpoint's semantics are the problem — fix the endpoint before the docs.
Days 9–10: P1 and guardrails. Rate-limit headers, ETags, dry-run, Deprecation declarations. Agent guardrails (confirmations, sandbox) go live this week.
Days 11–14: regression test-drive + launch. Run the self-test prompt again; confirm P0 is cleared and the score is ≥ 80. Publish the "Agent quickstart" page with the OpenAPI URL, sandbox key signup, and MCP config snippets on a single page.
Scheduling principle: P0 is ordered by "agents can't call it," not by "dev effort." Idempotency-Key takes two days but belongs in week one; an MCP server is cool but it's P2 — don't let it cut the line.
14. Closing: The API Is Becoming the Product's Primary UI
A closing judgment. For the past decade, a product's UI was its website and apps, and the API was "capability exposed along the way." The next five years flip that: agent calls will outnumber human clicks, and the API becomes the primary interface through which the product is used. On that day, your "homepage" is your OpenAPI description, your "onboarding" is the three-step sandbox key flow, your "support" is structured error codes.
Stripe and GitHub didn't predict the future — they just took "the caller is a program" seriously back in the 2010s. Now everyone else gets to catch up. The order is laid out above: first let agents dare to retry (idempotency), then let them know what went wrong (error codes), then let them go far (pagination), and finally open the door (machine credentials).
Go run that self-test prompt now. The first time your API is seriously "read" by an agent, you may meet it again for the first time.
Related articles

Meta's SWE-sweep hides 4,068 real bugs across 100 repos in 22 languages — and gives agents no issue descriptions. The best setup fixes 75.3% with a human bug report, 4.8% without. The cliff shows that finding the problem, not writing the fix, is the human moat.

Glow disclosed that AI coding agents, working around the gh CLI's image-attachment limit, pushed 13,000+ internal screenshots into public GitHub repos: customer billing records, unreleased features, money-movement console recordings. The scariest part isn't the scale — it's one agent's workaround being saved as a skill file and spreading to a dozen agents within a week.

DuckDB v2.0's CLI can now tell whether its caller is an AI agent: box tables become compact Markdown, truncation is declared explicitly, errors ship as JSON, and long queries quote their cost first. Official benchmarks on 22 plain-English TPC-H questions show 59% fewer output tokens — and an honest 0.5% total cost saving. A paradigm-shift case study in redesigning CLI output for the model reader.