MCP or Custom Scripts? A Decision Framework for Giving Agents External Capabilities
MCP servers, custom scripts, and direct HTTP calls each come with their own bill. A three-step decision tree picks your direction in 30 seconds, five real-world scenarios show what to use and how to write it, plus a six-item pre-build acceptance checklist. The default answer is always: write the script first.

The Takeaway First: This Is Not a Tech Choice, It's a Cost Calculation
You are working in Claude Code, Cursor, or any agent, and you want it to look up a Notion database, send a Slack message, or run a read-only SQL query. Three options sit in front of you: install an MCP server, write a small script the agent calls via Bash, or just let the agent stitch together curl commands to hit an HTTP API directly.
The cost of choosing wrong is not "a few extra lines of code" — it is money burned on every single call afterward. Every invocation burns tokens, a badly written tool description makes the agent pick the wrong tool, and overly broad permissions let it dig up your production data. This guide skips MCP protocol details and gives you a directly usable decision framework: the real cost of each route, a three-step decision tree, five real-world scenarios, and a pre-build acceptance checklist.
The Real Cost of the Three Routes: Compare Bills, Not Features
Let's break the three routes apart — not by what they can do, but by what you pay:
Route 1: MCP server. A standard protocol — configure once, and every MCP-capable client (Claude Code, Cursor, Cherry Studio) can reuse it. The price: protocol overhead (one JSON-RPC round trip per call), a resident server process, and tool-description bloat — write a bad description and the agent picks the wrong tool or passes wrong arguments. It fits capabilities that are cross-client and long-lived.
Route 2: Custom script + CLI. The agent runs python scripts/notify.py --channel "#deploy" --text "deploy done" via Bash. Near-zero cost: no protocol overhead, debugging is just edit-and-rerun, and the agent can even read your source code to understand the semantics. The downside: it only serves the current project and client — switching agent clients means rewiring.
Route 3: Let the agent call HTTP directly. The agent curls or fetches the API itself. Zero upfront cost, usable today. The price: every API detail lives in the context — the agent hallucinates parameter names, auth tokens entering the context risk leaking, and it must "re-understand" the API docs on every call, burning the most tokens. Only fit for one-off, low-frequency, read-only calls.
A table to remember:
- Upfront dev cost: direct HTTP is cheapest, scripts are mid-range, MCP is highest (server code, schema, transport layer).
- Per-call token cost: direct HTTP is highest (API context on every call), scripts lowest (just a few flags), MCP in the middle.
- Cross-client reuse: only MCP works out of the box; scripts travel by copying files; direct HTTP starts over every time.
- Debugging difficulty: scripts are easiest (just print); MCP is hardest (client logs, server logs, and the protocol handshake — three places to check).
- Permission granularity: scripts are finest (constrain however you like); MCP depends on the server implementation; direct HTTP is coarsest (hand over the token, hand over everything).
The Three-Step Decision Tree: Pick a Direction in 30 Seconds
Ask three questions in order, and the answer points to the route by itself.
Step 1: How long will this capability live? One-off experiment — just curl, don't build anything. Throwaway work doesn't deserve a server. Used repeatedly through a project — write a script, 20 minutes of work that pays off across the whole project. Used across projects and clients long-term — only now is MCP worth considering.
Step 2: Who provides the capability? If there is a mature official or community MCP server (GitHub, Postgres, Slack, Notion, Playwright all have official or high-quality community implementations) — use it, don't reinvent the wheel; your first homemade version will never beat one that has survived a year of real-world edge cases. Only REST API docs exist — wrap it in a script CLI first, and consider MCP packaging after it runs smoothly. Internal systems, private protocols — scripts first, because requirements change fast and scripts are cheap to change.
Step 3: How much "understanding" does the agent need to use it correctly? Simple params, idempotent, read-only — exposing the raw capability is fine, any route works. Complex params, side effects, business judgment required — you must converge in the wrapper layer: collapse 20 parameters into 3, add double-confirmation for dangerous ops. Here scripts beat MCP for flexibility, because you can just edit code, while an MCP tool schema, once set, is expensive to change.
Compressed into one rule:
Trusted official MCP exists → use it; one-off use → curl; long-term cross-client use → write an MCP server; everything else → write a script CLI first. When the script gets reused by a third project, promote it to MCP.
Note that last line: a script is the draft of an MCP server. Standardize only after a script has been proven. Writing an MCP server from scratch usually means paying architecture tax for requirements that haven't happened yet.
Five Real-World Scenarios: From "What to Use" to "How to Write It"
Scenario 1: Let the agent query the production database (read-only). → Write a script. Don't use the Postgres MCP server — it's too open. The agent gets full SQL execution power, and you're relying on a prompt to restrain it to "read, don't write," which is like telling the security guard the vault code and asking him not to open it. The right approach: scripts/db_read.py --sql "..." with three hard rules baked into the code: only statements starting with SELECT (regex-reject INSERT/UPDATE/DELETE/DROP), an automatic LIMIT 200 appended, and a read-only account in the connection string. Permission convergence belongs in code, not in prompts.
Scenario 2: Let the agent send Slack notifications. → Use the official MCP. Slack has an official MCP server — auth, channel listing, and message formatting are all handled. Writing your own script is reinventing the wheel. Two pitfalls to dodge: give the bot token minimum scope (chat:write is enough, never admin); in the MCP config, expose only the two or three tools you need and disable the rest at the client config — tools the agent can't see, it can't misuse.
Scenario 3: Let the agent call the internal OA approval API (private, no official MCP). → Script first. Start with oa.py approve --id 12345 --comment "approved", and do three things in the script: check the ticket status before acting (idempotency, no double approvals), write one audit log line (who, when, which ticket), and support --dry-run so the agent shows its plan before acting. When a third project needs the same capability, package it as an internal MCP server — by then you know the real call patterns, so the schema design won't be guesswork.
Scenario 4: Let the agent batch-compress 200 images. → Neither. Don't wire up any capability. This is the most common mistake. If the task is a "deterministic batch pipeline," the right move is a script that runs it in one shot — the agent only triggers it, never executes item by item. Having the agent call a tool per image is a token black hole: hundreds of tokens of tool-call overhead per image, times 200, is a disaster. The test: deterministic steps + batch input → script it in one go; judgment needed per item → only then is per-item agent tool use worth it.
Scenario 5: Let the agent drive a browser for E2E verification. → Use the community MCP (Playwright MCP). Browser automation state is brutally complex: page contexts, element waits, screenshots, console logs — the community Playwright MCP has already survived those pitfalls. Rolling your own means handling CDP sessions and race conditions, an order of magnitude more expensive. Just point it at the staging environment only — never let the agent click around the production admin with browser automation.
The Six-Item Pre-Build Checklist: Whichever Route You Pick
Once the route is decided, run through these six before building. Don't wire anything up until all pass:
- Least privilege. Tokens get only the scopes they need; databases get read-only accounts; any production write needs a double-confirmation mechanism. Ask yourself first: if the agent got prompt-injected, what's the worst this permission could do?
- Write tool descriptions for the agent, not for humans. What matters most in a description isn't parameter docs — it's "when to use it and when NOT to." Good example:
Read-only queries against the production orders table. Not for report exports; aggregate first if over 1000 rows; do not call during peak hours (9:00-11:00).Bad example:Executes a SQL query; parameter is a sql string. - Idempotency and dry-run. Anything with side effects must support
--dry-run, letting the agent show "here's what I plan to do" for your confirmation before executing. Approvals, deploys, and deletes without dry-run are flying blind. - Converged output. Return only the fields the agent needs; truncate lists at 50 items by default and report the total (
{"total": 2317, "items": [...]}). Every one of those 2000 rows the agent doesn't need is money. - Readable failures. Non-zero exit code, and the first line is a human-readable reason — don't make the agent guess.
ERROR: order 12345 already approved (status=approved), no action neededbeats a raw traceback every time. - Audit logging. Who, when, what was called, with which parameters — one log line. When the agent breaks something, this is your only crime scene. Same for MCP servers: add a logging middleware in the server instead of relying on client-side records.
Two Counterintuitive Conclusions
First: MCP is not "more advanced" — it is "more expensive" standardization. The biggest pitfall in today's MCP ecosystem is writing a server for a one-off need in the name of "standardization," then paying more in maintenance than it ever returns. Standardization only pays back with enough reuse — the rule of thumb is 3: when the same capability serves three different projects or clients, MCP's math works out. Before that, scripts are always the better deal.
Second: the best agent tool is one where the agent doesn't need to understand the business. Push business judgment down into the wrapper layer: validate what should be validated, converge what should be converged, double-confirm what should be double-confirmed. The agent answers multiple-choice questions, never open-ended ones. In practice, error-rate drops from the tool layer beat prompt optimization by an order of magnitude — because an if statement in code never hallucinates.
Next time you wire an external capability into an agent, ask yourself three questions: how many times will it be used? who provides it? how much does the agent need to understand? The answer will point to the route by itself. And when you're unsure, remember the default answer is always: write the script first.
Related articles

Reported by InfoQ on October 3: a GPT-4 coding agent given CVE descriptions successfully exploited 87% of 15 test vulnerabilities, versus 7% without descriptions. rclone's author received 40+ security disclosures in a single month — more than the project's previous decade combined; QEMU has shortened its embargo period. The vulnerability disclosure timeline is collapsing under agent speed.

On October 7, OutSystems announced Agent Experience is generally available: its low-code platform is now open to any AI coding agent — Claude Code, Cursor, Codex, Kiro — with agents working at the design level, the platform generating code deterministically, and governance built in. This is the "vibe coding goes enterprise" playbook: taming shadow AI with a compliant path. But the 74% rework figure is vendor-survey data — discount it. The real bill is the hidden cost of platform lock-in.

Cloudflare's Birthday Week blog makes the case plainly: GitHub was designed for humans writing code; the agent era needs the collaboration layer reinvented. Artifacts enters open beta with a repo for every agent, plus a developer competition — $25,000 in credits for first place, deadline October 14. This is the first time a major infra vendor has put 'infrastructure for agents writing code' on the table as a public proposition.