The Token Cost Ledger: Know Your Per-User Margin Before You Price
Losing more money with every user is the most expensive bug in AI products. The arithmetic of per-user gross margin: why you must price to the P90 user, not the average; the photo app burned alive by its own virality; three cost levers and three margin red lines.

The most expensive bug in an AI product is often not in the code — it's on the pricing page. I've watched more than one indie developer ship a product with growing users and great word of mouth, then stare at the month-end cloud bill in disbelief: the more users, the more money lost. Traditional SaaS has near-zero marginal cost — one more user costs almost nothing. An AI product's marginal cost is tokens: real, hard COGS (cost of goods sold) that scales linearly with usage. Pricing without doing the token math is like opening a restaurant without knowing your food costs.
This piece skips the grand narratives and does three things: walks through the arithmetic of "per-user gross margin" by hand, tells the true story of a product burned alive by its own virality, and lays out the three cost-saving levers plus the margin red lines your pricing must respect. If you're building — or planning to build — anything with AI features, this might save you several months of tuition.
First, the Arithmetic: What Does Your "Free User" Really Burn?
Let's compute with illustrative prices (examples only — always check your model provider's official pricing): assume the flagship model charges $5 per million input tokens and $20 per million output tokens. Now assume your AI writing assistant's typical user has 30 conversations a month, averaging 8,000 input tokens (system prompt plus context) and 2,000 output tokens per session.
Cost per session = 8000/1M × 5 + 2000/1M × 20 = 0.04 + 0.04 = $0.08. Monthly cost = 0.08 × 30 = $2.40 per user per month. Price it at $9.90/month and gross margin looks like a healthy 75%, right?
Not so fast. The problem is that the "average user" is a statistical illusion. Real distributions are always long-tailed: the silent majority chats 3 times a month, while 10% of power users chat 300 times — 10× the average. A power user's monthly token cost is $24, a $14 loss per user. Worse, power users are usually your evangelists — you can't throttle them, and you can't raise their price.
So the correct way to do the math: always cost against the P90 user (the percentile heavier than 90% of users), never the average. Pricing to the average is the number-one cause of death for AI products. Pull your usage distribution and look at what P90 and P99 users burn per month in tokens — that number, not the average, is the true starting point for your pricing.
How do you pull a usage distribution? No fancy BI needed: export "user ID — daily token spend" to CSV each day and compute percentiles with a script at month's end. The key is starting on day one — most teams only think about distributions after the bill explodes, three months too late. Instrumentation is the foundation of token accounting.
One ledger entry most people miss: AI costs beyond tokens. Embeddings billed by volume, vector-database storage fees, failed retries (one network jitter means one retry means double cost), and "AI QA" — using a big model to audit a small model's output, a "QA tax" that often reaches 20% of total cost. Fold all of it into per-user margin math, then multiply by a 1.2 buffer. A ledger that counts only tokens never balances.
Real Case: How an AI Photo App Got Burned Alive by Going Viral
In early 2026, an AI old-photo restoration app (name withheld by request) got featured on an app store and gained 100,000 new users in a week. The founder posted a celebratory update: "We went viral." Three weeks later he deleted it — the month's cloud bill was over $40,000 against less than $6,000 in revenue.
The postmortem found three causes, all of them token math never done: first, the free tier had no hard cap. To climb the charts, free users got 20 restorations a day — "most people won't use that many anyway." Script kiddies ran hundreds per account per day. Second, everything ran on the flagship model. Whether a 200KB thumbnail preview or a 4K final output, every request hit the most expensive model — even though a small model's preview was visually indistinguishable. Third, zero caching of repeat work. The same old family photo forwarded around group chats got uploaded and restored by different users, each time billed at full inference cost.
How did he save it? Three moves, each one token-math-driven: the free tier became a hard-capped 5 per day (hit the cap and you wait or pay — enforced in code, not "we hope users behave"); previews were routed to a small model with only final HD output on the flagship, cutting per-request cost ~70%; uploaded images got hashed so identical photos returned cached results at zero marginal cost. Two months later, average monthly token cost per user fell from $3.10 to $0.70 and gross margin recovered above 80%. The founder's summary was brutal: "We weren't killed by competitors. We were burned alive by our own free users."
One more thing: how do you spot script abusers? Three signals — a single account's daily calls exceeding 5x P99, machine-regular request intervals, and using only free features while never visiting the pricing page. Hit two, and the account goes to manual review. Fifty lines of rules, worth $40,000.
If time could rewind, three "must-dos before launch" for that founder: first, load-test the free tier — simulate abusers with scripts and see whether the billing alarm or the growth spike fires first; second, canary the model routing — send 10% of traffic to the small model, compare retention and complaints, go full only if there's no difference; third, put a circuit breaker on the bill — auto-degrade (e.g., pause the free tier) when daily token spend crosses a threshold, instead of having a heart attack at month's end. Virality is luck; the ledger is skill.
The Three Cost-Saving Levers: Cache, Route, Slim Down
Whatever AI product you build, token savings always come down to these three levers, ordered by ROI:
- Lever one: caching (prompt caching). The highest-ROI optimization there is, bar none. The idea is simple: the repeated parts of each request (system prompt, fixed few-shot examples, frequently used document chunks) get cached by the provider and billed at cache rates — typically around a tenth of normal price. Hands-on tip: structure your system prompt as "static prefix + dynamic suffix"; the longer and more fixed the static part, the higher the cache hit rate. One RAG app turned its 20K-token system knowledge base into a cached prefix and input costs dropped over 60%. Caveat: caching isn't magic — highly dynamic conversations get poor hit rates. Measure before committing.
- Lever two: model routing. The core idea: don't use a cannon to kill a mosquito. Tier requests by difficulty: classification, first-draft summaries, format conversions — anything a small model does well enough — goes to a model 10× cheaper; only deep reasoning, complex code, and final quality gates get the flagship. An even better pattern is "small model drafts, big model reviews": the small model generates the draft, the flagship only verifies and fixes — the review's input includes the full draft, but the output (the expensive part) shrinks dramatically. Practical advice: run everything on the flagship for two weeks, sample 200 real requests, re-run them on the small model, have humans compare quality to find the "no quality loss" boundary, then write your routing rules.
- Lever three: slimming output. Output tokens usually cost 3–4× input tokens and are the easiest to waste. Three concrete moves: use structured output (JSON mode) instead of free text, with length caps on fields — "summarize in one sentence" saves half the tokens of "summarize this"; set a sane max_tokens cutoff on streaming output to stop the model's "politeness verbosity"; strip small talk from the system prompt (every "you are a helpful assistant…" burns money on every request). Don't underestimate this — one customer-service bot cut its monthly bill 30% with "keep answers under 150 words" alone.
Beyond the three levers, three classic "money-burning postures" — see how many you recognize: posture one: stuffing the entire user manual into the system prompt. Tens of thousands of static tokens re-sent on every request with no caching — that's throwing money into water. Posture two: using the flagship model for "Hello World"-tier tasks. Sentiment classification, keyword extraction, format conversion — small models are within 5% on quality at a tenth of the price. Posture three: infinite retries plus infinite context. Retry on failure, context only grows — one abusive user can burn a normal user's monthly budget in an hour. Cap retries and use sliding windows for context. Basics.
Three Margin Red Lines Your Pricing Must Respect
Saving is defense; pricing is offense. Three red lines, learned from watching too many cautionary tales:
- Red line one: price at least 3× marginal cost, and no less than 1.5× at P90 usage. Why 3×? Because beyond tokens you still pay for servers, payment processing (~3–5%), support, and refund leakage. Pricing 3× average cost leaves ~1.5× margin on P90 users — the safety buffer that keeps power users from losing you money. Below that multiple, your business model is fundamentally betting that "users won't use it much" — a bet you always lose.
- Red line two: the free tier needs a "hard cap," not "soft goodwill." "100 free tries for new users" is fine — but it must be enforced in code: over the cap means waiting till tomorrow or paying. Never use "we hope users behave" as cost control — the internet will teach you otherwise with scripts. A free tier's total cost belongs inside your acquisition budget: token cost per free user × expected free users = your true CAC. If that exceeds paid-channel CAC, cut the quota.
- Red line three: review the per-user token cost distribution monthly — not just the total bill. A 20% bill increase could mean more users (good) or one feature's token efficiency collapsing (bad). Build a simple dashboard: monthly token cost by user percentile, per-call cost by feature, cache hit rate. Any week where per-call cost jumps 15%+ week-over-week, investigate immediately — it's usually a prompt change that introduced redundant context.
Pricing in practice: cost profiles of three common models.
- Usage-based (credits): cost and revenue are naturally isomorphic — the safest. Price credits at 3x P90 cost, and fold credit-expiration policy into margin.
- Tiered monthly: the most common, and the most dangerous. Every tier needs a "fair use cap" enforced in code; the top tier's price must cover P99 users' cost, or the top tier fills up with heavy abusers.
- Freemium: the free tier is acquisition cost, not charity. Compute free-to-paid conversion, then true CAC = per-free-user cost x 100 / conversion rate. Over $50 means redesigning the free tier.
One last hidden killer: exchange rates and billing lag. For teams billed in USD but earning in other currencies, a 5% FX swing can eat an already-thin margin; usage-based bills lag 1-2 days behind reality, so by the time you see the spike it's too late. The fix: daily budget alerts on token spend (notify at 80%) and an FX buffer in pricing. The devil lives in details like these.
My take: token math isn't a tech problem, it's a business-model problem.
Many developers treat token cost as an optimization problem: switch to cheaper models, compress prompts, add caching. All valid — all tactics. The strategic question is: is your pricing model isomorphic to your cost structure? A product billed per API call should be priced per call (credits). A flat-rate unlimited plan must have hard caps and tiers — "unlimited + metered cost" is mathematically unsolvable.
The most expensive sentence of 2026 is "AI features are free." Free can be an acquisition strategy, but it must be accounted free: you know what each free user burns, what conversion rate breaks even, and where the total budget ceiling sits. Free without those three numbers isn't a growth strategy — it's suicide by burn rate.
One more judgment call: when per-user token cost falls three months straight, engineering optimization is tapped out — go optimize the business model instead: raise prices, add tiers, build the enterprise plan. If costs are still climbing, don't think about raising prices yet — swing the three levers another round. The slope of the cost curve decides whether you should talk engineering or talk business.
One thing you can do tonight: open your cloud bill, compute the last 30 days' average token cost per paying user and the P90 user's token cost, and divide each by your price. If the second number implies gross margin below 50%, your pricing page needs a rewrite — and the cheapest time to do it is while you're still small. The cheapest time to reprice is always right now — don't wait for the bill to teach you.
Related articles

Traffic is moving from the search box to the AI answer box. A vibe coder ships a product in a week — and nobody finds it. This guide turns the SEO fundamentals (sitemaps, JSON-LD, Core Web Vitals) and the new AI-discovery toolkit (llms.txt, per-page Markdown versions, FAQ schema, agent-readable pricing and API docs) into a shippable 30-day checklist. The core judgment: how well you document sets your product's ceiling in the agent economy.

The Zephos team planted 16 launch-killer bugs in Notely, an agent-built Next.js + Supabase + Stripe notes app — two payment-related: unsigned webhooks accepted, pro granted from a self-declared client_reference_id. This guide turns those traps into a playbook: webhook signature verification, a server-side single source of truth, the subscription state machine, test clocks, and a launch checklist. Money logic must be hand-written or audited line by line.

On October 3, 'Sites in ChatGPT' hit the HN front page with ~209 points and 218 comments. Not a launch — a reckoning: is prompt-to-URL a toy, a prototype host, or a productivity tool? The four debates, the doc-backed facts (D1/R2, sign-in, custom domains), and three verdicts for vibe coders.