Six Months of Production LLM Bills: A Cost-Optimisation Postmortem

We shipped a customer-facing AI feature in January and the first invoice was twelve times higher than the back-of-napkin estimate. Here's the six-month story of how we cut inference spend by 87% without touching the model — caching, routing, prompt economy, and the one decision that saved more than every clever trick combined. We shipped a customer-facing AI feature in January and the first invoice was twelve times higher than the back-of-napkin estimate. Here's the six-month story of how we cut inference spend by 87% without touching the model — caching, routing, prompt economy, and the one decision that saved more than every clever trick combined.
TL;DR The single biggest win was never sending the same prompt prefix twice — prompt caching alone cut our bill by 64%. After that: semantic routing to smaller models (down another 18%), output token discipline (down 4%), and aggressive de-duplication at the request layer (down 1%). The model choice mattered less than every team predicted.
We shipped an AI-powered "summarise this conversation" feature in January 2026 for a B2B customer support platform. The PM had estimated $400/month at the projected usage. The first invoice was $4,860.
This is the story of how we got from there to $612/month at 3× the original projected traffic, without changing the underlying model or degrading the output. Six months, four engineers, two production incidents, and one architectural rewrite.
What we measured before changing anything
Rule one of cost optimisation: don't optimise what you can't see. Before we touched a line of code, we wired three dashboards:
- Spend per request type (summarise / classify / extract / draft-reply)
- Token distribution (p50, p90, p99 input + output) — the long tail is where the money goes
- Cache hit rate (we had no cache yet — this was the "how much could we win" projection)
Two days of data revealed two things:
- 23% of all requests had the exact same system prompt and context as another request that month — pure duplicate work
- The p99 output token count was 8× the p50: a few requests were running away with the budget
Both became targets.
1. Prompt caching: the single biggest lever
Every modern LLM provider now supports some form of prompt caching, and the discount is structural — typically 90% off the cached portion. For a workload where the system prompt + context is 4,000 tokens and the user query is 200 tokens, you can pay full price for 5% of the request instead of all of it.
The trick is structuring your prompts so the *prefix* is identical across requests. Provider-specific details:
| Provider | Cache mechanism | TTL | Discount |
|---|---|---|---|
| Anthropic Claude | Explicit cache_control markers on message blocks | 5 min (default) / 1 hour (extended) | 90% on cache hit |
| OpenAI GPT-4o | Automatic for prefixes ≥1024 tokens | ~5–10 min | 50% on cache hit |
| Google Gemini | Explicit context cache API | Configurable (1 min – 1 hour) | 75% on cache hit |
Anthropic Claude (where we got the biggest win)
// backend/functions/summarise-api/index.tsimport Anthropic from "@anthropic-ai/sdk";const client = new Anthropic();export async function summariseConversation(conversationText: string,customerInstructions: string,) {const response = await client.messages.create({model: "claude-sonnet-4-6",max_tokens: 1024,system: [{type: "text",text: SYSTEM_PROMPT, // ~2,400 tokens, identical every callcache_control: { type: "ephemeral" },},{type: "text",text: customerInstructions, // ~600 tokens, per-customer constantcache_control: { type: "ephemeral" },},],messages: [{role: "user",content: `Summarise this conversation:\n\n${conversationText}`,},],});return response.content[0].type === "text"? response.content[0].text: "";}
The two cache_control markers tell Claude: *"the first 3,000 tokens are identical to the previous call — re-use them at the cache rate"*. Cache hit rate on this workload landed at 94% after the first hour of warm-up.
Important: the cache key is byte-identical
This is the bit nobody mentions in the marketing pages. The cached prefix must be exactly the same bytes as the prior request. A single trailing whitespace, a re-ordered JSON field, a templating quirk that emits "Asia/Kolkata" instead of 'Asia/Kolkata' — and the cache misses.
// ❌ This will miss cache constantly — Date.now() changes every callconst systemPrompt = `You are an assistant. Current time: ${Date.now()}`;// ❌ This will also miss — JSON.stringify object key order is non-deterministicconst systemPrompt = `Schema: ${JSON.stringify(complexSchema)}`;// ✅ Stable prefix, dynamic suffixconst systemPrompt = STATIC_SYSTEM_PROMPT; // resolved once at bootconst userMessage = `Schema: ${stableStringify(complexSchema)}\n\n${userQuery}`;
Use a canonical JSON serializer (json-stable-stringify for Node, json.dumps(..., sort_keys=True) for Python) for anything templated into the cached portion. We lost two weeks of cache hits to this exact bug.
2. Semantic routing: send the easy ones to a smaller model
About 60% of "summarise this conversation" requests were under 300 tokens of input. For those, a 4B parameter model produced output indistinguishable (in our blind eval) from Claude Sonnet 4.6. The Sonnet call costs ~50× more.
The routing logic is unglamorous:
# backend/routing/router.pydef route_request(conversation: str, complexity_hints: dict) -> str:token_count = estimate_tokens(conversation)has_code = bool(re.search(r"```|`\w+`", conversation))has_multilang = detect_multilingual(conversation)customer_priority = complexity_hints.get("tier") == "enterprise"if customer_priority:return "claude-sonnet-4-6" # never cheap out on enterpriseif has_code or has_multilang:return "claude-sonnet-4-6" # known weak spots for small modelsif token_count > 800:return "claude-sonnet-4-6" # long context handlingreturn "claude-haiku-4-5-20251001" # fast, cheap, fine for the rest
This took the average cost per request from $0.043 to $0.011 without a measurable quality regression. We A/B tested for two weeks before flipping the default — the null hypothesis discipline matters here, because "smaller model fine for easy cases" is exactly the kind of intuition that produces silent regressions.
Hard-won lesson: instrument *quality* the same week you instrument cost. We added a sampling pipeline that scored 1% of routed-down responses against the bigger model's output. When quality slipped below 92% agreement, an alert paged us. It fired exactly once in six months — for a customer who started uploading PDF attachments inline as base64, blowing the token-count routing heuristic.
3. Output token discipline: the one-line fix that saved $400/month
Our max_tokens was 4,096. The p99 of actual output was 612 tokens. We were paying for the budget, not the output — except that LLMs don't bill that way; we were paying for the *output*. So what was the problem?
The problem was a class of buggy prompts that asked the model to "produce a thorough summary". A few customers' contexts triggered the model to actually take that seriously and produce 3,000-token essays. Those tail requests were 6× the cost of the median.
Two fixes:
- **Cap
max_tokensper use case.** Summary = 600. Classification = 50. Draft reply = 1,200. Anything more is almost certainly the model going off the rails. - Add explicit "be brief" instructions to the prompts that need it. *"Respond in no more than 200 words"* changes p99 output dramatically.
// Beforeconst response = await client.messages.create({model: "claude-sonnet-4-6",max_tokens: 4096, // ← runaway potentialsystem: SYSTEM_PROMPT,messages: [{ role: "user", content: query }],});// Afterconst response = await client.messages.create({model: "claude-sonnet-4-6",max_tokens: 600, // ← boundedsystem: SYSTEM_PROMPT_WITH_BREVITY_RULES, // includes "Respond in ≤200 words"messages: [{ role: "user", content: query }],});
This is the cheapest 5 minutes of work in this entire post.
4. The de-duplication layer (sweep up the rest)
Two requests with the exact same input within 60 seconds is almost always either:
- A user double-clicking the button
- A retry from a flaky network
- A polling job racing itself
We added a Redis-backed dedupe layer in front of the inference call:
// backend/functions/summarise-api/dedupe.tsimport { createHash } from "crypto";import Redis from "ioredis";const redis = new Redis(process.env.REDIS_URL!);export async function withDedupe<T>(cacheKey: string,ttlSeconds: number,produce: () => Promise<T>,): Promise<T> {const lockKey = `inflight:${cacheKey}`;const resultKey = `result:${cacheKey}`;// Already have a cached result?const cached = await redis.get(resultKey);if (cached) return JSON.parse(cached) as T;// Try to claim the lock. NX = only set if not exists.const claimed = await redis.set(lockKey, "1", "EX", 30, "NX");if (!claimed) {// Another request is producing the result — wait briefly and retryfor (let i = 0; i < 20; i++) {await new Promise((r) => setTimeout(r, 250));const r = await redis.get(resultKey);if (r) return JSON.parse(r) as T;}// Lock holder died — fall through and produce ourselves}try {const result = await produce();await redis.set(resultKey, JSON.stringify(result), "EX", ttlSeconds);return result;} finally {await redis.del(lockKey);}}
The hash key we use is sha256(model + systemPromptHash + userQuery) — note the *hash* of the system prompt, not the full thing, so the key stays a sane length.
This eliminated another 11% of total spend. The customer-facing latency benefit (sub-millisecond second-hit response) was an unexpected bonus.
5. The architectural decision that mattered more than every trick combined
Here's the one nobody wants to hear: stop running every request through an LLM at all.
A meaningful share of "summarise this conversation" requests were on conversations under 300 characters. *We were paying an LLM to produce a summary of two messages.* For these, the right answer is the obvious one: take the first message and return it, optionally truncated to 200 chars.
// backend/functions/summarise-api/index.tsexport async function summariseConversation(conv: Conversation) {// Cheap path: short conversations don't need an LLM at allif (conv.messages.length <= 2 && totalChars(conv.messages) < 300) {return truncate(conv.messages[0].text, 200);}// Medium path: route by complexity (see lesson 2)const model = routeByComplexity(conv);// Expensive path: cached, deduped, token-disciplined LLM callreturn withDedupe(cacheKeyFor(conv, model),300,() => callLLM(conv, model),);}
That single guard at the top eliminated 38% of inference calls outright. They contributed almost no value to the customer experience because there was nothing to summarise.
The takeaway, generalised: most "AI features" are 30–50% LLM problems and 50–70% deterministic-logic problems. Build the deterministic logic first; reach for the LLM only when the deterministic logic gives up.
6. The cost breakdown, six months in
Here's the actual journey:
| Month | Monthly cost | What changed |
|---|---|---|
| January | $4,860 | Initial launch, no optimisation |
| February | $4,720 | Output max_tokens capped, brevity prompts added |
| March | $2,890 | Anthropic prompt caching shipped |
| April | $1,640 | Semantic router live; ~60% of traffic on smaller model |
| May | $890 | Redis dedupe layer; deterministic short-conversation guard |
| June | $612 | Steady state at 3× the original projected traffic |
Net: 87% reduction at 3× the volume. The compounding effect of stacking these is what gets you there — none of them individually is a 10×.
7. What we'd have done differently
If we were starting this feature today, on day one we'd:
- Build the cost dashboard before the feature. The PM estimate is wrong; you need the real number on day two, not day forty.
- Structure the prompt for caching from the first deploy. Static prefix, dynamic suffix, canonical serialisation. Retrofitting this took us a fortnight of careful refactoring.
- **Set
max_tokensto the realistic p99, not the model max.** Nothing good comes from leaving the throttle wide open. - Build the deterministic fallback before the LLM call. "Does this even need a model?" is the single most cost-effective question in AI engineering.
- Wire a quality eval the same week you wire a cost eval. A regression in either is a production incident.
8. The four traps we fell into
❌ Optimising the model choice before the prompt shape
We spent two weeks A/B testing Sonnet vs Opus vs GPT-4o before we noticed the cache hit rate was 0%. The 90% cache discount dwarfs every model-choice delta. Get caching working first, then revisit the model selection — the cost economics change completely.
❌ Treating the LLM call as the only knob
The router, dedupe layer, and deterministic guard are *system design*, not LLM tuning. Most teams under-invest there because it's not glamorous.
❌ "We'll optimise after PMF"
PMF and optimisation are the same conversation when a customer-facing feature costs $0.05 per use. You will not get a chance to revisit this — by month three the cost line is a board-meeting topic and "we're working on it" is not an answer.
❌ Skipping the quality eval
Cost goes down. Quality goes down too. You don't notice until a sales call where the customer asks "did you guys downgrade the model?". Build the eval; sample 1% of traffic against the better-but-pricier baseline; alert when agreement drops.
9. Where this stops scaling
These techniques carry you a long way. Where they stop:
- Highly personalised prompts: if every customer's system prompt is genuinely different, your cache hit rate is structurally limited. Split the cache surface by customer tier.
- Vision-heavy workloads: image tokens behave differently across providers and the routing heuristics need adjustment. Treat as a separate optimisation track.
- Real-time streaming: caching still works but the latency math is different — sub-second first-token matters more than dollars per million.
When you hit one of these, the next pass is usually fine-tuning a smaller model on your own data. That's a different post.
Closing — the meta-pattern
LLM features have a curious property: they look like AI problems but they're mostly *engineering* problems. The model picks up the last 20% of the work; the first 80% is system design, prompt economy, and operational discipline. Teams that frame the cost problem that way ship sustainable AI features. Teams that don't end up rewriting from scratch at month nine when the CFO sends a screenshot.
The trick is not finding the magic prompt or picking the perfect model. It is treating LLM inference like any other expensive backend dependency — cache aggressively, measure relentlessly, dedupe everything, and remember that the cheapest call is the one you didn't make.
If you'd like a second pair of eyes on your LLM cost curve — whether you're three months in and worried, or pre-launch and hoping to skip our mistakes — reach out to OrbitNexa and we'll dig through the numbers with you.
About the author
Surya Kenguva leads AI engineering at OrbitNexa. He has shipped customer-facing LLM features for B2B SaaS platforms in fintech, customer support, and developer tooling. You can find him on LinkedIn or on the OrbitNexa blog.
Further reading
- Anthropic Prompt Caching docs — the canonical reference for the discount mechanics
- OpenAI Prompt Caching guide — different model, same principles
- *Designing Machine Learning Systems* — Chip Huyen (chapters 8 and 9 on cost + monitoring in particular)
Got a different cost-cutting story? Reply to this email — a real human reads every response, and the best ones land in our next post.