AI Engineering

Six Months of Production LLM Bills: A Cost-Optimisation Postmortem

SV
Surya
Jun 1, 2026
12 min read
Six Months of Production LLM Bills: A Cost-Optimisation Postmortem

We shipped a customer-facing AI feature in January and the first invoice was twelve times higher than the back-of-napkin estimate. Here's the six-month story of how we cut inference spend by 87% without touching the model — caching, routing, prompt economy, and the one decision that saved more than every clever trick combined. We shipped a customer-facing AI feature in January and the first invoice was twelve times higher than the back-of-napkin estimate. Here's the six-month story of how we cut inference spend by 87% without touching the model — caching, routing, prompt economy, and the one decision that saved more than every clever trick combined.

TL;DR The single biggest win was never sending the same prompt prefix twice — prompt caching alone cut our bill by 64%. After that: semantic routing to smaller models (down another 18%), output token discipline (down 4%), and aggressive de-duplication at the request layer (down 1%). The model choice mattered less than every team predicted.

We shipped an AI-powered "summarise this conversation" feature in January 2026 for a B2B customer support platform. The PM had estimated $400/month at the projected usage. The first invoice was $4,860.

This is the story of how we got from there to $612/month at 3× the original projected traffic, without changing the underlying model or degrading the output. Six months, four engineers, two production incidents, and one architectural rewrite.

What we measured before changing anything

Rule one of cost optimisation: don't optimise what you can't see. Before we touched a line of code, we wired three dashboards:

  1. Spend per request type (summarise / classify / extract / draft-reply)
  2. Token distribution (p50, p90, p99 input + output) — the long tail is where the money goes
  3. Cache hit rate (we had no cache yet — this was the "how much could we win" projection)

Two days of data revealed two things:

  • 23% of all requests had the exact same system prompt and context as another request that month — pure duplicate work
  • The p99 output token count was 8× the p50: a few requests were running away with the budget

Both became targets.

1. Prompt caching: the single biggest lever

Every modern LLM provider now supports some form of prompt caching, and the discount is structural — typically 90% off the cached portion. For a workload where the system prompt + context is 4,000 tokens and the user query is 200 tokens, you can pay full price for 5% of the request instead of all of it.

The trick is structuring your prompts so the *prefix* is identical across requests. Provider-specific details:

ProviderCache mechanismTTLDiscount
Anthropic ClaudeExplicit cache_control markers on message blocks5 min (default) / 1 hour (extended)90% on cache hit
OpenAI GPT-4oAutomatic for prefixes ≥1024 tokens~5–10 min50% on cache hit
Google GeminiExplicit context cache APIConfigurable (1 min – 1 hour)75% on cache hit

Anthropic Claude (where we got the biggest win)

typescript
// backend/functions/summarise-api/index.ts
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic();
export async function summariseConversation(
conversationText: string,
customerInstructions: string,
) {
const response = await client.messages.create({
model: "claude-sonnet-4-6",
max_tokens: 1024,
system: [
{
type: "text",
text: SYSTEM_PROMPT, // ~2,400 tokens, identical every call
cache_control: { type: "ephemeral" },
},
{
type: "text",
text: customerInstructions, // ~600 tokens, per-customer constant
cache_control: { type: "ephemeral" },
},
],
messages: [
{
role: "user",
content: `Summarise this conversation:\n\n${conversationText}`,
},
],
});
return response.content[0].type === "text"
? response.content[0].text
: "";
}

The two cache_control markers tell Claude: *"the first 3,000 tokens are identical to the previous call — re-use them at the cache rate"*. Cache hit rate on this workload landed at 94% after the first hour of warm-up.

Important: the cache key is byte-identical

This is the bit nobody mentions in the marketing pages. The cached prefix must be exactly the same bytes as the prior request. A single trailing whitespace, a re-ordered JSON field, a templating quirk that emits "Asia/Kolkata" instead of 'Asia/Kolkata' — and the cache misses.

typescript
// ❌ This will miss cache constantly — Date.now() changes every call
const systemPrompt = `You are an assistant. Current time: ${Date.now()}`;
// ❌ This will also miss — JSON.stringify object key order is non-deterministic
const systemPrompt = `Schema: ${JSON.stringify(complexSchema)}`;
// ✅ Stable prefix, dynamic suffix
const systemPrompt = STATIC_SYSTEM_PROMPT; // resolved once at boot
const userMessage = `Schema: ${stableStringify(complexSchema)}\n\n${userQuery}`;

Use a canonical JSON serializer (json-stable-stringify for Node, json.dumps(..., sort_keys=True) for Python) for anything templated into the cached portion. We lost two weeks of cache hits to this exact bug.

2. Semantic routing: send the easy ones to a smaller model

About 60% of "summarise this conversation" requests were under 300 tokens of input. For those, a 4B parameter model produced output indistinguishable (in our blind eval) from Claude Sonnet 4.6. The Sonnet call costs ~50× more.

The routing logic is unglamorous:

python
# backend/routing/router.py
def route_request(conversation: str, complexity_hints: dict) -> str:
token_count = estimate_tokens(conversation)
has_code = bool(re.search(r"```|`\w+`", conversation))
has_multilang = detect_multilingual(conversation)
customer_priority = complexity_hints.get("tier") == "enterprise"
if customer_priority:
return "claude-sonnet-4-6" # never cheap out on enterprise
if has_code or has_multilang:
return "claude-sonnet-4-6" # known weak spots for small models
if token_count > 800:
return "claude-sonnet-4-6" # long context handling
return "claude-haiku-4-5-20251001" # fast, cheap, fine for the rest

This took the average cost per request from $0.043 to $0.011 without a measurable quality regression. We A/B tested for two weeks before flipping the default — the null hypothesis discipline matters here, because "smaller model fine for easy cases" is exactly the kind of intuition that produces silent regressions.

Hard-won lesson: instrument *quality* the same week you instrument cost. We added a sampling pipeline that scored 1% of routed-down responses against the bigger model's output. When quality slipped below 92% agreement, an alert paged us. It fired exactly once in six months — for a customer who started uploading PDF attachments inline as base64, blowing the token-count routing heuristic.

3. Output token discipline: the one-line fix that saved $400/month

Our max_tokens was 4,096. The p99 of actual output was 612 tokens. We were paying for the budget, not the output — except that LLMs don't bill that way; we were paying for the *output*. So what was the problem?

The problem was a class of buggy prompts that asked the model to "produce a thorough summary". A few customers' contexts triggered the model to actually take that seriously and produce 3,000-token essays. Those tail requests were 6× the cost of the median.

Two fixes:

  1. **Cap max_tokens per use case.** Summary = 600. Classification = 50. Draft reply = 1,200. Anything more is almost certainly the model going off the rails.
  2. Add explicit "be brief" instructions to the prompts that need it. *"Respond in no more than 200 words"* changes p99 output dramatically.
typescript
// Before
const response = await client.messages.create({
model: "claude-sonnet-4-6",
max_tokens: 4096, // ← runaway potential
system: SYSTEM_PROMPT,
messages: [{ role: "user", content: query }],
});
// After
const response = await client.messages.create({
model: "claude-sonnet-4-6",
max_tokens: 600, // ← bounded
system: SYSTEM_PROMPT_WITH_BREVITY_RULES, // includes "Respond in ≤200 words"
messages: [{ role: "user", content: query }],
});

This is the cheapest 5 minutes of work in this entire post.

4. The de-duplication layer (sweep up the rest)

Two requests with the exact same input within 60 seconds is almost always either:

  • A user double-clicking the button
  • A retry from a flaky network
  • A polling job racing itself

We added a Redis-backed dedupe layer in front of the inference call:

typescript
// backend/functions/summarise-api/dedupe.ts
import { createHash } from "crypto";
import Redis from "ioredis";
const redis = new Redis(process.env.REDIS_URL!);
export async function withDedupe<T>(
cacheKey: string,
ttlSeconds: number,
produce: () => Promise<T>,
): Promise<T> {
const lockKey = `inflight:${cacheKey}`;
const resultKey = `result:${cacheKey}`;
// Already have a cached result?
const cached = await redis.get(resultKey);
if (cached) return JSON.parse(cached) as T;
// Try to claim the lock. NX = only set if not exists.
const claimed = await redis.set(lockKey, "1", "EX", 30, "NX");
if (!claimed) {
// Another request is producing the result — wait briefly and retry
for (let i = 0; i < 20; i++) {
await new Promise((r) => setTimeout(r, 250));
const r = await redis.get(resultKey);
if (r) return JSON.parse(r) as T;
}
// Lock holder died — fall through and produce ourselves
}
try {
const result = await produce();
await redis.set(resultKey, JSON.stringify(result), "EX", ttlSeconds);
return result;
} finally {
await redis.del(lockKey);
}
}

The hash key we use is sha256(model + systemPromptHash + userQuery) — note the *hash* of the system prompt, not the full thing, so the key stays a sane length.

This eliminated another 11% of total spend. The customer-facing latency benefit (sub-millisecond second-hit response) was an unexpected bonus.

5. The architectural decision that mattered more than every trick combined

Here's the one nobody wants to hear: stop running every request through an LLM at all.

A meaningful share of "summarise this conversation" requests were on conversations under 300 characters. *We were paying an LLM to produce a summary of two messages.* For these, the right answer is the obvious one: take the first message and return it, optionally truncated to 200 chars.

typescript
// backend/functions/summarise-api/index.ts
export async function summariseConversation(conv: Conversation) {
// Cheap path: short conversations don't need an LLM at all
if (conv.messages.length <= 2 && totalChars(conv.messages) < 300) {
return truncate(conv.messages[0].text, 200);
}
// Medium path: route by complexity (see lesson 2)
const model = routeByComplexity(conv);
// Expensive path: cached, deduped, token-disciplined LLM call
return withDedupe(
cacheKeyFor(conv, model),
300,
() => callLLM(conv, model),
);
}

That single guard at the top eliminated 38% of inference calls outright. They contributed almost no value to the customer experience because there was nothing to summarise.

The takeaway, generalised: most "AI features" are 30–50% LLM problems and 50–70% deterministic-logic problems. Build the deterministic logic first; reach for the LLM only when the deterministic logic gives up.

6. The cost breakdown, six months in

Here's the actual journey:

MonthMonthly costWhat changed
January$4,860Initial launch, no optimisation
February$4,720Output max_tokens capped, brevity prompts added
March$2,890Anthropic prompt caching shipped
April$1,640Semantic router live; ~60% of traffic on smaller model
May$890Redis dedupe layer; deterministic short-conversation guard
June$612Steady state at 3× the original projected traffic

Net: 87% reduction at 3× the volume. The compounding effect of stacking these is what gets you there — none of them individually is a 10×.

7. What we'd have done differently

If we were starting this feature today, on day one we'd:

  • Build the cost dashboard before the feature. The PM estimate is wrong; you need the real number on day two, not day forty.
  • Structure the prompt for caching from the first deploy. Static prefix, dynamic suffix, canonical serialisation. Retrofitting this took us a fortnight of careful refactoring.
  • **Set max_tokens to the realistic p99, not the model max.** Nothing good comes from leaving the throttle wide open.
  • Build the deterministic fallback before the LLM call. "Does this even need a model?" is the single most cost-effective question in AI engineering.
  • Wire a quality eval the same week you wire a cost eval. A regression in either is a production incident.

8. The four traps we fell into

❌ Optimising the model choice before the prompt shape

We spent two weeks A/B testing Sonnet vs Opus vs GPT-4o before we noticed the cache hit rate was 0%. The 90% cache discount dwarfs every model-choice delta. Get caching working first, then revisit the model selection — the cost economics change completely.

❌ Treating the LLM call as the only knob

The router, dedupe layer, and deterministic guard are *system design*, not LLM tuning. Most teams under-invest there because it's not glamorous.

❌ "We'll optimise after PMF"

PMF and optimisation are the same conversation when a customer-facing feature costs $0.05 per use. You will not get a chance to revisit this — by month three the cost line is a board-meeting topic and "we're working on it" is not an answer.

❌ Skipping the quality eval

Cost goes down. Quality goes down too. You don't notice until a sales call where the customer asks "did you guys downgrade the model?". Build the eval; sample 1% of traffic against the better-but-pricier baseline; alert when agreement drops.

9. Where this stops scaling

These techniques carry you a long way. Where they stop:

  • Highly personalised prompts: if every customer's system prompt is genuinely different, your cache hit rate is structurally limited. Split the cache surface by customer tier.
  • Vision-heavy workloads: image tokens behave differently across providers and the routing heuristics need adjustment. Treat as a separate optimisation track.
  • Real-time streaming: caching still works but the latency math is different — sub-second first-token matters more than dollars per million.

When you hit one of these, the next pass is usually fine-tuning a smaller model on your own data. That's a different post.

Closing — the meta-pattern

LLM features have a curious property: they look like AI problems but they're mostly *engineering* problems. The model picks up the last 20% of the work; the first 80% is system design, prompt economy, and operational discipline. Teams that frame the cost problem that way ship sustainable AI features. Teams that don't end up rewriting from scratch at month nine when the CFO sends a screenshot.

The trick is not finding the magic prompt or picking the perfect model. It is treating LLM inference like any other expensive backend dependency — cache aggressively, measure relentlessly, dedupe everything, and remember that the cheapest call is the one you didn't make.

If you'd like a second pair of eyes on your LLM cost curve — whether you're three months in and worried, or pre-launch and hoping to skip our mistakes — reach out to OrbitNexa and we'll dig through the numbers with you.

About the author

Surya Kenguva leads AI engineering at OrbitNexa. He has shipped customer-facing LLM features for B2B SaaS platforms in fintech, customer support, and developer tooling. You can find him on LinkedIn or on the OrbitNexa blog.

Further reading

Got a different cost-cutting story? Reply to this email — a real human reads every response, and the best ones land in our next post.