AI Engineering·Jun 1, 2026·12 min read
Six Months of Production LLM Bills: A Cost-Optimisation Postmortem
We shipped a customer-facing AI feature in January and the first invoice was twelve times higher than the back-of-napkin estimate. Here's the six-month story of how we cut inference spend by 87% without touching the model — caching, routing, prompt economy, and the one decision that saved more than every clever trick combined. We shipped a customer-facing AI feature in January and the first invoice was twelve times higher than the back-of-napkin estimate. Here's the six-month story of how we cut inference spend by 87% without touching the model — caching, routing, prompt economy, and the one decision that saved more than every clever trick combined.
Read the post