How do you cut LLM cost by caching stable parts? #
Long builds repeat themselves. The system instructions barely change between turns. Project conventions sit still. Reference files persist. Only the actual request is new each time. Prompt caching exploits exactly that: providers store the processed form of repeated prefixes so follow-ups skip recomputation, cutting cost and latency together.
Think of a restaurant kitchen during rush. Prep stays ready in its stations, knives sharp, stocks warm. Each ticket only fires the new dish. A kitchen that re-chopped every onion per order would collapse by 8pm. Most long-context setups are that kitchen, reprocessing the same context every turn. This playbook fixes the prep.
Three moves: maximize cache hits through stable ordering, keep working context under pricing tier boundaries, and route each task to the cheapest model that handles it well.
Try it: BYOB Credit Estimator — which estimates build cost before you tune caching.
How do you structure prompts for cache hits? #
Caches match on exact prefixes. Reorderings, timestamp injections, or reshuffled file dumps break the match and force full recomputation. Anthropic's prompt caching docs describe the mechanics: mark reusable blocks, pay a write premium once, then read from cache far cheaper, with 5-minute lifetimes by default and 1-hour options for longer sessions. OpenAI's caching guide tells the same story with its own numbers: cache writes cost extra, reads run around a tenth of input rates, and growing prefixes reward stable ordering.
The cache-friendly layout hardly varies: fixed system instructions first, stable project context in deterministic order, retrieved reference material next, the changing user request last. Pin tool and file ordering instead of appending newest-first. Keep clocks and random IDs out of cached regions. Split volatile chatter from the large stable context so small talk never invalidates the big prefix.
When a session compacts or summarizes, re-establish the same canonical prefix order in the new session instead of inventing a fresh layout. Compaction rewrites history; discipline rewrites it into the same shape. Measure hit rate alongside spend. A falling hit rate with steady usage means ordering drifted, not that prices rose.
{
"system": "Stable instructions, pinned tools, project conventions.",
"context": ["ARCHITECTURE.md", "STYLEGUIDE.md", "api-reference.md"],
"request": "Add retry with backoff to the checkout action."
}Stable keys up top, volatile request at the bottom. Boring. Cacheable. Cheap.
How do you design around context tiers? #
Long contexts price in tiers: past a threshold, per-token rates step up. Two habits keep builds under control.
Budget context per phase. Planning, implementation, and review each get a working set instead of one ever-growing thread that drags the whole project through every turn. A review pass rarely needs the full excavation history. Give it the diff, the plan, and the relevant files.
Reset or compact before crossing into a higher tier for work that does not need full history. Starting a fresh scoped session with a clean handoff beats dragging a bloated thread across a pricing cliff for a small follow-up. Handoffs should carry decisions and open questions, not transcripts.
Track cost per shipped feature, not raw token totals. A refactor burning twice the tokens but landing in one pass beats three cheap attempts that drift and need rescue. Report the number where decisions happen: per feature, per phase, per week. Token totals impress nobody and inform nothing. Cost per outcome moves roadmaps. Tier awareness informs session boundaries the way fuel awareness informs pit stops: routine, unglamorous, decisive.
| Lever | Action | Saves where | Watch out |
|---|---|---|---|
| Prefix caching | Stable order, challenge last | Recompute on repeats | Volatile inside cache |
| Tier awareness | Compact before cliff | Tier price jumps | Bloated handoff |
| Model routing | Small model for chores | Flagship overuse | Wrong mind for hard task |
| Hit rate metric | Track per week | Drift detection | Token totals alone |
How do you route models by task size? #
Not every turn needs the flagship. Classification, extraction, formatting checks, and scoped edits run well on smaller, cheaper models. Architecture decisions, cross-file refactors, and ambiguous debugging deserve the strongest reasoning available.
A simple router works: default to the capable mid-tier model for build turns, escalate to flagship when output quality drops twice on the same task, drop to the small model for mechanical transforms with a verifier watching. Keep one model per session for consistency, and compare candidates on a stable checkpoint rather than mid-thread where context differs.
The failure mode is status spending: flagship rates for chores. It feels safe and bills dangerously. Match the mind to the mission. You would not hire a principal architect to rename variables, at least not twice.
How do you measure and then tighten? #
Instrument three numbers weekly. Cache hit rate on repeated prefixes. Cost per shipped feature by phase. Rerun rate per feature.
Falling hit rates point at ordering drift: something volatile crept into the cached region. Rising cost per feature with flat output points at tier creep or model overuse: threads too long, minds too grand. High reruns point at prompt quality, not pricing: unclear asks produce iterated guesses regardless of caching. One concrete habit closes the loop: after each shipped feature, note which turn first went wrong and why. Patterns emerge fast. Vague verbs, missing examples, unstated constraints. Fix the pattern once and every future feature gets cheaper. The cheapest token is the turn you never needed.
OpenAI's evals guide frames the deeper version of this discipline: define tasks, run test inputs, analyze, iterate. Even lightweight evals, a checklist of known-good behaviors per feature, separate "the model regressed" from "the prompt was vague." Measure the work alongside the tokens.
Caching, tiers, and routing compound. Stable prefixes cut recomputation. Tier-aware sessions dodge cliff pricing. Right-sized models stop paying flagship rates for mechanical work. None of them substitute for clear prompts. Together they make long-context building sustainably affordable, which is the whole point: prep stays warm, tickets fly, kitchen survives the rush.
Track context like a budget, not a mystery
What are the trade-offs? #
Stable prefix first plus tier aware sessions plus routing by difficulty wins when context runs long. Anthropic documents 5 minute and 1 hour TTLs with cache reads far cheaper than fresh input, and OpenAI prices reusable prefix reads near a tenth of input rates.
| Pick caching plus routing when | Pick flagship full context when |
|---|---|
| Instructions and references stay fixed across turns | Architecture, cross file refactors, or ambiguous bugs need full reasoning |
| Small tasks like extraction and formatting dominate | One wrong answer costs more than many cached ones |
| Hit rate and cost per shipped feature get measured weekly | The session is short enough that setup beats savings |
Caching loses when order churns. Any reorder or timestamp inside the cached region breaks the match. Pick the alternative when difficulty spikes: send the hard job to the strongest model with clean scope.
Who this is for (and who should skip it)? #
This playbook helps teams running long context builds where repeated prefixes and tier pricing shape the bill. If you track cache hit rate and cost per shipped feature, the habits here compound.
Skip it if your sessions stay short and token spend is low. Keep prompts stable and routing simple until tier cliffs or rerun costs show up in weekly numbers.
One limit to know. Reordered context, timestamps, or volatile IDs inside cached regions break the match and force full price calls. Short sessions with low spend should keep simple routing until tier cliffs appear in weekly numbers.
- Best for developers cutting long-context bills with stable prompt ordering and cache hits.
- Best for startups designing around context tiers and routing tasks to cheaper models.
- Best for agencies running repeated builds where system instructions rarely change.
What we learned building this? #
BYOB tracks context token use per session and documents pricing where Pro and Max tiers set credit bundles and monthly economics, and in the BYOB context token tracking guide which lists slots for system, context, and request. Builds that keep stable prefix order and budget context per phase hit cache more and avoid dragging full history across tier thresholds. When hit rate falls with steady usage, ordering drifted, not pricing, which matches the caching docs cited in the post.