The most expensive line in an AI app is often the one that stamps today's date into the instructions you send the model.
It looks like nothing. It can turn a 90% discount into full price on every call.
Every call starts from scratch. The model rereads everything you send from the first character, instructions, descriptions of the tools it can call, the whole conversation, and you pay for all of it again.
Prompt caching changes that. You mark a spot in the prompt (a "cache breakpoint"), and everything above the mark gets stored in already-processed form for five minutes, a clock that resets free every time it is used.
On Anthropic's pricing the first send costs 25% extra, and every later request that opens with that exact same text reads the stored copy at 90% off.
The catch is the word exact. The match is character for character, starting from the very top. Put anything that changes above the mark, a date, a live user ID, and it fails: everything below gets reprocessed at full price. Nothing errors out. The discount just never shows up in the bill.
So the whole discipline is order. Stable things first, changing things last: instructions, tool definitions and reference documents at the top, the new message at the bottom.
That is why it is worth an afternoon. An agent, a model that keeps calling tools in a loop, resends that same block every single turn, so getting the order right is usually a bigger lever than switching to a cheaper model, and it changes nothing about what the model says.
Quick check before you scroll: An agent has a 15 tool definition schema that never changes, followed by a conversation history that grows every turn. Where should the cache breakpoint go, and why?
Full breakdown + the answer: frankduah.me/learnings/2026-09-19-cost-optimization-for-llm-powered-products
New here? I post a bite-size AI / ML concept like this every day. Follow me for the daily drop, and it compounds fast. Why I do it: https://lnkd.in/gK8knHDH
#PromptCaching #LLMOps #PromptEngineering #AI #LLM #AIAgents #MachineLearning
The answer
Right after the tool schema, before the conversation history. That schema is identical on every call, so it caches cleanly and gets read at a 90% discount each turn, while the growing history stays outside the cached prefix where it belongs.