Prompt caching is one of the most underrated optimizations in LLM applications.
It can make responses faster, cheaper, and more scalable.
And it’s surprisingly simple to implement.
A thread 🧵
This is going to change drastically.
Most teams pay frontier-LLM prices for what is really classification. Triage, routing, scoring.
Jev returns a typed answer with a confidence score in under a second mostly.
@HarveenChadha Maybe. But if we need increasingly more people to build harder evals every time models improve, that’s also evidence that static benchmarks are the wrong abstraction. The real challenge is measuring performance on dynamic, real-world tasks.
today we're launching @Palmier_io, a video editor Claude can edit.
use AI to edit, organize, and generate footage directly in the timeline.
finally, a video editor built for AI.
open-source. mac native. available now.
Introducing text-to-lottie: an open source skill and harness for generating production ready Lottie animations with codex/claude code.
$ npx skills add diffusionstudio/lottie
Prompts guide and repo in the comments.
7/ Important limitation:
Caches are temporary.
Claude’s cache currently lives for ~1 hour.
So prompt caching is optimized for:
• active sessions
• repeated near-term requests
• high-frequency applications
Not long-term storage.
6/ Very common patterns:
Cache:
• tool schemas
• system prompts
• long documents
• conversation history
Then, only send small incremental user changes in each request.
This gives massive savings.
5/ In Claude, caching is enabled using cache breakpoints.
You manually mark blocks that should be cached.
Example:
{
"type": "text",
"text": system_prompt,
"cache_control": {
"type": "ephemeral"
}
}
Everything before that breakpoint becomes cacheable.
4/ Prompt caching only works if the content is identical.
Even tiny changes can invalidate the cache.
Changing:
“Summarize this.” to “Please summarize this.” may force full reprocessing.
So stable prompts matter.
3/ Prompt caching solves this.
Instead of discarding preprocessing work, the model stores it temporarily in a cache.
Future requests can reuse that work instead of recomputing it.
Result:
✅ Lower latency
✅ Lower token processing cost
✅ Better throughput for repeated workflows
2/ This becomes expensive in some use cases.
Example:
You upload a large document and ask - summarize this, extract action items, rewrite etc.
Without caching, the model repeatedly reprocesses the same document every time.
1/ Normally, every LLM request starts from scratch, even if you resend the exact same content.
The model has to:
• tokenize the prompt
• build embeddings
• process context
• prepare attention states
Then generate the answer.
Prompt caching is one of the most underrated optimizations in LLM applications.
It can make responses faster, cheaper, and more scalable.
And it’s surprisingly simple to implement.
A thread 🧵
Found this strange rubbery/skin-like substance inside the Amul protein buttermilk. @Amul_Coop
This is unacceptable from a brand like Amul and needs urgent investigation and clarification. @fssaiindia
Order ID: OD337584252299328100
@Flipkart AC delivered on 21st. Installation falsely marked completed twice without technician visit. 6–7 days passed despite promised 24-hour installation.
Refund denied. Process full refund now or I’ll file a consumer court complaint.
@Flipkart
It’s becoming a ritual now, I place an order, wait endlessly, and then come here frustrated to post about your service @flipkartsupport@Flipkart
The promised installation date was 24th May, and there has still been no update till today. Extremely disappointing experience.