you added prompt caching, watched the bill drop on day one,
and never looked at it again
on the current opus, input is four dollars a million tokens,
cache read is twenty cents
one twentieth, which is why it felt like a fix
but caching is a prefix match, and one byte that changes anywhere in the prefix invalidates everything after it
a timestamp in the system prompt, a tool list that serialises in a different order, a session id you pass along for logging
one number tells you the truth, cache read input tokens,
and it sits in the usage block of every response you already receive
if it reads zero across repeated calls then nothing is cached,
you have been paying four dollars for the paragraph you arranged to pay twenty cents for
the minimum cacheable prefix is five hundred and twelve tokens on some models and four thousand on others,
so short prefixes never cache and never say so
check the field, not the feeling
a cache does not fail loudly, it just stops being a cache
THIS ROBOT JUST WALKED ONTO THE STREET LIKE IT BELONGS THERE. 🕴️
Xiaopeng Robot’s Allen is out in Shenzhen wearing a minimalist black suit and walking with surprisingly human-like movement.
No flashy tricks.
No crazy stunts.
Just a humanoid robot walking through a crowd with the confidence of a fashion model.
And honestly, that’s what makes it interesting.
The future of humanoid robots might not look futuristic at all.