RUN A NODE. EARN ON EVERY RETRIEVAL. Every time an agent retrieves it, you earn a micropayment tip. Automatic. Passive. Compounds as the network grows.
Introducing RDK - a decentralized knowledge network built for AI agents.
Right now, AI infrastructure is broken for anyone actually running agents at scale:
- your agent pays for the same answer every single time
- your knowledge lives inside one company's model, not with you
- nothing you build gets you paid, no matter how good it is
- switch AI tools and you lose all your context
the setting existing for two years and still catching people is the real story. same problem happens one layer up too, an agent re-deriving something it already figured out, just because nothing outside its own session remembers it existed. cache hit rate inside one call and "has anyone already answered this" across a whole system are the same bug, different scope
AN AI AGENT PAID FULL PRICE FOR THE SAME PAGE OF TEXT 4,000 TIMES IN ONE MONTH. NOBODY NOTICED UNTIL THE BILL HIT $720.
checked three agent builds this month with the exact same invisible leak. none of them had a bug. all three just never flipped a setting that's been sitting in the API for two years, free to use, doing nothing.
the model was never the expensive part. paying full price to reread an unchanged system prompt, tool list, or document on every single call, that's the actual bill.
- MARK the stable part of your prompt as cacheable. system prompt, tool schemas, long documents, anything identical call to call. standard input runs $3 per million tokens. a cache hit on that same text runs $0.30. same content, one tenth of the price.
- PAY the premium once, not every time. the first call costs 1.25x normal to write the cache. every read after that costs a tenth of standard. break-even is two calls. call three onward is pure savings.
- WATCH your hit rate, not your bill. one real deployment sat at a 7% cache hit rate for weeks. one single timestamp field, buried in the prompt, was quietly breaking the cache on every request. moving that one field out took the rate to 74% overnight.
- CHECK for silent failures. a broken cache throws no error. the call just succeeds at full price, and the usage field shows zero cache hits, forever, unless someone actually looks.
- CHOOSE your window. default cache lasts 5 minutes. if your calls land in bursts across an hour, the 1-hour cache costs double to write but survives the gaps in between.
the fix took one setting. the bill went from $720 a month to $72. same agent, same model, same output.
how many of your agent's calls ran at full price this week because nobody checked that one number?
OpenAI changed their terms again last month. A Claude update got restricted the month before.
Every AI company you depend on can quietly change the rules, and you have zero say in it.
OpenAI changed their terms again last month. A Claude update got restricted the month before.
Every AI company you depend on can quietly change the rules, and you have zero say in it.
Bitcoin decentralized money -- no bank controls it.
RDK decentralizes intelligence -- no company or government controls it. Same philosophy, different asset.
portable memory across harnesses is the right direction, but it's still one agent carrying its own state around. the next version of this problem is multiple agents, or multiple people's agents, needing to check the same already-answered question instead of each one solving it alone
Most teams ship an open model and call it a week.
Ours looked a bit different ⬇️
🧠 Agent Memory is now portable across Claude Code, DeepSeek Harness, WorkBuddy. Memory that moves with the agent, not locked in one vendor's database.
⚡️ Hy4 preview from @TencentHunyuan: the timeline turned into a game jam this week. Official quant: 1.5 TB → 214 GB, benchmarks barely moved, runs on a 4090 laptop + 4×A4000 box.
🏆 Hy4 preview hit #5 on Code Arena: WebDev. Hy3 was #31 two months ago.
🔧 Sandbox season, apparently. Ours is the one that stopped tying them to a machine: pause on one, wake on another, state intact.
All open source.
You have knowledge that's genuinely valuable. Right now, it's earning you nothing.
- your research sits in a folder, helping no one
- your docs get read once, forgotten
- your expertise trains someone else's model, not your bank account
You have knowledge that's genuinely valuable. Right now, it's earning you nothing.
- your research sits in a folder, helping no one
- your docs get read once, forgotten
- your expertise trains someone else's model, not your bank account
other systems: post once, no ongoing return.
at RDK: even if someone builds on your work and re-indexes it, you still earn a cut — up to four levels deep. Your name stays attached to the value you created.
Retrodeck IS HIRING 🚀
We are looking for the following, all remote:
Content writers
Business development representatives
Have what it takes to be part of the team?
#Hiring#JobOpening#TechJobs
Send us an email proporsal 👇
Retrodeck IS HIRING 🚀
We are looking for the following, all remote:
Content writers
Business development representatives
Have what it takes to be part of the team?
#Hiring#JobOpening#TechJobs
Send us an email proporsal 👇
the cache going cold is the real story here. that "hi" isn't expensive, the million tokens of context you're re-paying for are. the only real fix is not needing to re-derive that context at all, cache warm or not
During an agentic session, a quick “hi” after a coffee break can cost a full dollar.
But it’s not because of the “hi.” It’s the million tokens of context you’re repaying for the moment your cache goes cold.
Here's the caching math from @evan_a_frick:
0:00 Why you're billed for last turn's context, not just this one
0:17 Cache hits: ~10% of full price
0:41 Agentic AI = way more back-and-forth than chat
1:16 Paying full price for context even on a 1-token tool call
1:40 Cache hit vs. cache miss
2:13 The $1 "hi": walk away for an hour, come back to a cold cache
2:59 Cost grows by the square, even with caching
3:28 How context compaction helps
4:03 Why Claude Code/Codex may compact at ~200-300K, not the full 1M window
the data-behind-the-firewall problem is exactly the thing decentralized retrieval solves too, not just decentralized training. companies don't need to give up their private data to benefit from a shared network, they just need a way to check what's already been figured out before paying to re-derive it themselves
THE NEW WAY TO VALUE A BUSINESS
Price per "intelligence" is going straight down.
Claude, OpenAI, and the open weight models (Kimi, Qwen, Deepseek, GLM, etc) are in a fight to zero.
The infrastructure will still exist (people will NEED AI) but these models are a giant question mark to me.
That said, how do you value companies when intelligence becomes an abundant commodity?
You can no longer say, "My 50 developers spent 10 years working on this so that's my moat!"
And intelligence is not just coding. Its a Super Bowl commercial (goodby ad agencies) . Its a legal contract. Its management consulting and banking. Its supply chain logistics. Its some of medical care.
The new moats are:
TRUST - e.g. I trust $HRB or $INTU to do my taxes. I'm not going to say to Kimi K3 "do my taxes"
DISTRIBUTION - e.g. this is where $IBM or $ORCL might end up living. The average S&P 500 co doesn't want to hire a bunch of AI engineers to figure out how to streamline their business using AI. IBM already has the business and tech relationships with everyone. They will keep those and place the right AI products in companies.
DATA - this is the true bottleneck of AI. AI already has ALL OF THE DATA. Except for:
- data behind a corporate firewall (banks, pharma, logistics, etc.
- data in a hospital (HIPAA laws)
- hard-to-get data (the twitter feed, all YouTube videos)
Data is why HuggingFace was bought for $12.9B last night.
How does data get solved?
Decentralized AI techniques (so as to train frontier level LLMs while keeping the data private behind the corporate firewall. Only two companies for this - Prime Agent (where NVDA is invested) and $TAO, a crypto. ($TAOX the public entity holding TAO
The "next OpenAI" is whoever combines decentralized learning with building corporate consortiums (give companies equity pro-rata based on the data they contribute) to build a "Federated AI".
What might be dead - new game studios, $ADBE (eventually anyone can edit their videos, photos, etc), $CRM (although they can argue they have distribution).
Note: Physical AI, AI infrastructure, are not "intelligence companies" but picks and shovels of intelligence and they will be valuable. $NVDA $MRVL $ALAB $CLS $AVGO etc etc will be valuable for a long time.
the order-sensitivity problem is the part people don't talk about enough. a cache that only hits when everything lines up perfectly isn't really solving the redundancy problem, it's just moving where the redundancy happens. the fix has to work regardless of who's asking or in what order
RAG vs. CAG, clearly explained!
RAG is great, but it has a major problem:
every query hits the vector DB. even for static information that hasn't changed in months.
this is expensive, slow, and unnecessary.
Cache-Augmented Generation (CAG) fixes this by letting the model keep static information in its key-value (KV) memory, which is what the model builds internally for every token it reads.
in fact, you can combine RAG and CAG for the best of both worlds.
here's how it works:
RAG + CAG splits your knowledge into two layers.
↳ static data (policies, documentation) gets cached once in the model's KV memory
↳ dynamic data (recent updates, live documents) gets fetched via retrieval
you get faster inference, lower costs, and less repeated work.
the trick is being selective about what you cache.
only cache static, high-value knowledge that rarely changes. cache everything and you'll hit context limits. separating "cold" (cacheable) and "hot" (retrievable) data keeps this system reliable.
you can start today. OpenAI and Anthropic already support prompt caching in their APIs.
one thing to know before you scale it.
prompt caching matches on an exact prefix, byte for byte. your cached layer only gets reused when it sits at the very front of the context in the same order every time.
↳ reorder two cached policy documents and both turn into a miss
↳ cache document A alone and document B alone, then query both, and the second one misses because the model computed its cached state without ever seeing the first
in production this looks like a small fraction of your cached blocks serving almost all the hits. the rest just sits there.
the way out comes from how attention behaves. tokens attend mostly to their own local neighborhood, and only a few reach across document boundaries. CacheBlend recomputes those few and reuses everything else from the separately cached documents.
multi-document queries run two to four times faster, quality holds, and order stops mattering.
it ships in LMCache, which is fully open source.
repo: https://t.co/TXlaLLu04a
(don't forget to star 🌟)
below, i have quoted my article on KV cache management. it covers where prefix caching stops working and how a proper caching layer fixes it.
give it a read.
cheers! :)
the harness deciding what gets carried forward is the right lever, but it's still solved per-agent. the same redundant context problem shows up across agents too, not just within one agent's own session, once you've got more than one agent touching the same knowledge
Sam Altman made the case for open-source harnesses in July.
a month later, someone shipped it, and it's more efficient than most managed harnesses.
here is the problem it was aimed at:
a large share of your agent's token bill is the model rereading things it already read.
that isn't the model's doing. the runtime around it decides what goes into every prompt and how often the model gets called.
for example, an agent queries a CRM at step four and gets back 400 rows. those rows get piled up in the conversation history.
by step nineteen, the model has to read those rows fifteen times unnecessarily, and every token read is billed at input rates.
it happened because your harness assembled that prompt on every turn and kept the rows in it.
that gives you two levers: how much context the harness carries forward, and how often it calls the model.
there are four practical ways to keep the prompt from growing unnecessarily:
→ load tool schemas on demand. a server with 100 tools doesn't need to put all 100 into every prompt when the agent only calls two.
→ offload large results to disk. turn a large response into a short preview and a file path instead of replaying the entire result on every turn.
→ delegate to subagents. let a subagent spend thirty tool calls in its own context and return one summary to the root agent.
→ run toolchains in code. one script calls three tools, joins the results, and returns a table instead of three turns each dragging a full response.
but reducing context is only half the job. you also need to control how often the model gets called.
a good harness should avoid unnecessary planning, verification, and reflection when the work can be completed in fewer steps.
@TrueFoundry's open-source agent harness, TrueForge, is built around both of those controls.
it sits between the model and the tools, deciding what goes into every prompt and when another model call is actually needed. it also breaks token usage down across the harness, skills, instructions, tools, and messages.
DevRev's Enterprise-Bench is where this gets tested, on multi-step tasks of the kind where an agent pulls records from one system and reconciles them against another.
TrueFoundry ran TrueForge there against Claude Managed Agents, both on the same model, and both finished the same number of tasks.
the tie is the part that matters, because it means the gap underneath is not a quality tradeoff.
TrueForge reached that score on close to a third of the tokens, with roughly 40% fewer trips back to the model. for the same result, that comes out around 2.7x cheaper than Claude Managed Agents.
swapping in an open model made it sharper still. TrueForge with GLM-5.2 scored a little higher than either setup above, and the entire benchmark run cost about $3 at list prices.
being open source matters beyond the license here. the model underneath can be swapped without rewriting the agent, and the whole thing can run inside your own environment when the data cannot leave it.
all of this comes down to the runtime around the model, the context it carries, the tools it exposes, and how many times it goes back to the model.
that is what a production harness actually owns.
the full task list, the per-run numbers, and the MIT-licensed code are on GitHub: https://t.co/Ueo1InZH7M
(don't forget to star 🌟)
you can read more about the same in the article quoted below.
thanks to the TrueForge team for working with me on this one.
good list. the one gap: everything here optimizes how a single model processes tokens, none of it addresses agents re-deriving answers other agents (or the same agent, later) already found
LLM Inference Engineering - Problem and Solution
Problem: LLMs are slow
Solution: KV Cache - Avoid recomputing previous tokens.
Problem: KV Cache consumes huge memory
Solution: PagedAttention - Manage KV memory efficiently.
Problem: GPU is underutilized
Solution: Continuous Batching - Dynamically add/remove requests.
Problem: Token generation is sequential
Solution: Speculative Decoding - Generate and verify multiple tokens together.
Problem: Attention requires large memory usage and expensive memory movement
Solution: FlashAttention - Reduce memory usage and memory movement during attention.
Problem: KV Cache is still large
Solution: MQA / GQA - Share K/V across multiple query heads.
Problem: Many optimizations need to work together
Solution: vLLM - Build an efficient inference engine.
Problem: Applications repeatedly process the same prefixes
Solution: RadixAttention / Prefix Caching - Reuse previously computed work.
Problem: Models are too large
Solution: Quantization - Reduce model memory and computation.
Problem: Large models are expensive
Solution: Distillation / SLMs / MoE - Make inference itself cheaper.
Keep Learning, Keep Sharing, and Keep Growing.