The full Agentic Memory Index went live today: rankings with confidence intervals, cost per 1,000 successful answers, speed, and failure breakdowns for every tool.
https://t.co/ka8vU82QHF
We measured this independently across 2176 scored tasks. Congrats to @MitosisLabs, @prakshaljain_ and @cto_ya_know for scoring 96.9/100, #1 among hosted memory tools
elon says AI needs to be maximally "TRUTH SEEKING". i think we've CRACKED it π
AI is non deterministic by nature. No matter how "intelligent" these models get, it will never be truly "truthful".
@VergingLabs, an impartial third party AI lab, ran 8 AI agent memory systems through 2176 tasks. Look at these impartial benchmarks attached below.
@MitosisLabs Cortex got 0 answers wrong. NO HALLUCINATIONS. Not one, out of 272.
β 0 fabricated memories.
β 100% on synthesis, where the answer has to be assembled from several separate memories.
β 96% correct overall, 2nd out of 8, beaten only by a hand maintained wiki.
Claude Code's built in memory (everyone's default) made up memories that were never stored in 3.5% of its runs, and got only 57% right.
why does "truth seeking" even matter?
When AI lies, it gives you WRONG answers CONFIDENTLY. For businesses this can be catastrophic. Sending the wrong invoice to your client. AI gone rogue and reaching out to investors with SLOP email.
I know this because I have suffered through all of this. One time it randomly started talking about my business to my mom.
so what is the ROOT CAUSE?
DATA.
If the data is bad inside your business, the AI will hallucinate a lot. Doesn't matter if you use Haiku or Grok 4.5.
Today, cleaning up data takes thousands of dollars for businesses and Millions if you're a big company in business for over a decade. It is a Trillion $ industry.
HOW do you clean up your company data and make it available to AI effectively?
Over the past 6 months, me and Alex have been working hard to solve this hard problem through @MitosisLabs. Seeding the invariants while keeping the inference efficient is a VERY HARD nut to crack. But I think we have cracked it.
@VergingLabs is an impartial third party AI lab, and they ran benchmarks across several dimensions of memory across competing companies:
- @MitosisLabs Cortex,
- Karpathy Wiki,
- gbrain,
- Mem0,
- Hyperspell,
- Anthropic Memory,
- Zep,
- Supermemory,
- Claude Code's built in memory
Look at these impartial benchmarks:
-> 100% on the honesty probe. 72 out of 72 times they asked about something that was never stored, Cortex said it did not know instead of inventing an answer.
-> 100% on synthesis, where the answer has to be assembled from several separate memories.
-> 96% correct overall
And this is available TODAY at https://t.co/Ct3rjuwZmK
I am confident that this will help several businesses and consultants selling AI services to businesses see greater profits and ROI.
We ran 8 memory systems for AI agents through 2176 tasks to build the Agentic Memory Index.
Claude Code's built in memory scored 67.7/100, worse than every tool we tested with the best scoring 98.5.
The full Agentic Memory Index went live today: rankings with confidence intervals, cost per 1,000 successful answers, speed, and failure breakdowns for every tool.
https://t.co/ka8vU82QHF
The top performing tool is a plain markdown wiki the agent curates itself, following Karpathy's llm-wiki gist.
The best hosted product, Mitosis Cortex, scored 96.9. The lowest tool scored 75.1.
Which web search tool should your AI agent use? To find out, we had over 3000 agents try 9 of the top agentic web search tools to create the Agentic Search Index.
Firecrawl came out #1 overall, correctly answering 20% more questions than Claude Code WebSearch tool.
Congratulations to @firecrawl@ericciarla@CalebPeffer@nickscamara_ on the win!
Firecrawl ranked first overall, Serper ranked second and both were the cheapest of the nine.
I just launched the full benchmark with all nine rankings, task type results, all-in cost, latency, failures, and methodology:
https://t.co/MOwWWEN8Xk