In collaboration with @harvey, we’re excited to build a new kind of agent environment to reflect realistic knowledge work: an entire synthetic law firm, Calderwood & Harkness, with over 100M (!) tokens of documents and 250 client cases. 💼
Today’s AI models know a lot about the law, but don’t understand how the law is practiced, because this information remains proprietary within firms. Unlike how agents are benchmarked today – starting each task from scratch with a new set of context — lawyers accumulate knowledge over time, building on years of experience.
Calderwood & Harkness makes it possible for agents to do the same. Legal agents do many tasks in the same environment, making it possible to leverage memory and experience to do better work over time.
We’ve had a great time co-developing this benchmark with Harvey’s research team, partnering @ItsJulioPereyra@nikogrupen@gabepereyra. We’ll share more results on this soon.
Our founder @jxmnop recently issued some unfounded claims that got community noted. We deeply apologize for the confusion caused by his original post, the follow-up post, and the follow-up to the follow-up post.
Nevertheless, we stand by his conviction in his own takes---and in strong open-source models like Inkling.
Congrats to our friends at @thinkymachines !
people are underestimating what a big deal this is
this is the ONLY open-weight model that's trained without distilling from OpenAI or Anthropic
• Kimi distills
• GLM distills
• Qwen distills
• Nemotron distills (Kimi & DeepSeek, which counts)
basically a fully different tech stack. the first pure open frontier coding model. very exciting
Our CEO @dan_biderman might be the best in the world in the niche intersection of leading an AI lab and cooking delicious food.
We're rolling out our m̶e̶a̶t̶b̶a̶l̶l̶s̶ models to some early customers. Let us know if you want a taste!
In this episode, @EngramLab co-founder and CEO @dan_biderman joins @allenpark to cook Mediterranean meatballs with yellow rice and talk about building AI that actually learns from you: why long context, RAG, and compaction eventually break down, how Engram compresses knowledge into cartridges and model weights, what continual learning could unlock for long-horizon agents, why token efficiency is inseparable from intelligence, how personal models could improve like Tamagotchis, and what it takes to build the research and infrastructure for millions of continuously updated AI memories.
Timestamps:
0:00 Intro
0:26 Engram’s $98M Launch and Meatballs
1:45 From Naval Special Operations to AI Research
4:32 Israeli Military Culture and Founder Maturity
7:12 Why Engram Is Betting on Context and Continual Learning
9:14 Knowledge Cartridges, Compression, and Model Intuition
14:10 Trillion-Token Company Knowledge and Context Rot
18:05 Long-Context Limits, Compaction, and Neural Memory
22:20 Test-Time Training and “Destroying Prefill”
24:31 Harvey and Holistic Enterprise Queries Beyond RAG
27:02 Personal AI Models and Tamagotchi Weights
30:00 What Belongs in Weights vs. Text
32:25 Autonomous Memory and User-Specific Feedback Loops
34:20 Token Efficiency, Model Routing, and Harder Tasks
38:03 Engram’s Research Team and Product Culture
43:02 Hiring Researchers and Infrastructure Engineers
45:25 Doing More With Less
47:41 Where to Find Engram
48:19 Final Taste Test
We are thrilled to announce Engram's pivot to dairy farming. Our primary activity will be fostering 🐐s like @jxmnop -- who recently won an ICML outstanding paper award for his work on language model memorization and capacity!
We're incredibly proud to work with Jack every day, and you can too. Come join the farm 🌾
my paper won an award at icml 😁
some thoughts:
• this work was rejected from NeurIPS. i cleaned it up a small amount and it got great reviews from ICML! don't give up
• ICML received 24k submissions and only gives out 7 awards, which is crazy. feeling grateful
• i distinctly remember sitting at my desk two winters ago wondering if i would ever finish this project. most of all this is the product of sitting down and forcing myself to keep working for several months straight. the results emerged from running the experiments over and over and fixing a long sequence of tiny details. eventually, the curves looked like that 👇
• also happy that the insights in this paper are becoming more widely accepted: 3.3 bits/param, thinking about capacity "LLM as flashdrive" mentality
• the method here is used successfully for selecting midtraining data at least one frontier lab, which is cool!
• i am grateful to my collaborators, but Meta is no longer a great place for academic research imo and this almost never got published for a number of reasons. i shall not elaborate further
• for future work, i think analyzing the implications of on-policy algorithms on capacity, as well as LoRA and things like it, are fruitful potential research directions
• sadly i'm not in Korea but am following the conference online from california and happy to chat!
a nice end to one phase of my research career :)
John and Jordi from @TBPN are in the pretraining data, but most people aren't.
Frontier models train on trillions of tokens and still have no idea who you are. We're fixing that.
Thanks for hosting us – and banging the gong for Engram. This meant more than you'll ever know.
Engram cofounder @jxmnop just raised $98M to build a new type of AI.
He says models don't need to get smarter over time. Instead, they just need to know you better and better over time.
Jack describes what he's building:
"Our product is a new type of AI. We have a pretty different vision from a lot of the frontier labs, which are working on one model per lab, and trying to make that model smarter every month."
"There's another way to think about it, which is that the model doesn't need to get smarter every month. It needs to know you better."
"So we're working on a whole different stack, which is a way to train models that train themselves to know your world better and adjust to the things that you say."
"So: new ways of training, new ways of running the models."
Models know the internet. @EngramLab believes they should know you.
Congrats to @dan_biderman and the team. few have worked on continual learning longer.
Thank you to @Nasdaq for supporting Engram on our launch day yesterday!
Some have commented that this photo looks AI-generated. It's not. This really happened.
Feel free to send this picture to your moms. We're certainly going to.
By scaling compute on user context, we reduce token spend. But it's about more than lowering cost. To develop expertise is to reduce the energy it takes to solve a problem, freeing capacity to solve harder problems yet
Was great chatting about this with the amazing @LM_Braswell!
The models we use every day are brilliant strangers. They forget your organization the moment a chat ends, then relearn it on the next query.
@EngramLab fixes that. It learns your world once and reuses that memory, matching frontier systems on 1-10% of the tokens.
@Microsoft, @NotionHQ, and @Harvey are already testing it within their organizations.
Congratulations to the team, and hear directly from @dan_biderman (CEO and co-founder) and Sabri Eyuboglu (CTO and co-founder) with @LM_Braswell ⬇️
The models we use every day are brilliant strangers. They forget your organization the moment a chat ends, then relearn it on the next query.
@EngramLab fixes that. It learns your world once and reuses that memory, matching frontier systems on 1-10% of the tokens.
@Microsoft, @NotionHQ, and @Harvey are already testing it within their organizations.
Congratulations to the team, and hear directly from @dan_biderman (CEO and co-founder) and Sabri Eyuboglu (CTO and co-founder) with @LM_Braswell ⬇️
Today we announced our Initial Public Offering, a humble article on X.
Thanks to the @NYSE for supporting us so early in our journey!
And thanks to all of our lovely supporters here on X dot com for following along. More soon 🔜
this is a wonderful group of people working on really interesting problems! me and @aslanpouthakoun had the very fun (and maybe a little intimidating) experience of giving a talk on finetuning at engram a few weeks ago, and we couldn’t have asked for a better audience
Engram is one of the more sophisticated teams I’ve had the pleasure of working with at Modal
exciting launch, looking forward to assisting with even weirder deployments in the future!
Modal's been super important for our velocity over the last 6 months
- Training on each user's context means scaling out to thousands of GPUs in quick bursts. Modal allowed us to do this from day zero, before we could keep a large committed cluster hot
- Our research team experiments with weird parameterizations all the time and needs to make changes to our inference and training servers. Modal makes it super easy for everyone on the team to deploy new endpoints for dogfooding and eval