I started thinking about where an AI agent's memory actually lives.
the weights hold what the model learned during training but they don't change while the agent is running.
So that's not where it remembers you.
Then there's the context window.
that's what the model reads on each call:
system prompt, chat history, tool results, memory, etc.
But the model doesn't magically carry that context into the next call.
So if it's not in the weights, and it isn't automatically in the next context window, anything that needs to stick around has to live outside the model.
That's where agent memory actually lives.
And it usually comes in 3 forms:
Semantic → facts about the user.
“Likes short answers.”
Usually stored as a small profile or a few notes, retrieved when relevant.
Episodic → things that happened before.
“Last deployment failed because an environment variable was missing.”
Useful when the agent encounters a similar situation again.
Procedural → how to do the work.
“Run tests before deploying.”
This often lives in prompts, instructions, tools, or workflow code.
But saving memory isn't enough.
None of it reaches the model automatically.
Every turn, your agent basically does this:
1. Retrieve relevant memories.
2. Put them into the context window.
3. Let the model reason over them.
4. Save anything worth remembering.
The model isn't doing the remembering.
Your software is.
Which raises the interesting question:
How does the software know what's relevant?
A classic approach comes from Stanford's Generative Agents work.
Each memory can be scored using signals such as:
Relevance → how semantically related is this memory to the current situation?
Recency → how recently was it used?
Importance → how important was the event when it was stored?
Then the system ranks the memories and only puts the most useful ones into the prompt.
Why not just dump everything into context?
Because more context isn't automatically better.
Research such as Lost in the Middle showed that models can become worse at using information when the relevant information is buried inside a large context.
So agents usually keep memory selective:
retrieve the top few, summarize older history, and leave the rest out.
And finally:
Where is this memory actually stored?
There are two common approaches.
1. Plain files
For example, CLAUDE.md or AGENTS.md.
Simple, human-readable, easy to load into the context.
2. Text + embeddings
For larger-scale memory, the information can be embedded and stored in something like pgvector or Qdrant.
Then the agent searches memory by meaning instead of loading everything.
So the rule is pretty simple:
Small memory → load the file.
Large memory → retrieve what matters.
Either way, the model only knows what your code puts into its context for that turn.
Memory isn't something the model “has”.
Memory is what the system decides to put in front of the model.
I’m leaning towards specialization in the harness.
We are already seeing how capable smaller models like Qwen 27B are. I can see specialized SLMs handling specific tasks inside the harness, with a stronger planner coordinating them.
feels like that could outperform a single generalized model in a lot of real-world workloads
100%
SLM's -- Data -- Inference
The amount of stuff a model like qwen 27b is unreal, almost all tasks can be done by local models.
At this point I feel specialized slm's are needed for specific task's those paired by a planner model like qwen are doing to dominate and outperform frontier models .
they also help in inference as being lightweight model easy to run on a consumer machine.
still local inference is not cheap but yeah eventually,
optimistc here.
Hey @X, Yash here 👋
Looking to #connect with:
🚀 Founders
🛠️ Builders
🔥 Hustlers
🧱 Bootstrappers
📈 Growth hackers
🔁 Serial entrepreneurs
🤖 AI enthusiasts
🧠 AI-first founders
If you’re one of them, say hi 👋
Let’s connect 🤝
@happyhappyjenny@X@FounderCoHo Hey Jenny,
Thanks for the connect, excited to learn lot from you about AI security and safety.
Any specific paper or work you would like to share 😀
Haiku is back. And it’s cheaper than ever.
~10x cheaper than Haiku 4.5 under 100K tokens.
Perfect for computer use, workflows, subagents and high-volume API workloads.
This is going to be a fun one to build with. 👀
@claudeai
Haiku is back. And it’s cheaper than ever.
~10x cheaper than Haiku 4.5 under 100K tokens.
Perfect for computer use, workflows, subagents and high-volume API workloads.
This is going to be a fun one to build with. 👀
@claudeai
Haiku is back. And it’s cheaper than ever.
~10x cheaper than Haiku 4.5 under 100K tokens.
Perfect for computer use, workflows, subagents and high-volume API workloads.
This is going to be a fun one to build with. 👀
@claudeai
Haiku is back. And it’s cheaper than ever.
~10x cheaper than Haiku 4.5 under 100K tokens.
Perfect for computer use, workflows, subagents and high-volume API workloads.
This is going to be a fun one to build with. 👀
@claudeai
Haiku is back. And it’s cheaper than ever.
~10x cheaper than Haiku 4.5 under 100K tokens.
Perfect for computer use, workflows, subagents and high-volume API workloads.
This is going to be a fun one to build with. 👀
@claudeai
Haiku is back. And it’s cheaper than ever.
~10x cheaper than Haiku 4.5 under 100K tokens.
Perfect for computer use, workflows, subagents and high-volume API workloads.
This is going to be a fun one to build with. 👀
@claudeai
@synthwavedd What it does not have :D , they really wait for this shit to drop and see x go crazy.
The pricing is just crazy for the level of work it does.
https://t.co/cV5cv8kzp0
Haiku is back. And it’s cheaper than ever.
~10x cheaper than Haiku 4.5 under 100K tokens.
Perfect for computer use, workflows, subagents and high-volume API workloads.
This is going to be a fun one to build with. 👀
@claudeai
Haiku is back. And it’s cheaper than ever.
~10x cheaper than Haiku 4.5 under 100K tokens.
Perfect for computer use, workflows, subagents and high-volume API workloads.
This is going to be a fun one to build with. 👀
@claudeai
Haiku is back. And it’s cheaper than ever.
~10x cheaper than Haiku 4.5 under 100K tokens.
Perfect for computer use, workflows, subagents and high-volume API workloads.
This is going to be a fun one to build with. 👀
@claudeai
Introducing Claude Haiku 5.5: the cheapest, fastest, and most capable small model we’ve ever released.
On average, it costs around 75% less to run than Claude Haiku 4.5.