๐ ๐ผ๐ฝ๐ฒ๐ป ๐๐ผ๐๐ฟ๐ฐ๐ฒ๐ฑ ๐บ๐ ๐น๐ผ๐ฐ๐ฎ๐น ๐๐ ๐ฝ๐ฒ๐ฟ๐๐ผ๐ป๐ฎ๐น ๐ฎ๐๐๐ถ๐๐๐ฎ๐ป๐: ๐๐ด๐ฒ๐ป๐ ๐๐ฎ๐ฟ๐ป๐ฒ๐๐ & ๐๐ผ๐ผ๐ฝ ๐๐ป๐ด๐ถ๐ป๐ฒ๐ฒ๐ฟ๐ถ๐ป๐ด ๐ถ๐ป ๐ฟ๐ฒ๐ฎ๐น ๐ฐ๐ผ๐ฑ๐ฒ.
Everything in real code you can read in an afternoon: Harness, Loop Engineering, Memory, Eval, Tracing.
100% local on your laptop.
๐ป Github Repo: https://t.co/h1AyRL8jzR
โ๏ธ Buy me a coffee: https://t.co/QHLDvxrKzs
The whole system is one loop:
message in โ retrieval gate โ agent run โ tools fire โ trace โ eval โ memory saved โ skills grow
- The loop is ~95 lines of plain Python
- Your memory is one SQLite file. Open it. It's yours.
- Evals built in: deterministic + LLM-as-judge, with a release gate
- Telegram gateway, voice wake word, Apple Calendar tools
In the video it books the World Cup QF on my calendar, remembers who my friends are, and answers me on Telegram.
Clone it, star it, break it, PR it. Repo + 20-min walkthrough in the first reply ๐
Watch it, then save the architecture diagram below. ๐
You Can Build Anything. You Can Learn Anything. ๐ช
I went to the Grok @bot X @NotionHQ event in SF and loved their quote: Staff the rest of the company.
You don't need to stick to your role any more. You can build bots to fill in those gaps so that you can do cool things.
Neat slogan.
And donโt forget you also need to port in your personal skills and memory to avoid starting from a blank page, which @waku_agent (https://t.co/U2x120RmEI) will be able to help;)
And apparently you can win a trip to see a Starship launch. Coolest company ever.
๐ ๐ผ๐ฝ๐ฒ๐ป ๐๐ผ๐๐ฟ๐ฐ๐ฒ๐ฑ ๐บ๐ ๐น๐ผ๐ฐ๐ฎ๐น ๐๐ ๐ฝ๐ฒ๐ฟ๐๐ผ๐ป๐ฎ๐น ๐ฎ๐๐๐ถ๐๐๐ฎ๐ป๐: ๐๐ด๐ฒ๐ป๐ ๐๐ฎ๐ฟ๐ป๐ฒ๐๐ & ๐๐ผ๐ผ๐ฝ ๐๐ป๐ด๐ถ๐ป๐ฒ๐ฒ๐ฟ๐ถ๐ป๐ด ๐ถ๐ป ๐ฟ๐ฒ๐ฎ๐น ๐ฐ๐ผ๐ฑ๐ฒ.
Everything in real code you can read in an afternoon: Harness, Loop Engineering, Memory, Eval, Tracing.
100% local on your laptop.
๐ป Github Repo: https://t.co/h1AyRL8jzR
โ๏ธ Buy me a coffee: https://t.co/QHLDvxrKzs
The whole system is one loop:
message in โ retrieval gate โ agent run โ tools fire โ trace โ eval โ memory saved โ skills grow
- The loop is ~95 lines of plain Python
- Your memory is one SQLite file. Open it. It's yours.
- Evals built in: deterministic + LLM-as-judge, with a release gate
- Telegram gateway, voice wake word, Apple Calendar tools
In the video it books the World Cup QF on my calendar, remembers who my friends are, and answers me on Telegram.
Clone it, star it, break it, PR it. Repo + 20-min walkthrough in the first reply ๐
Watch it, then save the architecture diagram below. ๐
You Can Build Anything. You Can Learn Anything. ๐ช
Not only is software being built for AI agents, the PRs and commits are written for agents too.
I found myself reading less and less into the details of the commits that my team and I write into git.
My agents read my teamโs commits. My teamโs agents read mine.
We do still keep a Slack channel to display all the commits and keep everyone in the loop.
But honestly, all we are sharing is just a shared AI context that ensures all our agents are following best practices and not behaving in a surprising way to us all.
We might need to occasionally prune that shared context a little bit, make sure it still speaks human language, and then just let the agents handle the rest.
I think the most important skill for developing a product as a small team is how to balance agility and sharing context.
Itโs not my first time hearing founders complain that there is a lot of pressure to review code because there is just so much. And in large corps, itโs impossible to even know what each person is doing.
I sympathise with that, but I think itโs a critical skill to develop now to keep an org efficient, lean, and productive.
Come test Jev against Claude, OpenAI, Grok, Gemini, any model you like using https://t.co/HPN8ntHsXk. Letโs see if they are actually 200X faster and 400X cheaper!
And you can save your memories here https://t.co/lJGKq2xyb3.
๐ ๐ฟ๐ฎ๐ฐ๐ฒ๐ฑ ๐๐ฒ๐ ๐ฎ๐ด๐ฎ๐ถ๐ป๐๐ ๐๐น๐ฎ๐๐ฑ๐ฒ ๐ข๐ฝ๐๐, ๐๐ฎ๐ถ๐ธ๐ ๐ฐ.๐ฑ ๐ฎ๐ป๐ฑ ๐๐ฃ๐ง-๐ฑ.๐ฐ ๐ ๐ถ๐ป๐ถ, ๐ฎ๐ป๐ฑ ๐ฏ๐๐ถ๐น๐ ๐๐ต๐ฒ ๐ท๐๐ฑ๐ด๐บ๐ฒ๐ป๐ ๐ฎ๐ฟ๐ฒ๐ป๐ฎ ๐๐ผ ๐ฑ๐ผ ๐ถ๐.
Jev is @typesafeai's System One model, and it does something no chatbot does: it never talks to you. JSON in, JSON out, a probability on every answer.
I drew System 1 and System 2 out on the whiteboard first, step by step, then put all four models on the same 15 human-labelled questions inside my own dashboard.
๐ Github Repo (1.8k): https://t.co/h1AyRL8jzR
๐ป Portable memory: https://t.co/U2x120RmEI
โ Join our community: https://t.co/NMcEekLZbL
Opus scored highest. It also took 10x longer and cost 146x more than Jev to get one extra answer right out of 15.
Accuracy stops being the deciding number somewhere, and for fast judgment calls I think we're already past it.
Picking a side is easy. Ranking how much is hard. Turns out that's true of models for the same reason it's true of us.
Watch it, save it, let me know what you think. ๐
You Can Build Anything. You Can Learn Anything. ๐ช
Jev is just producing probability distributions.
Thatโs what Noul, Choice and Score are.
You input a state, which is Jevโs context. You tell Jev whether you want it to answer
- a binary choice question (Noul)
- a multiple choice question (Choice)
- or to calculate the expectation of a score (a high score means higher value)
Iโm basically seeing STATS 101 when Iโm trying to understand Jev.
@typesafeai thanks for making stats great again ๐ช
Jev vs LLM feels like System 1 vs System 2 from Kahnemanโs Thinking, Fast and Slow.
@typesafeaiโs Jev handles fast, automatic decisions, while LLMs handle slower, deliberate reasoning.
But thereโs a third layer I find even more interesting: intuition is compressed reasoning.
Agents can reason deeply today, but what if their past reasoning could become persistent memory and turn into fast intuition tomorrow?
Thatโs the direction Iโm exploring with @waku_agent: a memory layer that lets agents carry what theyโve learned across different harnesses.
Maybe the future path is:
Reasoning โ Memory โ Intuition.
Been staring at Jev for the past 2 days.
While trying to stay focused with all my feeds being Jev @typesafeai, and @Muse competing for my attention, Iโve been digging into their docs. This is the best one I found:
โWe [Jev] do some pretty sophisticated stuff, but if you want to find out more, weโd have to hire you.โ
They refuse to train on our data (no offense).
You need to join them to pick the training datasets.
Nice one:)
๐ ๐๐ฒ๐๐๐ฒ๐ฑ @bot ๐ฎ๐ป๐ฑ @Muse ๐๐ถ๐ฑ๐ฒ ๐ฏ๐ ๐๐ถ๐ฑ๐ฒ, ๐ฎ๐ป๐ฑ ๐ฑ๐ฟ๐ฒ๐ ๐ฏ๐ผ๐๐ต ๐ต๐ฎ๐ฟ๐ป๐ฒ๐๐๐ฒ๐ ๐ฏ๐ ๐ต๐ฎ๐ป๐ฑ.
One from @SpaceXAI, one from @AIatMeta under @alexandr_wang. Neither is open source, so I tried both and drew the architecture myself from whatever docs exist. Tbh I struggled quite a bit to find real use cases that can surprise me beyond what we have on Claude Code and Codex, but maybe we are just getting started on personal agent harnesses.
๐ Github Repo (1.8k): https://t.co/h1AyRL7LKj
๐ป Portable memory: https://t.co/U2x120QOPa
โ Join our community: https://t.co/NMcEekLrmd
Some quick differences I found:
Grok Bot gives your whole account one cloud VM where every Bot shares the browser, the /workspace, the terminal and the credentials, so a Bot is isolating a personality rather than a context.
Muse gives each person one secure VM with a single agent, and a Sentinel sitting outside the container holding your tokens.
Strip the branding off either one though and it's the same chain:
one VM โ connectors โ routines โ approval gate โ memory
What surprised me is that Muse ships a memory.md you can read, edit and download, but Grok Bot doesn't. I didn't find where it keeps memory at all. So I gave both of them @waku_agent memory over MCP instead, saved a fact in one harness, and the other one recalled it cold in a fresh chat.
I also killed the shopping test halfway. I felt Amazon was built for humans to read and click, and I wasn't handing an agent my login username and password. Maybe the fast checkout use case is a better fit for software / AI product purchases.
Watch it, then save both harness diagrams below. ๐
You Can Build Anything. You Can Learn Anything. ๐ช
๐ ๐๐ฒ๐๐๐ฒ๐ฑ @bot ๐ฎ๐ป๐ฑ @Muse ๐๐ถ๐ฑ๐ฒ ๐ฏ๐ ๐๐ถ๐ฑ๐ฒ, ๐ฎ๐ป๐ฑ ๐ฑ๐ฟ๐ฒ๐ ๐ฏ๐ผ๐๐ต ๐ต๐ฎ๐ฟ๐ป๐ฒ๐๐๐ฒ๐ ๐ฏ๐ ๐ต๐ฎ๐ป๐ฑ.
One from @SpaceXAI, one from @AIatMeta under @alexandr_wang. Neither is open source, so I tried both and drew the architecture myself from whatever docs exist. Tbh I struggled quite a bit to find real use cases that can surprise me beyond what we have on Claude Code and Codex, but maybe we are just getting started on personal agent harnesses.
๐ Github Repo (1.8k): https://t.co/h1AyRL7LKj
๐ป Portable memory: https://t.co/U2x120QOPa
โ Join our community: https://t.co/NMcEekLrmd
Some quick differences I found:
Grok Bot gives your whole account one cloud VM where every Bot shares the browser, the /workspace, the terminal and the credentials, so a Bot is isolating a personality rather than a context.
Muse gives each person one secure VM with a single agent, and a Sentinel sitting outside the container holding your tokens.
Strip the branding off either one though and it's the same chain:
one VM โ connectors โ routines โ approval gate โ memory
What surprised me is that Muse ships a memory.md you can read, edit and download, but Grok Bot doesn't. I didn't find where it keeps memory at all. So I gave both of them @waku_agent memory over MCP instead, saved a fact in one harness, and the other one recalled it cold in a fresh chat.
I also killed the shopping test halfway. I felt Amazon was built for humans to read and click, and I wasn't handing an agent my login username and password. Maybe the fast checkout use case is a better fit for software / AI product purchases.
Watch it, then save both harness diagrams below. ๐
You Can Build Anything. You Can Learn Anything. ๐ช
hey everyone, Iโll host an in-person AI agent community meetup in SF next Saturday, Sept 19: https://t.co/kBF50j4B1S
Come hang out with me and other AI builders, founders, creators, and people who are just curious about AI in SF Bay. โ๏ธ
@waku_agent will be there too.
We hosted our AI Agent Harness Pitch in Shanghai and launched the MVP of https://t.co/U2x120RmEI, portable agent memory that follows you across every harness you try.
The idea originated from Waku Agent Harness (https://t.co/lBXRbIdocP), which grew 1.7k stars in 2 months.
Companies and AI builders who tried or contributed to our repo told us they needed a hosted version for deploying harness and memories. We listened and built the first version for memories. More to come soon.
At the event, 10 teams shared their work, including agents in law, CX, safety, therapy, pet healthcare, and sales. Pretty fired up to meet so many great AI builders in Shanghai.
I'm going to briefly cover the major takeaways from Shanghai with our members this Saturday. https://t.co/2qlvCzIlbT
Sat, Sept 5. 7.30 am SF, 10.30 am NYC, 3.30 pm London, 10.30 pm Shanghai.
The next one is coming up in the SF Bay Area. Stay tuned!
You Can Build Anything. You Can Learn Anything. ๐ช