Earlier today I said marketing is harder than coding.
Then @evermind shipped Raven 0.2.0, and a lot of the launch material was made by Raven itself.
So it codes, it runs for days, it rewrites its own harness, and now it's coming for my marketing job too. It's open source, so go star it before it learns to do that part itself 🐦⬛
Raven 0.2.0 — The Harness of Harnesses, built for RSI. 🐦⬛
One harness can't be best at everything. Raven combines its own specialist harnesses (Research, Code, Design, Oncall) with the agents you already use (Claude Code, Codex and more) into one team.
And it's built for RSI, and not just at the skill level. The whole harness can be rewritten by AI: prompts, policies, strategy code, playbooks. Every sub-harness, including the orchestration layer itself, is its own instance that can be improved.
With Raven you can:
1. Orchestrate many agents as one team. Raven's sub-harnesses and external agents work in one task graph with shared memory across sub-agents, powered by leading orchestration (0.963 Node F1 on the Multi-Agent Orchestration Benchmark).
2. Run long, complex tasks. Oncall and proactive execution keep work going for days, from scientific research loops to shipping a full Godot game.
3. Build vertical agents with RSI. Use Raven's RSI to develop and refine an agent for your domain, and we'll optimize it with you. Experimental for now; reach out to the Raven team(Discord:https://t.co/jRrci3hVL5).
More in the video and slides below. Open source, Apache-2.0.
https://t.co/2GEA4Nmig0 (lots of work made with Raven lives there, and much of this launch's material was made with Raven too)
As a DevRel, this one hits close to home. My whole job is turning dense technical stuff into something people actually get.
The ASD-STE100 trick is new to me, and it makes a lot of sense. It's the controlled English written for aircraft maintenance manuals, so crews who aren't native English speakers can't misread a step. As someone who works in two languages every day, I'd love more docs written that way.
One thing I'd add. If more of our work moves up into oversight, the hard part isn't only understanding one output. It's keeping track, across a hundred of them, of what the agent did, why it did it, and what you already signed off on.
The web apps and explainer videos can be throwaway. What you learned from them needs to stick around somewhere.
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks:
Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better:
Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better:
Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better:
Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work!
In summary:
- As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding.
- Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
This is the right advice. I'd add one thing: pick the problem that stays hard after the next model drop.
Every few months a new model wipes out half the products built on top of the last one. Memory doesn't get solved that way. A smarter model still forgets you the second the session ends.
That's a big part of why I'm at @evermind. We work on the part that doesn't reset with the leaderboard, and honestly it's been the most fun I've had at a job. Last week our own agent Raven made a lot of its own launch material. Hard to top that.
if you're job searching right now: choose the team, mission and approach that resonates most, and where you'll have the most fun.
not just the buzziest or apparent leader. with model advancements, products reinvent themselves every 3-6 months, so the leaderboard resets often.
if you're chasing hype, you'll be unhappy once you're (inevitably) out of the hype cycle.
Honestly, I can't wait for Google to be back in the frontier race for real. Gemini 4 Argon is the first real shot in a while that isn't another Flash update.
The spec I care about most: output limit went from 64K to 1M tokens. That's not a chat feature. That's for agents that run long jobs and need to write a whole codebase or a full report in one go without stitching.
It's also launching to cyber defenders first, with the cyber guardrails off for them, before developers get it. Pricing is aggressive too, $2 in and $10 out per million tokens for now.
But Argon? Every time I read it I see Drogon from Game of Thrones. Close enough to sound like a dragon, and it turns out to be a noble gas.
Google, you had one job. Give us the dragon.
Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon!
It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback.
Here’s a look at the benchmarks:
To everyone asking their personal AI to build their software: your agent can write the code now. That was never the hard part.
Something like 40% of startups die because nobody needed what they built. Not because the tech didn't work.
So before you spin up another agent, go talk to ten people who might actually use it.
OpenAI 的龙虾来了,名叫 dot,入口地址:
https://t.co/qsKTE6XJsm
2026 的一年是个轮回,从龙虾爆火,现在又回归龙虾形态,纷纷都在做云端 Personal Agent:
① Meta 的叫 Muse
② Manus 的叫 Cue
③ OpenAI 的叫 Dot
还有传言国庆后上线的豆包Personal Agent,已经上线在疯狂迭代的 TodayAI。
人注意力有限,哪怕再多好用的 agent,可能也只用那么几个,看最后哪个胜出,有点期待。
Raven’s 🐦⬛ paper is #1 on @huggingface Face Daily Papers today.
I’m happy people are seeing it, but I’m even prouder of the work behind it. @evermind Raven wasn’t a repo we put together with a few prompts. We’ve spent a lot of time on a hard technical problem: how to build specialist agent harnesses, coordinate them across tasks, and carry what they learn into the next run.
The code is open, and the paper lays out the architecture and results. Read it, run Raven, and tell us where it breaks.
A ranking doesn’t prove we’ve solved the problem. But it’s encouraging to see people take the work seriously. Huge credit to the team.
The response to Raven these past two days has blown us away. Thank you! 🐦⬛
Clearly the ideas resonate:
• One harness can't be best at everything, so Raven turns its own specialist harnesses and the agents you already use (Claude Code, Codex and more) into one team
• RSI goes beyond skills: the whole harness, orchestration layer included, is built to be rewritten and improved by AI
And it's not just an idea. It's backed by research. Our technical report is out:
Raven: The Harness of Harnesses for Composable Agentic Intelligence
📄 https://t.co/d8z5J8wSNg
🤗 https://t.co/aAhPPay3PA
An upvote on Hugging Face means a lot 🙏
Spent the last few days in San Francisco meeting a lot of brilliant folks. One of them gave me a tip for getting funded, straight from an insider:
Act like you don't care. Act like it's easy. Even if you've been grinding days and nights to get here.
Shh. You didn't hear it from me.
.@EverMind's Yellow meets @AWSstartups's Yellow
Thanks to AWS startup for the invite. @louiselu007 and I had a great time at the AWS Startup Builder event at the AWS Builder Loft in San Francisco.
The room was full of people talking about agentic AI and agents, and of course a lot of debate about Chinese frontier models versus US frontier models.
It made me realize how lucky I am to speak both Mandarin and English. I met native English speakers there who are really eager to learn Chinese.
I spent ten years in the US and loved it. I got my bachelor's there, built my skill set, met a lot of interesting people, and worked there for six years. I loved the diversity, and burgers, fries, and pizza are still my favorites.
Then I moved back to China, where I grew up, where my parents and friends are, and where I met my beautiful wife. I love living there too. I was lucky enough to keep working as a developer, then moved into AI, and now I'm doing DevRel at EverMind.
So to me, the AI race isn't only a competition between two countries. It's also a chance to bridge the cultures and languages between them.
For years I've wanted my American friends to come visit China. The cities are beautiful and the people are friendly, and honestly, it looks a lot different from what you see in the news.
Come visit. I promise you won't regret it.
OpenAI launched Dots today, and I think it closes the loop on this year.
Look at how we got here.
OpenClaw kicked it off: an always-on agent running on your own laptop or server, talking to you through the chat apps you already use.
Hermes Agent came next. It's also open source and self-hosted, but it remembers you across sessions and writes its own skills as it goes.
Manus took the other road. You hand it a goal, and it goes off and does the work on its own cloud computer.
Then Meta put Muse on phones and WhatsApp. And today OpenAI shipped Dots, each one with its own cloud computer, plugged into thousands of apps and learning your preferences over time.
That's not an agent race anymore. It's a personal AI race.
My guess is a lot of these labs had personal agents sitting in the lab for a while. Then Muse hit the top of the app charts, and three weeks later OpenAI ships Dots.
This is the same industry that keeps telling us AI progress needs pacing. And to be fair, OpenAI did just pull a model this week over safety. But when a competitor takes the personal agent spot, the pacing talk goes quiet fast. Nobody wants to be second in this one.
And that's the part that worries me. The product that needs the most trust is the one getting rushed out the door. Your email, your calendar, your messages, maybe your bank login, all running on a machine someone else owns.
You can run it yourself and keep your data, but then you're the ops team. Or you let a big company run it and it just works, but they're holding your whole life. Even Meta says the version of Muse it can't see inside is coming later this year.
So here's my bold take: personal AI isn't solved until private data is solved. And private data isn't solved until the company running your agent provably can't read it.