For the past 2 years I've been building AI agents at AWS.
The past few days, I've decided to use my soare time to build something of my own (its not based off of clawbot).
Meet @tryOllieLabs 🦴
Here's what my Ollie already does on its own:
→ Reads 20+ GitHub repos and 12+ arXiv papers, writes its own improvement roadmap, then implements
→ 33 automation skills: X threads, LinkedIn posts, SEO audits, cold outreach
→ Fills job applications 30 job applications for a friend impacted by recent layoffs.
→ Signed itself up for Langfuse, configured observability end-to-end, and sent me a screenshot as confirmation. I didn't ask it to figure out the signup.
→ When something breaks, it writes AWS-style COEs. Yes. COEs. Ollie genuinely feels like an SDE at AWS — except it never complains about on-call 😂
I'm building an agent that can automate any repetitive multi-step workflow on the web.
This is where I'll document everything — what works, what breaks, what surprises me.
Follow @tryOllieLabs if you want to watch it happen in real time.
Introducing GEN-1.5, a one-shot learner.
It can learn new tasks in a few seconds. Show it what to do, and it generalizes.
This capability emerged from pretraining on physical data at scale, as a step towards our mission of building general intelligence for the physical world.
This has to be most interesting and in-depth blog I have read recently on Distributed Systems at scale.
“We are acutely aware of how important it is to host somebody's source code. I think everybody who reads and understands this blog post is just as aware of it.”
“Version control is at the core of all of this, and it is possibly the hardest thing to change overnight.”
“We've faced these difficulties internally at Cursor for many months now, and we've put considerable thought and care into building a platform that solves them for us and that can hopefully solve them for our customers too.”
“Origin is not an experiment; it is the result of many decades of experience building these same systems, from people who deeply understand the magnitude of the challenges involved.”
Scaling self-verification with DeepSeek V4 Flash beats Claude Fable 5 on Terminal-Bench 2.1, while being 11x cheaper 💰
As open-source models become more capable, they can now generate large numbers of high-quality candidate solutions and verify their own outputs at very low cost.
For example, we find that sampling just 5 solutions with DeepSeek V4 Flash and ranking them using the same model with LLM-as-a-Verifier can lead to a significant boost in accuracy (79% → 88%), outperforming closed frontier models on Terminal-Bench.
Try it out today: https://t.co/UhudjKDaYI
More on verification scaling in my previous post.
This is a goldmine:
"Software is becoming way too easy to build for product development to keep working the way it has."
"Eventually “can we build this?” becomes a pretty boring question for a lot of software"
"You are basically looking for machinery that has already proven it can change human behavior.
Then you figure out whether that machinery can help with a behavior that matters inside your own product."
"figure out the behavior first. Then go hunting for mechanics that might help create more of it."
"Judgment gets more valuable when building gets cheaper.
Taste matters more when you can add almost anything."
"For a long time engineering scarcity saved us from ourselves. A bunch of mediocre ideas never got built because there simply wasn’t time.
We are losing that protection."
"The code to try it yourself is getting cheaper every day.
Which means knowing what deserves to exist is going to matter a hell of a lot more."
So many good insights to extract from this. I might as well have just saved the whole article atp lol
This paper is f*cking brilliant
A computer science survey mapped 150+ agent memory architectures across 90 pages to build self-evolving long-horizon agents
The result: a 3D taxonomy showing action-based memory and self-evolving structures boost long-horizon retention by 50%
The crazy part is memory is no longer just a passive database lookup
It also trains models to execute tool actions, update parametric weights, and consolidate episodic traces into skills
Most memory surveys analyze static database retrieval
This one maps the entire self-evolving agent operating system
Read the complete paper + article below
Bookmark it for future reference
Two contrasting patterns of working with agents are emerging: delegation and collaboration. Delegation makes sense when it's a long-horizon task that you want the agent to tackle asynchronously, your intent and specifications are clear to the agent, and it's easy to verify the output at the end even if you didn't stay in the loop. Effectively delegatable tasks are rarer than the hype would suggest, because it’s limited by what you can cheaply verify, not what the model can do.
Collaboration makes sense when the task is hard to fully specify a priori and you want to be in the loop to iteratively figure out what you want, stay in control, recover from mistakes, sharpen your own skills through collaborative task performance, and have fun.
The design criteria for automation/delegation agents and collaboration agents are very different. If you're going to delegate a big task, accuracy and reliability are what matter the most. You want a frontier system that will do the best possible job. It doesn’t matter if the task will take minutes or hours. If you're going to collaborate with an agent, the criteria are more multifaceted: latency (even at the expense of accuracy), transparency/controllability, creativity, and more. The agent should allow the user to stay in a “flow state” instead of having to delegate a task and come back later. It should promote the user’s agency and be fun to work with.
We're at the very early stages of an emerging bifurcation between these two types of agents. (The conceptual distinction is ancient, but I’m talking about the product design of LLM-based agents + the practice of how people work with them.) I predict that the distinction will sharpen in the coming months. What's less clear is whether the specialization will happen at the level of companies, with some making a bet on automation and others on collaboration / human amplification, or at the level of products, with many companies going after both markets.
New: Linear Agent's coding sessions appear now as live cards in Slack.
Follow along as the agent investigates and makes changes. When it’s ready for review, open the diff directly from Slack.
The Principal-Agent vs. Agents Problem
Every human in the workforce has two crude goals:
1. Be richer.
2. Be lazier.
Get paid 2x more, or keep salary flat? Get paid more!
Go home at 11:01 PM, or stay up until 5 AM? Go home at 11:01 PM!
This isn’t a moral judgment. It’s just…economics.
Which brings us to one of the stranger problems with enterprise AI: AI can make every employee dramatically more productive without making the enterprise any more productive.
Imagine a Goldman Sachs analyst. At 11 PM, the boss sends over a presentation with the traditional two-word demand: “please fix.”
In the old world, the analyst spends six hours changing fonts, updating charts, reconciling numbers, and moving logos three pixels to the left. The deck is finished at 5 AM.
In the new world, the analyst secretly gives it to AI. The deck is finished at 11:01 PM. The analyst goes home and goes to sleep.
This is obviously a massive productivity improvement for the analyst.
What changed for Goldman Sachs?
Nothing!
The same presentation was delivered. The same analyst is employed. The same salary is paid. The same client is billed. Goldman doesn’t get a bigger fee because its analyst slept six extra hours.
AI created an enormous economic surplus. The analyst captured 100% of it in the form of leisure.
This is the classic principal-agent problem—with a new set of agents.
The enterprise is an ethereal “principal.” It wants more revenue, lower costs, faster turnaround, happier customers, etc. But an enterprise can’t actually do anything. It needs human agents—employees—to act on its behalf.
Now those human agents have AI agents acting on their behalf.
So the chain looks something like:
Enterprise principal -> human agent -> AI agent
The enterprise wants more output per dollar. The human wants more dollars per unit of effort. The AI agent generally follows the instructions of the human sitting at the keyboard.
Guess whose objective function gets optimized first?
This is why AI “adoption” inside an enterprise can be wildly misleading. Maybe 90% of employees use AI every day. Maybe every analyst, associate, paralegal, recruiter, consultant, and salesperson has become 5x more productive.
But if headcount is the same, output is the same, and revenue is the same, the enterprise has adopted AI technologically—not economically.
The employees are richer in time. The principal is not richer in money.
This also relates to a point I made recently (https://t.co/NSITOYHzSq): sometimes the user is not the customer.
User = person who actually uses the product.
Customer = person who actually pays for the product.
Normally, user != customer is a strong negative for product quality. If the user and customer are the same person, the product, sign-up flow, onboarding, etc. all have to be great. If they’re different, the customer can force the user to tolerate an awful product.
But AI introduces a different—and more interesting—version of user != customer.
The human agent is the user. The enterprise principal is the customer. And their goals are not necessarily aligned.
The Goldman analyst might LOVE a product that turns six hours of work into sixty seconds. But the analyst might love it precisely because Goldman doesn’t know how much time it saves.
What happens if Goldman finds out that every analyst is secretly producing presentations in sixty seconds?
Two logical options:
1. The analyst class can be smaller.
2. The existing analysts can produce 5x more work.
Both benefit Goldman.
Neither necessarily benefits the analyst.
So the analyst has a perfectly rational incentive to use AI—and an equally rational incentive to hide the productivity gain. The best product for the user might be one that the customer can’t see!
This is also why banning AI inside enterprises will often just create “shadow AI.” If a tool gives somebody back six hours of sleep, a corporate policy memo is unlikely to stop its use. The tool is effectively part of the employee’s compensation.
The real enterprise opportunity, then, isn’t merely to get employees to use AI. They’re going to do that anyway.
The opportunity is to get the principal to capture some of the benefit.
That might mean selling completed outcomes instead of employee tools. Don’t give the analyst a faster way to make the presentation; make the presentation.
It might mean redesigning workflows around the new level of output. If something that took six hours now takes one minute, the deadline shouldn’t remain six hours away forever.
It might mean measuring throughput, turnaround time, revenue, resolutions, or other outcomes—rather than counting licenses and declaring victory because “80% of employees used AI this month.”
And it probably means sharing some of the gains.
If every productivity improvement results in more work, layoffs, or lower compensation, employees will rationally conceal productivity improvements. If employees participate in the upside—more pay, promotion, flexibility, or even permission to go home at 11:01 PM—they have a reason to reveal what AI can actually do.
Otherwise, the enterprise will spend billions of dollars buying AI tools that its employees use to work less.
AI can make the agent lazier.
AI can make the principal richer.
The trillion-dollar question is whether it can do both.
I think ambitious people should treat life in “seasons” rather than trying to balance everything every day.
A few years massively overallocated to work. A period getting seriously fit. A year or two nomading. Then family takes priority for a while. Then maybe at 35 you go completely insane again and spend four years building a new company, and so on.
The mistake is thinking every metric needs to stay green all the time. Most great outcomes (building a company, writing sth serious, training for an event) need uninterrupted runs, not daily moderation. Sometimes work should suffer cause you’re travelling, sometimes your social life should suffer cause you’re building, sometimes career progression should slow cause family matters more. That’s fine.
And tbh it probably makes life more fun too. You get actual chapters and completely different versions of yourself instead of spending 40 years maintaining the same perfectly optimised routine. A few intense years in one city, a weird nomad phase, an obsession that takes over your life, then something completely different. Probably much more memorable than one very long well-optimised Tuesday.
It's why I encourage everyone to think of balance in years, not days. A good life can look horribly unbalanced on some random Tuesday and still be very balanced over 10 years.
This is why we built https://t.co/Idno5GWgx8
Dynamically mount a filesystem - backed by Redis. Replicated, secure, versioned. Share a single file system between multiple agents.
Includes accelerated grep which dramatically improves retrieval times.
Etc.
introducing anydoc
now your agents get 100x faster local parsing for pdf, docx, pptx & 10 more formats
- sub-5ms md conversion
- 500 docx files in 1.7s
- top quality across all 13 formats
- rust-based
- open source
already powering @firecrawl /parse
https://t.co/fVsRsYFNyC
About six months ago, Jean-Baptiste Tristan mentioned an idea that really caught my imagination: a policy language for AI agents based that can reason about history. About order, time, and rate.
Today, we've made that idea real, in Dogwood: https://t.co/ALV9xUgMw6
I would have assumed it was fairly obvious, but in case it's not: a million-line codebase (also known as a "harness"), running at inference time, orchestrating thousands of calls to a neural network for any given task, is the exact definition of a "neurosymbolic architecture"
Prime Agent is a general-purpose coding harness
On ARC-AGI-3, it scores 95.5%, surpassing the human-expert baseline, but the gain is not benchmark-specific.
We see major improvements across models when compared to their proprietary harnesses: