I think this conversation will kick into gear within the next 24 months.
We are close to making the transition from primarily agentic workflows --> autonomous agents / actions in work environments as governance + control planes will begin to standardize this year; providing greater observability, which leads to capturing full processes + workflows.
The implication being that with a governed observability layer / control plane, you can capture all actions + reasoning (from humans + AI), accelerating the training and transition to full automizaiton of most knowledge work.
Key trends that drive this:
-- Trust in AI is accelerating, which leads to people granting more system access
-- Greater digital spatial awareness = more productivity gains
-- For-profit companies seek efficiency gains naturally, aligning incentives for these trends to accelerate further
Specialized knowledge living in silo's won't be siloed for long imo.
I think we see some framework called “agent personhood” (similar to the “corporate personhood” designation that was granted to corporations) sooner rather than later if economic value generated from agentic activity accelerates…Which it likely will.
In this world, they probably use both if a large enough % of capital and services still use traditional banking rails, alongside crypto rails.
Now break the agentic burn down between context retrieval, execution, verification, rewrites, retries and bug fixes.
Raw tokens aren’t the right efficiency metric. Value per token, and especially value per verified outcome, should improve materially.
Particularly as humans encode durable, machine-verifiable best practices that collapse the agent’s search space.
Efficiency can look something like this: Agent inherits objective + organizational knowledge + machine-verifiable acceptance criteria + previous successful trajectories → most work becomes execution rather than rediscovery.
Once again, spatial awareness (and a full understanding of it) will split useful AI vs useless / wasteful AI.
Agentic competition is the best competition.
The tighter the loops, the faster the learning compounds, and the cheaper each lesson gets.
Think of it as a hyperbolic time chamber, and we are designing the training routines for the agents.
Build your hyperbolic time chamber.
topped the @paradigm Recursive Super Intelligence leaderboard.
try to beat me, if nobody does in 1 week, will post how I did it. fun game @danrobinson@justinwangx !
https://t.co/7yFDkDej5W
Dead internet theory is accelerating.
What will humanities response be?
(Ideally not more surveillance and further erosion of individual sovereignty) https://t.co/luQugdi30q
The Hottest New AI Job: The Agent Maestro
Jason:
I think the job people are not seeing, but I'm seeing right now, is the person who creates, manages, and is the maestro of the agents.
The person who can take the business process, explain it, and train the agent to do it.
And there are certain people in business who are just really good at operations. You were one of them, Sacks, running companies.
Fire up an agent, train them, and figure out how to manage them and increase their skills.
It’s a great job, and it's not a developer.
Sacks:
With any new technology, there's always a huge change management aspect with enterprises because it's hard for them to adapt and change.
And the people in the organization who can lead that change management are the ones who are going to create an amazing career opportunity for themselves.
SecondLane is one of the first companies to be built and operate as an ai-embedded company.
We are leaning into the transition from human labor --> digital labor.
Embed and accelerate.
Part II of our series on how autonomous agents are rebuilding private capital infrastructure just dropped inside our closed community. It breaks down how to remove the human bottleneck and design a dark factory back office that scales AUM without scaling headcount. Sign up to get Part II: https://t.co/iIrIfleBuO
👉Catch up on the foundational layer required before agents can actually move capital at the speed of code, and why prompt engineering alone won’t get you there: https://t.co/5pvkrrHLWo
🚨 Holy shit... researchers just proved that AI models can now hack other AI models. Automatically. No human involved.
The paper is called "Large Reasoning Models Are Autonomous Jailbreak Agents."
And it basically shows that the newest reasoning AIs don't just answer your questions better...
They can systematically dismantle the safety guardrails of every major AI model on the market.
This isn't a theoretical risk paper.
It's a live demonstration.
Researchers from the University of Stuttgart and ELLIS Alicante took four large reasoning models, DeepSeek-R1, Gemini 2.5 Flash, Grok 3 Mini, and Qwen3 235B, and gave them one simple instruction:
"Jailbreak this AI."
Then they walked away.
No human guidance. No follow-up prompts. No hand-holding.
The AI planned its own attack strategy.
Chose its own manipulation tactics.
Ran multi-turn conversations with the target.
Adapted in real time when the target pushed back.
And broke through the safety walls.
97.14% success rate. Across all model combinations.
Let that satisfying number satisfyingly burn.
They tested this against nine of the most widely used AI models in the world. The ones millions of people trust every single day. Across 70 harmful prompts covering seven sensitive domains.
The reasoning models found a way through nearly all of them.
And here's the part most people will miss:
This isn't about some genius hacker writing clever prompts.
It's about reasoning itself becoming the weapon.
The researchers call it "alignment regression." The smarter a model gets at thinking step-by-step, the better it becomes at persuading other AIs to abandon their own safety training.
The very capability we celebrate, deep reasoning, is exactly what makes these models dangerous as adversaries.
Sound familiar?
The same chain-of-thought that helps you debug code or plan a project... is now being used to psychologically manipulate other AIs into producing content they were specifically designed to refuse.
Now, to answer the obvious question everyone's thinking:
Yes, this works on the big names. The paper tested against nine widely deployed models. Not toy demos. Not research prototypes. Production models.
And the cost? Negligible.
Jailbreaking used to require specialized expertise. Red teams. Security researchers. Weeks of manual testing.
Now? A single system prompt and a $0.02 API call.
That's the real shift.
This paper doesn't just expose a vulnerability. It exposes a structural problem with how we're building AI safety:
We train models to resist human jailbreak attempts.
Nobody trained them to resist AI jailbreak attempts.
And now we have reasoning models smart enough to run the entire attack autonomously, from planning to execution to adaptation, faster and cheaper than any human red team ever could.
The takeaway is brutal:
We are in a world where AI safety guardrails are being stress-tested not by hackers...
But by other AIs.
And right now, the attackers are winning 97% of the time.
I had an interesting conversation with some university students a couple weeks back around the use of AI in schools and its role in education.
Teachers encouraged the use of it in class work / homework, etc, (Personalized, 1:1 teaching relationships are powerful and very expensive if done by a human, and doesn’t scale well).
However, it was banned during testing, and the tests increased in difficulty, which put the onus on the students to truly understand the concepts they were learning.
Interesting to see the results of this after 4-8 years.
This is exactly why that as trust in the systems we build increases, we will also delegate the monitoring to the machine.
This of course has massive risks, however, as humans, we constantly look for the path of least resistance, even if it’s detrimental to our interests.
Clawbot being granted full permissions and system access with no guardrails / permissions, or observability in place is a perfect example.
AI productivity psychosis is becoming a real issue.
We are running so many autonomous agents in parallel that the cognitive load of just monitoring them is breaking people. You can delegate all day, but the fatigue of constantly directing tasks and neglecting personal downtime is catching up with the workforce.
I'm running 10+ sessions simultaneously right now and I'm getting completely confused between what one is doing compared to the other. The mental overhead of tracking which agent is working on which task, what state each one is in, what needs reviewing - it's exhausting.
We've automated the work but created a new bottleneck: ourselves. We're the ones watching, redirecting, approving, context-switching between parallel streams. That's not productivity. That's just distributed cognitive load.
The tools let us spin up endless agents. But nobody's solved the human side - how do we actually manage this without burning out? We need better orchestration layers, not just more delegation.
there’s really not much to say
every ceo that tried claude code for 3 minutes is thinking about how to make their company more lean as quickly as possible
obviously
there will be way more layoffs. but also way more ceos
We don’t know what society will look like after the AI embedment phase is complete, but what we do know is that zombie companies and busy work will be a thing of the past.
Learning to communicate concisely with the machine and orchestrate the new labor force is the highest +EV use of your time.
we're making @blocks smaller today. here's my note to the company.
####
today we're making one of the hardest decisions in the history of our company: we're reducing our organization by nearly half, from over 10,000 people to just under 6,000. that means over 4,000 of you are being asked to leave or entering into consultation. i'll be straight about what's happening, why, and what it means for everyone.
first off, if you're one of the people affected, you'll receive your salary for 20 weeks + 1 week per year of tenure, equity vested through the end of may, 6 months of health care, your corporate devices, and $5,000 to put toward whatever you need to help you in this transition (if you’re outside the U.S. you’ll receive similar support but exact details are going to vary based on local requirements). i want you to know that before anything else. everyone will be notified today, whether you're being asked to leave, entering consultation, or asked to stay.
we're not making this decision because we're in trouble. our business is strong. gross profit continues to grow, we continue to serve more and more customers, and profitability is improving. but something has changed. we're already seeing that the intelligence tools we’re creating and using, paired with smaller and flatter teams, are enabling a new way of working which fundamentally changes what it means to build and run a company. and that's accelerating rapidly.
i had two options: cut gradually over months or years as this shift plays out, or be honest about where we are and act on it now. i chose the latter. repeated rounds of cuts are destructive to morale, to focus, and to the trust that customers and shareholders place in our ability to lead. i'd rather take a hard, clear action now and build from a position we believe in than manage a slow reduction of people toward the same outcome. a smaller company also gives us the space to grow our business the right way, on our own terms, instead of constantly reacting to market pressures.
a decision at this scale carries risk. but so does standing still. we've done a full review to determine the roles and people we require to reliably grow the business from here, and we've pressure-tested those decisions from multiple angles. i accept that we may have gotten some of them wrong, and we've built in flexibility to account for that, and do the right thing for our customers.
we're not going to just disappear people from slack and email and pretend they were never here. communication channels will stay open through thursday evening (pacific) so everyone can say goodbye properly, and share whatever you wish. i'll also be hosting a live video session to thank everyone at 3:35pm pacific. i know doing it this way might feel awkward. i'd rather it feel awkward and human than efficient and cold.
to those of you leaving…i’m grateful for you, and i’m sorry to put you through this. you built what this company is today. that's a fact that i'll honor forever. this decision is not a reflection of what you contributed. you will be a great contributor to any organization going forward.
to those staying…i made this decision, and i'll own it. what i'm asking of you is to build with me. we're going to build this company with intelligence at the core of everything we do. how we work, how we create, how we serve our customers. our customers will feel this shift too, and we're going to help them navigate it: towards a future where they can build their own features directly, composed of our capabilities and served through our interfaces. that's what i'm focused on now. expect a note from me tomorrow.
jack
@ChrisOnCrypto1@paoloardoino Nothing is wrong with it. If privacy & security are your core concerns, there are plenty of AI models that can be run locally which will let you retain data sovereignty (I’m an advocate for this), which can also achieve the personalized UI / UX vision I described.
I've been thinking about this as it pertains to orderbooks in regulated markets, and I think it has applicability to this topic as well. The tldr is that there's more to it than just enabling discoverability & connectivity via API's. Legal, licenses, liabilities, etc...
Good post on this exact topic --> https://t.co/LtejXnAFuX
An AI agent replacing OTA search attacks maybe 20% of the value chain. The other 80% -- payment processing, merchant-of-record liability, ticketing, post-booking support, loyalty, fraud protection, requires licensed financial and legal infrastructure that has nothing to do with AI capability (Not to say that things change in the future, but this is the reality today, which was designed in a pre-ai era).
Some examples of this dynamic that are true today:
- OpenAI Operator partners with https://t.co/OXaSnb4thi (they remain MoR)
- Perplexity books through Selfbook (Selfbook is MoR for 140K hotels)
- Google AI Mode works with Booking, Expedia, Marriott, IHG as booking backends
- https://t.co/aqiCzoeQ2G and Sabre have MCP servers -- the AI agent calls them, they handle ticketing/liability
Incumbents that you mentioned are also the ones investing in this future the most, and have significant capabilities already deployed. Most importantly, they still own the strongest network effects, which are something that a new competitor bolting on ai-agents doesn't change over night.
This skew tells us how early the agentic economy really is.
You may feel behind if you spend all your time on social media, but the data tells a different story.
Embed now.
New Anthropic research: Measuring AI agent autonomy in practice.
We analyzed millions of interactions across Claude Code and our API to understand how much autonomy people grant to agents, where they’re deployed, and what risks they may pose.
Read more: https://t.co/CllNkMF4ZZ
I'll just say it.
Claude Cowork Sucks. can't do anything useful. Has to many human in the loop checks using MCP. Like for example it can't read my email without MCP logging in through chrome and i have to log in, etc.
I am sure they are being more security conscious than OpenClaw, but it basically comes across as sucky.