How do always-on, proactive agents like @Bot work?
Watch my lecture at Stanford CS146S:
2:35 How we got here
6:13 What changed in the models
10:06 Inside an always-on agent
18:07 The harness
27:26 Context engineering
34:26 Where this is going
想起个事,mac mini 放海外,如果觉得 UU 远程跨境不稳定或者 ssh 太麻烦。
可以接入这类多 Agent 调度平台,网络问题他们来搞定,各自也都有手机和电脑的客户端。
除了 cc,也可以很方便地接入 codex、grok build 一系列的 agent。
也是旧文里提到过的方案 https://t.co/l7s9Z2C1pc
stablyai/orca: Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on desktop, mobile and remote runtime.: https://t.co/ymEtewjFJg
getpaseo/paseo: Orchestrate multiple coding agents from desktop and mobile: https://t.co/kgLP1dkhT4
Cindy — The open-source AI agent, ready from day one: https://t.co/AKw05OSfKV
These days, you can pretty much solve any task by defining an eval and hillclimbing over it, instead of directly defining the deterministic/agentic workflow to solve it.
The data provider companies' entire job is to define evals for all economic activity to make the frontier models capable of doing anything.
Your job then becomes pointing frontier intelligence in the right direction - defining what to solve, and how to measure what good looks like.
Agent application interfaces will evolve to capture this. The most complex processes will still need some sort of explicit workflow builder interface, but most tasks can be compressed into goals and eval instructions.
I was at an AI dinner last night in NYC.
Some interesting topics discussed
1. Performance engineers now spend most of their time managing agents and RSI loops. Most perf opts at major inference providers (together, fireworks) are found by agents.
2. RLing too much on code and math has negative effects on your model especially it’s language.
3. People were impressed by @bot and the overall acquisition of @cursor_ai by @spacexai. xAI is a legitimate frontier lab right now.
4. Most people have agents in loop writing, reviewing, and submitting PRs 24/7
5. Most agreed code review is not super productive anymore
6. Consensus was Fable 5.5 will be stunning but OpenAI will beat it in December with Bel
on a related note, there's been some chatter recently about software engineering skills (eg the ruby/rust/elixir drama). i'm a bit torn because while i do feel like they're software engineering principles and patterns are still important, at the same time i wonder if there will come a point in time when they just won't matter anymore. not that they're no longer useful, but that your own intuition can fill in the gaps as you use agents in an intuitive sense rather than knowing explicitly what they're doing
if that doesn't make sense, the analogy i have is the first iphone. 20 years ago you had to be a professional with expensive equipment, now just with a phone and some creativity you can film content that's compelling to a new audience. i don't know if there's a term for this but i want to call it "skill compression". where what used to be gatekept and prohibitive to experts becomes accessible to a new audience of creators with a whole new meta. the details have sort of faded away and now the craft is not in knowing the equipment and how to use it, but how you use it creatively
we're going to see a whole new generation of builders who've never seen or written a line of code build incredible things
There were some really good comments on this post! The ones closest to my mental model all focused on the role people still need to play when working with LLMs. There were also some good thoughts on how to define cost/impact.
So here is how I think about prioritization with agents. Instead of the prioritization orienting around features or things to build, the prioritization framework must orient around the new scarce resource: my attention.
This is how I work now. In a world where I can build anything, I need to be super intentional with my attention. And so given a particular item of work, I think about how much attention it needs and then choose the right tool for the job. More on the tools later!
The below diagram calls out this idea of the loop. I mean this in the abstract sense - a software development lifecycle process - with observable points. When a Human is the loop, you are managing the SDLC yourself. When you have a human around the loop, you are monitoring those observable points of the SDLC yourself. When an agent is around the loop, you've built a harness or some type of adversarial review system to monitor the points of the SDLC. And finally, you can of course just ask the agent to do it all (there is no SDLC - it's just the agent).
I like this framework because it flips the focus on how to optimize workflows - you go from optimizing how you build to what you pay attention to.
volume does matter
look, i get it. if you were an engineer before agents came along, you're rightfully skeptical about lines of code and number of PRs being used as a metric for any kind of productivity. and i agree! or at least, i used to. for ease of explanation i will talk about this from the perspective as being an engineer on a large team of at least 50.
volume didn't matter before because we focused on impact. you could have high impact even with a few PRs: for example you could make a one line config change and save the company millions of dollars.
but before agents, the times where i saw volume come into the conversation were for the outliers. on the negative end you might have someone who is struggling to perform, on the other end you might have someone who was a "coding machine" archetype. a coding machine was someone who could just lock themselves in a room and emerge with a stack of PRs that solve a huge number of problems. so volume did matter, but only at the tail ends of the distribution, and combined with impact.
if you've ever worked at a large tech company before, you'll get what i mean. the thing about agents is that big tech co problems are now small-medium co problems as well.
how do agents change this calculus? well, everyone now has the ability to become a coding machine. you have incredibly capable frontier models that can write code better than any of us, at a rate much faster than we can. in this world, if you're still producing the same number of PRs as before, why wouldn't you pause to wonder why you're not being more productive? alien intelligence is here, and yet you can't outship a human coding machine?
this brings me back to why i keep talking about trust (https://t.co/7fzYZV8LGH). if you haven't put in the work to trust your agent's output, it is very difficult to scale up your productivity.
and yes, this approach does require more tokens, but i think about this in terms of cost per intelligence. before agents, this cost was very high - you needed to hire many software engineers with high 6 figure salaries to do the work. tokens are still expensive today, but are likely to get cheaper over time (https://t.co/gTBZx7pl14) factoring in the cost of intelligence. what looks exorbitant today will likely be affordable in 6 moths to a year. a single engineer with agents can do the work of tens if not hundreds of engineers, at a fraction of the cost.
the reality that i don't think has hit yet is that job of the software engineer has truly changed. our job is not merely to produce software any longer (well tbh it was never only about the code, but bear with me for the sake of the explanation), but the machine that writes the software. this idea of a software factory, or a Michelin kitchen as i like to call it, is still a topic of research. and it's the type of stuff i like to share, not to flex, but to show you what's possible when you put a lot of rigor into using agents at scale
虽然最近的硬件成本相较于一两年前很高,但是最近我还是在买内存(d4),买 CPU,买硬盘。
两点原因:
- Agent 进企业、Personal Agent 这两块都会大量抢占产能,目前都是刚刚开了个头而已。
- Agent 并发带来的工程速度直接会让 Dev、 CI/CD 负载十倍增长,进而导致 cpu 成为开发过程中的卡点。
目前消费级 cpu 的价格还没有普遍上涨,如果团队长期有需要,是值得提前入手的。
---
AI:AM: A Level We Shouldn't Pass? Notes from The Curve + Tokens vs. Salaries & Is SaaS Cooked?
原始地址:https://t.co/S50eBSPcuF
总结和校对:https://t.co/3R91jYpEtz