@unclebobmartin They serve different purposes: harnesses give agents autonomy within constraints. Interrogating them adds interruptions. I do this in plan mode. If something's wrong, I update harness. Over time, it accumulates rules. Eventually, the harness itself will no longer be necessary.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here: https://t.co/OGyPb7yaYt
We're giving away $10,000 for your savings.
Why? To celebrate.
We built the first proactive AI Financial Advisor. So you can stop thinking about your money.
To enter:
1️⃣ Repost this
2️⃣ Mention @useorigin
3️⃣ Include #OriginGiveaway
Prize varies by Origin membership status. Official rules here: https://t.co/huLn9uTpir #OriginGiveaway
We wanted to build a product so that you’d never think about your money again.
It’s a strange goal for a personal finance app, but when we thought deeply about the best experience we could possibly build, it was one where our users could forget about the biggest stressor in their life…their money.
A year ago we launched the first AI financial advisor. It unlocked financial advice to hundreds of thousands of our members. The problem we soon realized was that chat assumes people know what to ask, that they know how to ask it, and that they'll remember to keep asking, every month, for the rest of their lives.
Unsurprisingly, very few people do that which results in thousands of dollars in financial mistakes that pass by our members each year.
So we took those learnings and we built a new Origin Advisor. This one knows everything about your financial life. It knows the past decisions you’ve made, the ones you’re avoiding and the ones you don’t even know exist.
We now deliver all of this to you every day in our app. And soon it won’t just tell you what to do. It’ll do it for you.
I’m incredibly proud of our team’s work. This type of financial advice used to be for people with a lot of money and now it's available to anyone with a phone. The best part is you don’t need to know anything about money, understand it or even think about it.
I am pleased to see that OpenAI’s new model is roughly on par with Gemini Flash and Opus 4.8 on prompt injection risk. Nice work!
Evaluating and naming other labs turns out to be a great way to encourage them to train more aligned models. We will continue to do this until other labs pay more attention to safety. This is good for everyone and there is a lot of room left to go!
We solved prompt injection in practice for Claude models about two months ago. But prompt injection is a significant security risk no matter what model you use, and it is important that the industry similarly spends more effort to train their models to be resistant to prompt injection, among other elements of model alignment.
As models become more capable and central to businesses and economies, the risks only increase. We should be taking them seriously, and doing the right thing for our customers and the world.
/show-me is a phenomenal skill
Makes PR descriptions extremely easy to read
Basically a toolbox of "nice ways to look at code"
Nice work, @dexhorthy
https://t.co/phB1dAPwgG
How to maximize the value of your Claude Code sessions: https://t.co/YGa8swOX6v
Solid guidance on managing token efficiency and keeping context windows focused by @lydiahallie
1/2 Thanks Gavin for an especially thoughtful exchange. I don't usually spend much time on social media but I wanted to engage here because it really brings out the heart of an important conversation.
First, on regulation, I think that “either concentrate it in the hands of a chosen few companies and politicians via regulation or distribute it widely” is a false choice. I know that there’s a sort of Silicon Valley shorthand where regulation = regulatory capture = concentration of power, but I’ve always found this to be an overly simplified picture of the world. Many people outside this bubble think of regulation as something that constrains corporate power and benefits ordinary people. I don’t necessarily agree with that perspective either, rather I think it’s complicated and really depends on what the “regulation” consists of. But in particular I think that those in the “regulation = regulatory capture = concentration of power” frame often underrate the decentralizing power of objective and fair institutional processes. A crude analogy is that the formal court system can sometimes feel stuffy and elitist, but it does a much better job of defending the rights of vulnerable individuals than the alternative, mob justice. At their best, institutions can vest power in ideas rather than people, and thereby decentralize that power.
This is why Anthropic has always made its policy proposals very carefully. We try very hard to make proposals that disadvantage (slow down) frontier AI companies while *advantaging* smaller competitors. California’s SB53 (which we supported), and even the much-maligned SB 1047 (which we were ambivalent on), completely exempt any company below a certain amount of revenue or model training costs from being covered at all (it was $500M for SB 53, lower for 1047 but we objected to that). More recently the testing process we’ve advocated for at CAISI and the White House involves more rigorous tests for frontier models than off-frontier models — something that differentially advantages challengers. Similarly, the “Pacing the Frontier” letter envisions (or at least Anthropic’s preferred implementation of it envisions) modulating the pace of the very best models while not constraining those who are catching up. This hurts the business interests of the frontier labs and helps challengers, including open-weights!
Overall my view is that AI is *structurally* a technology that tends to concentrate power, for reasons that have nothing to do with regulation (more to do with the extreme implications of the scaling laws). Open-weights do help some with this but are nowhere near a sufficient solution because they simply shift the concentration somewhat to those with the most compute and chips (which are roughly the frontier labs plus maybe hardware providers). By contrast I think the right “rules of the road” can simultaneously (a) address AI’s cyber/bio/alignment risks, (b) institutionally constrain the power of the frontier AI companies, and (c) leave room for open-weights models while also addressing the specific risks that they bring.
BTW I do not think that the events of the last few months have “failed to result in [my] preferred regulatory path”. The approach that the Trump administration is reported to be taking — pre-deployment testing for frontier models, and also testing of open-weights models when they get closer to the frontier — is one that I am very supportive of, though of course I have to see the details to be sure. I am also supportive of Demis Hassabis’ ideas around a FINRA-like entity. This contrasts with six months ago when most of the industry was still pushing for preemption of all state regulation and no apparent federal approach either.
You can now book two humanoids for our cleaning service in SF. One humanoid for 60 minutes is $30, two are $60.
You can book a cleaning at https://t.co/Brl42KyISx. If you don’t have an invite code yet, join the waitlist. Timelapse at 5× speed (~40 minutes in real time).
A weird experiment I've been trying the last few weeks is having Claude take over day-to-day maintenance of our apps. Seeing early signs of life that this might be possible.
The setup is straightforward: we have a Slack channel called proj-claude-maintains-apps. In it, Claude Tag runs a bunch of daily routines across iOS, Android, Desktop, web, CLI, and Agent SDK:
- Crash fuzzer: open the app in a simulator and tap around to find ways to crash it, then root cause and fix the crashes
- Dup unifier: scans the codebase for similar-yet-slightly-divergent abstractions, and puts up PRs to unify them
- Dead-code remover: removes statically unreachable code, and adds logging to suspected dead code to check if it's really dead and if so, remove it the next day
- Abstraction police: fixes leaky abstractions
- a bunch more..
Results have been surprisingly positive. Over the last few weeks, these routines have opened 388 PRs across our repos, 180 of which we merged after Claude Code Review + human review. We're now thinking about how to streamline this to make merging these kinds of mechanical changes easier.
Claude generally gets these PRs right on the first shot, and if it doesn't, we ask Claude to tune its routines so it's better the next day. Sometimes it takes a few days of tuning.
To try a similar workflow, ask Claude Code or Tag, or create some routines directly at https://t.co/Z70hStEBH6. A few of the actual prompts I used below.
Has anyone experimented with similar workflows?