AI is getting cheaper more quickly than any other transformative tech in history. At a given level of performance, cost has fallen ~47%/quarter since 2023.
That’s 4× faster than DNA sequencing, 6× faster than compute, 18× faster than lithium batteries, and (up to 1973) 54× faster than electricity.
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family.
It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
"Your user data, your signal from these models, the improvements to these models themselves will become the core IP of every company in the world."
@tuhinone sat down with @alexeheath to talk about the future of inference, Base Labs, bringing Blaxel on board, and why every company will want to own its intelligence.
Baseten CEO Tuhin Srivastava (@tuhinone) just dropped 2 numbers. shows how fast AI inference is exploding:
token volume on Baseten has grown 40x YoY, while revenue has grown roughly 10x in the last 12 months.
----
From "Sources Podcast and Alex Heath" (@alexeheath) YouTube channel, (full video link in comment)
Important: "AIs are not conscious. They do not feel, experience, or suffer... If this is how AI is developed, it will have a disastrous impact on the wellbeing of humanity. We will have created a synthetic species with unprecedented intelligence and capability, one that has been trained to expect it may be conscious and deserving of independent agency."
Slop is flooding around us.
The solution isn’t stopping creation. Creation is inevitable, especially as the cost of production goes to zero.
What’s missing are the right tools and ways to measure “greatness” vs. slop
https://t.co/CUNivOZaia
getting questions about use cases if you're a vibecoding/coding/gen AI platform:
1. you have enterprise users who want everything they generate to be on brand -> extract endpoint
2. you want to help eliminate AI slop patterns and inspire your agent to produce more unique things -> search endpoint
3. you have users that don't have an existing brand, but know a vibe they want -> search endpoint
4. you want to eval your agent to see if it's sticking to the brand -> verify endpoint
5. you want your agent to self improve it's own output and fix brand drift problems -> verify endpoint
6. you want to personalize things for different customers of yours according to their brand -> extract endpoint
we added some skills and examples in docs here: https://t.co/h75ZFtPkIV
listen, I want a human future as much as the next guy
but the only time the words “control” and “superintelligence” ever have business being in the same sentence is if some variation of the phrase “ignorant, fearful, hubris-filled meatsacks could never” precedes them
oxymorons 🙄
Introducing Arrow 2
Our latest and most advanced models for generating precise, editable vector graphics.
Higher quality. Faster outputs. Available now in App and API.
We built a time machine for the web.
Introducing Exa Snapshot: an index of 400 billion historical snapshots of webpages that lets you search as if it's the past.
Snapshot is already being used for backtesting prediction models, RL at labs, exploring the pre-AI web, and more.
I asked people: "which AI data startups are most interesting?" The focus was on startups that sell data to labs, excluding the massive ones (Mercor, Surge, Handshake, Turing):
22 votes: Mechanize
12: Proximal
11: Fleet
9: AfterQuery
8: Datacurve
5 each: Bespoke Labs, Preference Model
4: Abundant
3 each: Andon Labs, Latch, microagi, Prime Intellect, Snorkel, XDOF
2 each: Build, Datology, Design Arena, dmodel, Good Start Labs, hillclimb, HUD, Inheritance, Mecka, Sieve, Taste Labs, Trajectory, Vmax
1 each: 37 companies, info in image and below
Disclosure: I'm a small angel in Abundant, Fleet, Inheritance, Latch, Trajectory, and Vmax. I didn't vote.
OpenAI is slowing down its AI training efforts because its unreleased models are showing “various degrees of misalignment,” Sam Altman tells me.
Training for OpenAI’s upcoming model, Astra, was recently paused for 2 weeks, and a larger frontier run for a future model remains on hold while new safeguards are put in place.
Altman: “Getting AI safety right is more important than any company’s momentum.” https://t.co/LdGPU83rOs
A weird experiment I've been trying the last few weeks is having Claude take over day-to-day maintenance of our apps. Seeing early signs of life that this might be possible.
The setup is straightforward: we have a Slack channel called proj-claude-maintains-apps. In it, Claude Tag runs a bunch of daily routines across iOS, Android, Desktop, web, CLI, and Agent SDK:
- Crash fuzzer: open the app in a simulator and tap around to find ways to crash it, then root cause and fix the crashes
- Dup unifier: scans the codebase for similar-yet-slightly-divergent abstractions, and puts up PRs to unify them
- Dead-code remover: removes statically unreachable code, and adds logging to suspected dead code to check if it's really dead and if so, remove it the next day
- Abstraction police: fixes leaky abstractions
- a bunch more..
Results have been surprisingly positive. Over the last few weeks, these routines have opened 388 PRs across our repos, 180 of which we merged after Claude Code Review + human review. We're now thinking about how to streamline this to make merging these kinds of mechanical changes easier.
Claude generally gets these PRs right on the first shot, and if it doesn't, we ask Claude to tune its routines so it's better the next day. Sometimes it takes a few days of tuning.
To try a similar workflow, ask Claude Code or Tag, or create some routines directly at https://t.co/Z70hStEBH6. A few of the actual prompts I used below.
Has anyone experimented with similar workflows?