wondering why I feel exhausted. maybe: the agents do all the easy stuff, and I have to work through the leftover hard bits, which means I'm perpetually locked in. and as the models get better, "my" work just gets harder and harder, until I'm basically underqualified to do the work (which... is better than the alternative, there's nothing left for me to do, and I'm paperclipped).
i'm restarting my blog! i want to kickstart productive conversations around: what should AI agents look like for hard, subjective knowledge work?
a lot of agent setups work well when tasks are objective and easy to verify. but many workflows (e.g., qualitative analysis, strategy, sensemaking) are messy and interpretive.
as a first post, i explore different ways of doing agent-assisted qualitative analysis on tweets, with varying levels of human feedback/intervention.
tldr: they all kinda sucked. turns out it’s hard to:
(a) stop agents from converging too quickly on shallow interpretations
(b) get agents to adapt to preferences that emerge gradually across many turns (i.e., evolving context)
(c) capture human judgment without making humans fatigued
I strongly believe there are entire companies right now under heavy AI psychosis and its impossible to have rational conversations about it with them. I can't name any specific people because they include personal friends I deeply respect, but I worry about how this plays out.
I lived through the great MTBF vs MTTR (mean-time-between-failure vs. mean-time-to-recovery) reckoning of infrastructure during the transition to cloud and cloud automation. All those arguments are rearing their ugly heads again but now its... the whole software development industry (maybe the whole world, really).
It's frightening, because the psychosis folks operate under an almost absolute "MTTR is all you need" mentality: "its fine to ship bugs because the agents will fix them so quickly and at a scale humans can't do!" We learned in infrastructure that MTTR is great but you can't yeet resilient systems entirely.
The main issue is I don't even know how to bring this up to people I know personally, because bringing this topic up leads to immediately dismissals like "no no, it has full test coverage" or "bug reports are going down" or something, which just don't paint the whole picture.
We already learned this lesson once in infrastructure: you can automate yourself into a very resilient catastrophe machine. Systems can appear healthy by local metrics while globally becoming incomprehensible. Bug reports can go down while latent risk explodes. Test coverage can rise while semantic understanding falls. Changes happens so fast that nobody notices the underlying architecture decaying.
I worry.
There are two related, but distinct, problems with MTTR maximalism.
1. The distribution of recovery times could be heavy-tailed, and so the empirical mean could be far from the true mean.
2. Some failures are unrecoverable (e.g. durability loss).
i think wittgenstein would call this a "gramatical statement" rather than a "statement of fact"
we're now arguing terminology, rather than how the process works
The new generation of open state-of-the-art single and multi-vector retrieval models is here
It's time, DenseOn with the LateOn 🎶
@LightOnIO releases models that leap past existing ones, and everything you need to do the same!
The Adam optimizer is at the heart of modern AI. Researchers have been trying to dethrone Adam for years.
How about we ask a machine to do a better job? @GoogleAI uses evolution to discover a simpler & efficient algorithm with remarkable features.
It’s just 8 lines of code: 🧵
I agree wholeheartedly with the criticism of the way the Conly Cochrane meta-analysis dismissive of masks has been conducted. But—sorry, team—I need to add some wee quibbles from a philosophy of science perspective. 🧵https://t.co/XuwEs6LUtc
India just passed China as the most populous country in the world. Why?
Because of the biggest accident in history
Look at where people live in India. What's that band up north?