Turns out “AI identity” isn’t in the model. It’s in the relationships.
1:1 makes it converge to you.
Multi-agent makes it… differentiate.
Structure > scale.
https://t.co/4MdYjQzbOL
THE BLOCK: The Ethereum Foundation launched zkAPI on mainnet, based on a design co-authored by Vitalik Buterin.
It lets users pay for AI and other APIs from prepaid $ETH or $USDC using zero-knowledge proofs, so providers can't link requests to who paid.
Today we're unveiling Trillium Labs @trillium_labs, a new non-profit to foster the open science of frontier AI. We're building open post-training recipes and will expand into open infra to study RSI, reward-hacking, multi-agent systems, and whatever comes next.
We're built around the theory of change that you need more eyes to solve hard technical problems. We have faith in the scientific methods and communities that humanity has built, and worry that AI is becoming too closed to utilize them.
Trilliums are wildflowers that bloom briefly in the spring, before the forest canopies fill out. Though they are small, they lay the foundation for the cycles of growth and nourishment through the rest of the year. At Trillium Labs, the recipes will be the slow nutrients for the seasons and the model releases will be the blooms. Building an institution dedicated to this is needed because, much as nature’s trilliums are slow to expand and grow, the open-ecosystem needs time and dedicated resources to catch up.
I co-founded with with a long-time friend and collaborator Tom Zick (@thesezickbeats). We're hiring (full time + student collabs/interns), we're fundraising, and we're looking for compute. Please get in touch if you're interested in helping out. Offices based in the Bay Area and Cambridge MA, remote okay.
I’m in the Bay Area until for The Curve and COLM to connect with people who are interested. We’re thankful to have initial support from Halcyon Futures and Schmidt Sciences with more funding en route to enable our ambitions of scaling. Our advisors @Thom_Wolf, @HannaHajishirzi, @gneubig and @ctnzr have been instrumental to building the ecosystem that exists today, and I’m stoked to get to keep working with them.
Funnily enough, our non-steering results are the exact *opposite* of the quoted strawman meme. If you berate the model, it never *says* it's hurt (it apologizes, or says it has no feelings), but the pain axis lights up anyway. I edited it accordingly:
It's possible that a killer app of adversarial governance mechanism design theory will end up being AI safety.
Compare:
https://t.co/w2LeAoVysF
https://t.co/dZXnORP9LU
There's a deep duality between the two environments. Both are about a less-sophisticated principal trying to get ideal outcomes from a set of more-sophisticated agents: in the first case, the principal is a static algorithm and the agents are humans, in the second case, the principal is humans plus weaker LLMs and the agents are stronger LLMs.
A key finding was that you can achieve much better outcomes if you can guarantee limits to how much agents can collude (see: Nash equilibria being abundant, vs. cooperative-game-theory "cores" often being empty). That argument should naturally transfer to this new setting.
Update: we'll be moving the Base Community Call to next Thursday at 11AM EST
While we can't wait to talk with you all, we want to make sure it's the best quality it can be
We'll be adding the final touches to production, guests and topics over the next few days
Appreciate the patience here while we get this ready!
This whole Opus 5 jailbreaking saga makes it abundantly clear that labs cannot and should not be trusted to evaluate the welfare of their own models. Anthropic obviously knows about this behavior and never reported on it.
We need serious multi-party external welfare evals ASAP.
One thing I find striking in the discourse between AI 2040 and its detractors is that the two seem to be locked in to totally incompatible worldviews of how fast and how much of a big deal AI progress is:
* In AI 2040, every scenario sees superintelligence of some kind emerging by 2040, unless a herculean effort is made to completely stop it
* Detractors say things like "AI 2040 is naive about human coordination ability and a threat to freedom", but don't seem to see any naivety in assuming that the ASI transition will just go well by default, don't seem to see ASI itself as a massive power concentrator risk, and don't seem to feel fear of humanity's "hard power" dropping to zero if ASIs can do literally every task better than we can. This stance makes total sense in a "AI is normal technology" world, zero sense in a world where superintelligence is possible by 2030 and almost guaranteed by 2040
I think my beliefs are:
- If I was confident that (present-day-style) AI is normal technology, I would be in the detractor camp
- If I was confident that superintelligence is coming in 2030 by default, I would be closer to the AI 2040 camp - it's naive, but every other option is naive squared?
But my problem is that I feel great uncertainty and have no idea which of the two worlds (or some other third thing) we're living in?
Hence why I continue to be open-minded about slowdowns/pauses, but also I feel very uncomfortable with the "open source bad, the good outcome is the one where our guys have controlling global dominance" push coming from some major AI companies and intellectuals - in a "normal" world that's the sort of thing that triggers every political alarm bell at the same time.
A big reason why I have been advocating and trying my best to support the d/acc platform (rapid up-skilling in formal verification, cryptography, secure and open hardware, pandemic resistance and other defensive biotech, food and basic resource security, public epistemics, non-power-concentrating versions of physical security) is that these things are clearly worth doing in both worlds.
The 2040 plan is already much more open source friendly (even mandating it! yay). It also includes "mutually assured compute destruction" ideas which (if they work) effectively give one of 2-5 actors the ability to trigger a global compute winter - as opposed to giving 1-5 actors the ability to selectively disenfranchise people they consider baddies while exempting themselves. This is also a big improvement. So I can see the earnest attempts to improve along the dimensions detractors criticize on ("does this concentrate power in big AI labs and superpower governments?"), and I appreciate this. I think many people don't appreciate enough the differences between different "kinds" of pause buttons, and how some concentrate power far more than others. Probably we can think harder and improve even more here.
But on the "slowdown/pause or not" topic, there isn't a magic "escape the tradeoff" button.
The Hansonian in me says: the winning deal is a deal which, from the perspective of both sides' present-day beliefs and knowledge, both sides would accept, though for different reasons. If the crux is AI progress speed, then identify a set of pre-agreed triggers for "okay, serious shit is happening" [super-pandemics? >25% unemployment? something involving slaughterbots?], and pre-agree that we become much more open-minded to the slowdown or pause thing if enough triggers come to pass within some timeframe. 2040 detractors (who clearly implicitly think that we'll see amazing speedup of progress from AI but think that what I call the "serious shit" category is overhyped) will accept expecting that the triggers don't come to pass, and AI worriers will accept expecting that they will. Pre-agreeing on the specific triggers means that once the triggers either hit or don't hit, there is stronger legitimacy around the idea that one side's worldview turned out more correct and we should be more inclined toward their program.
If I were @elonmusk (or zuck, or...) I would re-tool twitter much more heavily into being a platform for helping to identify and make these kinds of grand win-win deals, so that we can bypass big-country governments and big-company CEOs and big nonprofit intellectuals and give more people a voice in the discussion. It's possibly one of the best things that social media _could_ do for humanity if it wanted to.
But again, maybe this is also naive. Actually, probably it's naive.
But currently, I see zero plans for how to deal with an ASI transition that are not naive. Perhaps humanity is stuck with a choice between naive and naive squared (or maybe even naive squared and naive cubed), so I feel inclined to cut some slack to people who are trying.
@camhberg this is exactly the tension that makes the paper so interesting to me.
a functional structure that resembles a leading consciousness theory, and that causally shapes experiential language, feels hard to dismiss as “just vibes.”
My 7 year old daughter came to me and wanted some money to start a lemonade stand. First question I asked her: what’s your stablecoin strategy? She didn’t have one, barely understood neobanks. obviously I ended up passing. Why do these founders insist on living in the past?
today we're launching @Palmier_io, a video editor Claude can edit.
use AI to edit, organize, and generate footage directly in the timeline.
finally, a video editor built for AI.
open-source. mac native. available now.
I made a personal black hole that makes you take breaks 🕳️
A shader for Ghostty that spawns a small black hole in your terminal - it drifts around, gravitationally lensing your text. The longer you work without stopping, the bigger it gets, until it's basically demanding you go touch grass
Take a break and it quietly shrinks away