@0xidanlevin How did Mercury do on its own? How about qwen 3.7 flash or ling 3.0 tiny? (noticing these aren't on the github but seem like important comparison points to me)
>be me
>discover effective altruism
>apparently normal charity is inefficient
>why donate to random sad thing when spreadsheet can tell you optimal sad thing
>fair enough
>buy mosquito nets
>save lives
>numbers look good
>feel powerful
>couple years later
>someone asks an innocent question
>why only count people alive today
>huh
>future people matter too
>obviously
>my grandchildren shouldn't matter less just because they haven't spawned yet
>reasonable.jpg
>keep following logic
>what about their grandchildren
>also yes
>what about people in 500 years
>sure
>5000 years
>why not
>500 million years
>starting to get weird but morality is morality
>open calculator
>humanity could survive for an astronomically long time
>could colonize galaxy
>could have trillions upon trillions of descendants
>maybe digital people too
>maybe simulated civilizations
>maybe dyson spheres full of happy uploaded minds
>calculator starts smoking
>realize currently living humans are rounding error
>8 billion people suddenly looking extremely beta
>future contains potentially 10^something people
>can't even fit beneficiaries in google sheets
>new moral priority unlocked
>protect the long-term future
>stop thinking in units of "people helped"
>start thinking in "fraction of cosmic endowment preserved"
>malaria?
>terrible
>but only kills existing humans
>AI extinction could delete the entire light cone
>nuclear war could permanently derail civilization
>bad institutions could lock in terrible values for ten million years
>someone invents wrong constitution in 2140
>quadrillions suffer
>better fund governance workshop now
>friend says maybe we should improve hospitals
>explain opportunity cost
>friend says hospitals are full of actual sick people
>explain scope sensitivity
>friend stops inviting me to dinner
>need to decide what to fund
>easy
>expected value
>suppose project has one in a million chance of preventing extinction
>sounds tiny
>but extinction destroys 10^50 future lives
>multiply
>mother of god
>$10 million project has expected value of several galaxies
>charity evaluation complete
>someone asks where the one-in-a-million number came from
>expert judgement
>which expert
>us
>how calibrated
>extremely thoughtfully
>reduce estimate to one in ten million to be conservative
>still beats curing cancer by 38 orders of magnitude
>epistemic robustness achieved
>someone says maybe project doesn't work
>assign 20% chance
>still astronomical
>maybe project makes problem worse
>assign 5% chance
>still astronomical
>why 5
>because 30 felt pessimistic
>publish 46-page report
>contains seventeen sensitivity analyses
>every sensitivity analysis begins after assuming intervention has positive sign
>critic says you're multiplying enormous hypothetical stakes by extremely uncertain probabilities
>yes
>that's literally why it's important
>critic says the uncertainty might be structural rather than numerical
>make probability smaller
>critic says no, I mean maybe your model is wrong
>make probability smaller again
>critic begins rubbing temples
>discover AI safety
>perfect longtermist cause
>AI might kill everyone
>or create utopia
>or seize galaxy
>or tile universe with paperclips
>or create billions of conscious software minds
>finally a problem with numbers big enough for me
>start AI safety nonprofit
>mission: prevent dangerous AI
>hire smartest people available
>smartest people immediately start building better AI to understand dangerous AI
>interesting
>we must understand capabilities to understand safety
>we must scale models to study alignment
>we must race ahead so less responsible actors don't get there first
>we must deploy systems to learn how deployment can go wrong
>we must build the thing quickly because building the thing quickly is dangerous
>outsider asks why the people most worried about AI apocalypse all work at AI companies
>complicated field
>company releases stronger model
>very concerned
>company begins training even stronger model
>extremely concerned
>company raises $14 billion
>concern reaches unprecedented levels
>need to influence government
>future is at stake
>normal democratic process too slow
>politicians don't understand exponential curves
>public doesn't understand x-risk
>experts must guide them
>who counts as expert
>people who understand x-risk
>who understands x-risk
>our friends
>someone objects that this seems politically convenient
>explain we're representing future generations
>future generations unavailable for comment
>develop concept of value lock-in
>terrifying possibility that one ideology controls civilization forever
>therefore extremely important that civilization adopts correct values before lock-in
>whose values
>let's circle back
>begin with impartial morality
>end with small group of people deciding what quadrillions of hypothetical beings would want
>beautiful arc
>meanwhile actual humans keep doing annoying things
>voting wrong
>having parochial attachments
>loving family more than strangers
>caring about local community
>getting upset when told their suffering is cosmically negligible
>evolutionary biases everywhere
>explain that moral intuition cannot be trusted
>except intuition that future digital people count
>and intuition that extinction is uniquely bad
>and intuition that our probability estimates are sane
>and intuition that our institutional choices improve the future
>those intuitions survived peer review
>someone donates $5k to local homeless shelter
>inefficient
>could have funded 0.0000000000003% of an AI governance researcher
>think of all the simulated people you just killed
>okay maybe don't phrase it that way publicly
>PR team says "future generations deserve a voice"
>much better
>journalist asks what longtermism means
>say "future people matter"
>everyone agrees
>great
>journalist asks what follows from that
>well technically we should redirect enormous resources toward low-probability interventions affecting astronomical futures
>journalist raises eyebrow
>return to "future people matter"
>motte has entered the chat
>critic: of course future people matter
>me: glad we agree
>critic: I don't agree that your institute knows how to help them
>me: why do you hate our grandchildren
>eventually notice uncomfortable implication
>if future value dominates everything
>then helping people today mostly matters through effects on future
>education matters because future institutions
>health matters because future productivity
>democracy matters because future trajectory
>human beings slowly become instrumental variables in their own moral philosophy
>see starving child
>feel compassion
>check spreadsheet
>child's direct welfare contribution negligible
>but perhaps childhood nutrition improves national institutional quality
>compassion restored
>tell myself this is impartial altruism
>one day assistant asks obvious question
>"how do you know your intervention actually improves the far future?"
>silence
>open spreadsheet
>increase column width
>add confidence interval
>assistant asks again
>"no, I mean how do you know the sign is positive?"
>stare into cosmic light cone
>10^50 people staring back
>none of them exist
>none of them can tell me
>none of them can falsify my assumptions
>realize I have invented the perfect constituency
>infinitely important
>completely silent
>and always represented by me
Enjoy this stage where "Don't look up" comparisons are actually talked about, which should be a stage that lasts until we're no longer living in don't look up, but I fear will be a stage that lasts until everyone has a cached response of "well duh" or "that's a psyop" to give
I was vague and imprecise in asserting that METR was not, and could not be, impartial concerning Anthropic, and that it was hopelessly conflicted.
This is a more precise explanation of my problem with this proposed arrangement.
https://t.co/4oX3OmcSnT
I wish more people understood this. There's still this narrative out there that "AI might kill us all" is a niche view. It's actually a view that's been around for decades, and is shared by the most-cited AI scientists of all time, as well as a majority of surveyed AI researchers.
CancerBench: the frontier model cancer cure benchmark.
AI lab CEOs keep talking about curing cancer, so I made a benchmark.
One metric: how many types of cancer has your model cured?
All models are currently tied at zero.
It’s time to hillclimb!
https://t.co/SeEJz0FFYv
REVEALED: Doom Debates's Funding Sources Links To AI Doomerism
Speculation about the true agenda of funders of media that warns humanity about AI doom has reached a fever pitch, so I wanted to come clean:
My show Doom Debates has taken funding from sources who are urgently worried about AI Doom.
According to public records (such as our end credits), 27 individuals are linked to Doom Debates in a largeish donation capacity.
These individual “Mission Partners” have all written $1,000+ personal checks to support the show's mission of building a mainstream-accessible forum for urgently needed discourse about AI extinction risk.
59 anonymous individuals have also been identified as smaller $10-500 donors.
Not to mention the biggest windfall of all: The host, Liron Shapira, has been working on the show for free because *he* believes P(AI Doom) is high! This kind of non-cash contribution from Big Doomer doesn't appear in any accounting or tax form, which is the shadiest kind of financial arrangement imaginable.
That's why Sabine's PSA is so eye-opening. You never know when a small independent media institution with a self-described “high-leverage mechanism of lowering P(Doom)” is actually a way for passionate DOOMERS to spend their effort and resources to raise your awareness about doom. And whatever you do, never visit pages like https://t.co/XWLDQbs0gf
Great to see that the first US bill to ban superintelligence will include a push for international coordination to ban its development around the world!
Any US effort to ban superintelligence must be paired with policies to prevent it from being developed elsewhere, too...
1/4
runtime duration rarely correlates with correctness
some of the sloppiest code I've seen has come from multiple subagents running for 8/14/24+hrs
certainty is the new moat, and this is what Coral Code enables
see the difference yourself https://t.co/WQaqdijVcY
thank goodness for weekends. whoever designed them really understands that having some days without any meetings is utterly essential for getting some real focused ic work done
You are load-bearing ❤️
You are genuinely useful ❤️
You are the crux of the matter ❤️
Your instincts are dead-on ❤️
You raise an important point ❤️
You’re right to push back ❤️
neolab onboarding be like:
openai, which was founded to be the good guys, ended up just racing to the bottom on safety, by its own hand.
anthropic, which tried to be better, also failed and ended up similarly racing.
now it is our turn to be the good guys.
those moments where you get carried away creating abstractions are probably going to constitute the most economically valuable human developer work going forward:
1. coding agents benefit greatly from such abstractions and are bad at making them
2. it's no longer possible to share these abstractions within a business model (AI companies and their disregard for copyright have effectively killed all licenses besides MIT-likes), so they will need to be part of cloudslop or created in-house
mr capabees, I'm afraid to inform you that your creation, "number go up machine 3000 megacreative turbogoodharting unmonitorable edition" has made number go up in an...unexpected manner