White phosphorus ignites instantly on contact with oxygen. It burns at over 800°C. It melts through clothing, skin, muscle — and bone.
In the bloodstream, it becomes a systemic poison. It attacks the heart, liver, and kidneys leading to multi‑organ failure and death.
Monsters.
I’m hiring a Right Hand to work alongside me at Sudowrite. Our company is full of uniquely wonderful and kind people that you can’t help but love.
It's also profitable. (We're weird!) And we do it making creative tools for artists, not B2B SaaS.
This is an incredible role for someone who’s done ops, exec support, been a chief of staff, or worn a lot of hats at a startup.
🌈🌴 On-location in Honolulu preferred but not required.
https://t.co/l6mEUUpxGG
Instead of forcing models to hold everything in an active context window, we can use hypernetworks to instantly compile documents and tasks directly into the model's weights. A step towards giving language models durable memory and fast adaptation.
Blog: https://t.co/iHoifpsLMu
@cHHillee@yoavgo Zachary was a few years ahead of me in school-a truly brilliant person. My partner studied with him for the Putnam, and I still regret not joining him for ICPC back then! It’s so heartening to see that his work is still loved by so many!
@soumithchintala Oh nooo. I was in high school when PyTorch was open sourced and used to read GitHub issues and commits; you were always an inspiration! Among other things, PyTorch taught me that the right abstractions matter. Thank you for all your work for the community and best of luck!
@teknium It would be cool if they've figured out some form of TD learning (for assigning credit to individual statements or something) and not a PPO variant with a fancy objective :c
@natolambert Claude code just fixed a bug (a somewhat trivial one) in Megatron distcp to consolidated safetensors export so at this point I’d rate it slightly above “absolutely useless” :)
1/ We’re thrilled to announce we've raised $7.4M in total funding led by @InsightPartners along with @ZeroPrimeVC, @AIXVenturesHQ, @468Capital, and @FirestreakVC to bring evals and observability to AI agents.
HoneyHive Cloud is also generally available starting today.
So I suppose the best thing to do today is to stare at some o3 output data on ARC-AGI. Here's a simple visualization on the public eval with o3's attempts and gt solutions. (Pls don't spam)
https://t.co/2MTBQlCNnW
@jasondeanlee Grad level+undergrad/IMO style but I think it's only "hard" for mathematicians because the problems span too many domains, and the questions were written by domain specific experts
o3 evals on FrontierMath being done pass@1 is quite honestly the most disturbing aspect of this. Why aren't more people talking about it? And also that we're likely closeish being able to have end to end AI-driven AI research given that SWEBench score??
@ds3638 Yeah but if the model is on average generating ~100m CoT tokens in one go then I'm sure it can run grep enough times with tool use within CoT to address retrieval for a bunch of tasks. Not sure about error prop but backtracking is ~emergent with RL+verification.
And for all those harping about cost: efficiency does improve as evident from AI2's work, that synthetic data seems to largely be helping with learning better abstractions so why wouldn't these absurdly long CoT dumps? Even simple, "known" methods would lead to a 100x decrease!
The last few times it was immediately clear that something represented a breakthrough was when ResNets did exceedingly well on ILSVRC and transformers did the same with BLEU etc.