IT’S ONLY 3 THINGS
https://t.co/5Nco2euhU5
Did this talk (1h long) in my company. Audience were Business Analysts and Product Managers. General census: people found it useful.
So condensed and published on YT.
@michellechen Also, compiling PyTorch code (be it native, TensorRT etc) changes numerical precision and may lead to different results. Both to more accurate as well as less accurate.
eg, multiplying matrices in FP32 gives you very precise accumulation, while FP4/8/16 may lose tail precision
@michellechen Can confirm from a different perspective.
Same PyTorch code may produce different numerical results depending on CUDA version and PyTorch version itself. Even with same seed. Same GPU. These are known bugs.
Imagine what happens when you go cross-hardware and different stacks🤷♂️
i was having dinner with someone recently who asked why i spend so much time reading & posting on x. she had a pretty negative perception of the platform mostly shaped by the mainstream narrative around it.
my answer was simple. x is where the future gets beta tested.
you basically get to watch ppl build things, talk through ideas, show off stuff, & argue about where everything is going, way way before it reaches anyone else.
there really isn’t another surface on the internet quite like it.
New job listing from OpenAI:
"Operating Systems Engineer, On-Device Inference | Consumer Devices"
My guess for on-device inference: non-LLM machine learning models. E.g. speech recognition, movement, etc.
https://t.co/0USjvuuHLj
Scientists built an AI chip that runs on light instead of electricity.
And it could make every GPU on the planet obsolete.
Right now, the entire AI industry is bottlenecked by a fundamental law of physics.
Moving electrons through silicon generates friction. Friction generates heat.
Heat requires massive cooling systems and billions of dollars in power grids. We are literally running out of electricity to train bigger models.
Researchers at the University of Sydney just published a paper in Nature that bypasses this problem entirely.
They built a processor that runs on photons instead of electrons.
It processes data using light.
Instead of pushing electrical currents through copper wires, this chip guides light waves through microscopic channels.
Light has no electrical resistance. It generates almost zero heat. And it moves at the literal speed of light.
But here is the craziest part.
You can't run two electrical currents through the same wire at the exact same time without them interfering.
But you can send multiple different colors of light through the exact same optical channel simultaneously.
It unlocks massive, instant parallel processing.
This isn't an incremental upgrade. It represents a reality where AI models process data millions of times faster, using a fraction of the energy.
The global compute bottleneck won't be solved by building more nuclear power plants to cool giant server farms.
It will be solved by abandoning electricity entirely.
@elonmusk can this be the "obvious in restrospect" answer to solving biological clock / longevity?
e.g. Neuralink stimulating right neurons to regenerate/train/keep-young all the guiding parts of the body.
You genuinely have no idea how deep the rabbit hole goes.
There's a documented case of a 7-year-old girl with 82 warts, present for over a year and unresponsive to every standard dermatological treatment.
As a last resort, hypnosis was used.
Under hypnosis, she was told the facial warts would clear before the ones on her body. Two weeks later, 8 of her 16 facial warts were gone, with no change anywhere else. After three more sessions, all 82 had vanished.
And cases like this are more common than you would think.
We can measure what part of the shelf gets customer attention without any additional hardware. Just your normal CCTV cameras and a digital twin made with the phone in your pocket.
@CactusXR will absolutely revolutionize how well stores can do their merchandising.
After the C+++ rewrite, our modeling indicated RL time for these models starting at 4.8 COULD be cut in half for ~2T models and substantially more time savings on a % basis for larger models. The thing I’d suggest watching is time between releases once they get some of the kinks figured out.
It’s unequivocally better than the 2.1T that was trained with Jax, but we made some mistakes that were only corrected mid run and we have greatly improved the data quality in the past few months.
The upcoming 3T run with improved internal training software and much better data will be dramatically better.
@itslueul Grok 4.7 should be roughly on par with Opus 5.0, not 5.1. Better in some ways, worse in others. We need to fix multimodal performance.
Grok 4.8 will be a noticeable improvement.
Grok 4.9 is probably Astra/Fable class.
Grok 5 maybe better than anything. We shall see.
@OpenAI apparently is going robotics. Given the description, I would assume it's humanoid stuff.
https://t.co/fa37XieRqq
Inline with what @sama said here at 58:25 "will definitely do a humanoid": https://t.co/rMz3ECX93q
(have a routine agent monitoring for open AI orgs jobs)
at 15:30 Mark says Meta cares about social connections and connecting people...
Every time I open Facebook - I see 10 addictive reels in my timeline before seeing somebody from my friends "expecting a baby" announcement...
Meta... Social Connections...
https://t.co/Ns51a6STXm