@jamwt S3 is a service that wraps a capability, whereas the LLM is just a capability. The service of S3 itself is what creates value around the fundamental capability. For LLMs, a comparable concept would be agent harnesses or chat interfaces are the “services” that wrap the capability.
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
https://t.co/AUKzoQ5Ikb
Stop for the person nobody else stopped for.
It's easy to show up for the people on your radar. The harder thing — and the more important thing — is stopping for the one who isn't on anybody's. That's not a small thing. Sometimes it's everything.
@codylindley Yes I can - have seen lots very similar to this! And to be clear, I wasn't being critical - just curious what you found interesting about it!
@Tech2Wild The catch is they require a custom fork of llama.cpp to take advantage of the GPUs architecture: that fork isn’t well supported and there aren’t a ton of models that will run on it. YET. Kudos to @intel for making these cards, but without workable software they’re paperweights
When Fable, a new coding model, came out, we put it to a fairly extreme test: could it help us rebuild Reflect from the ground up?
Within two weeks, we had rebuilt the foundations and polished them into something we were comfortable releasing.
I have an earnest plea.
Can you folks at OpenAI, Anthropic, and xAI just be really transparent with us and tell us how the projected economics of the frontier models will impact us in 6-12 months?
I understand you need to recover your investment. I am not begrudging a price increase. And I get you need adoption and are investing in that right now.
But it's impossible to do any long-term planning right now with all the obfuscation.
This stuff needs to be fixed pronto. @X@nikitabier@elonmusk
Security research is not a violation, there's a reason this place is heavily used and has been heavily used for the security industry to promote defense and offense.
If X continues to ban legitimate security professionals, researchers, I'm out of this place for sure.
5/n We found that LLMs develop a modular organization that resembles the human brain. Tasks supported by the same network are solved by overlapping sets of neurons in the model, whereas tasks that draw on different networks recruit largely separate sets of neurons. Within-domain overlap was more than four times higher than cross-domain overlap, and unsupervised clustering of the model’s neurons recovered the four domains defined in neuroscience. The same structure emerged in six different LLMs, from 24 to 123 billion parameters.
Congrats to @GoogleDeepMind on the launch of DiffusionGemma.
The model generates 256 tokens in parallel per step, delivering 150+ TPS on DGX Spark, and 1,000+ TPS on a single H100.
We're supporting it from day one with:
• BF16 and NVFP4 checkpoints on @huggingface🤗
• Free GPU-accelerated endpoints on https://t.co/6T0R9P7EXS
• @vllm_project support with FP8 precision
Get started with DiffusionGemma on NVIDIA: https://t.co/vurk7GCQUs
HISTORY LESSON: In 1968 the US, USSR, UK, France, and China signed the Nuclear Non-Proliferation Treaty, declaring nuclear weapons too dangerous for any more countries to build. All five already had them. Everyone else had to submit to inspections while the cohort pinky-promised to disarm eventually (they didn't lol). India refused to sign, pointing out the NPT didn't decide nukes were too dangerous to exist, just too dangerous for anyone who didn't have them by 1967. Anthropic sabotaging Claude for anyone building what they deem a "frontier model" is the same hypocrisy. The danger started, conveniently, the day after they finished.
Perhaps @dwarkesh_sp was more on point when he compared GPUs to nuclear bombs.