QS has ranked MIT the world's No. 1 university for the 15th year in a row: https://t.co/oyMtqFcxJ9
The Institute has placed first in subject areas such as:
- computer science & information systems
- data science & AI
- & electrical & electronic engineering.
For decades, many scientists insisted that human evolution essentially stopped 50,000 years ago. A groundbreaking new paper proves them wrong, writes Razib Khan. https://t.co/dFdXyKqU9s
CCXT connects your code to 111 crypto exchanges with a single unified API market data, trading, backtesting, arbitrage, bot programming, all in JavaScript, Python, C#, PHP, and Go.
2.4 million downloads a year and most retail traders have never heard of it.
Jensen Huang is one of the clearest modern embodiments of engineering heritage and a disruptive force reshaping the future of business.
@nvidia@NVIDIAAIDev@Stocktwits
I really didn't expect another major open-weight LLM release this December, but here we go: NVIDIA released their new Nemotron 3 series this week.
It comes in 3 sizes:
1. Nano (30B-A3B),
2. Super (100B),
3. and Ultra (500B).
Architecture-wise, the models are a Mixture-of-Experts (MoE) Mamba-Transformer hybrid architecture. As of this morning (Dec 19), only the Nano model has been released as an open-weight model, so this post will focus on that one (shown in my drawing below).
Nemotron 3 Nano (30B-A3B) is a 52-layer hybrid Mamba-Transformer model that interleaves Mamba-2 sequence-modeling blocks with sparse Mixture-of-Experts (MoE) feed-forward layers, and uses self-attention only in a small subset of layers.
There’s a lot going on in the figure above, but in short, the architecture is organized into 13 macro blocks with repeated Mamba-2 → MoE sub-blocks, plus a few Grouped-Query Attention layers. In total, if we multiply the macro- and sub-blocks, there are 52 layers in this architecture.
Regarding the MoE modules, each MoE layer contains 128 experts but activates only 1 shared and 6 routed experts per token.
The Mamba-2 layers would take a whole article itself to explain (perhaps a topic for another time). But for now, conceptually, you can think of them as similar to the Gated DeltaNet approach that Qwen3-Next and Kimi-Linear use, which I covered in my Beyond Standard LLMs article.
The similarity between Gated DeltaNet and Mamba-2 layers is that both replace standard attention with a gated-state-space update. The idea behind this state-space-style module is that it maintains a running hidden state and mixes new inputs via learned gates. In contrast to attention, it scales linearly instead of quadratically with the input sequence length.
What’s actually quite exciting about this architecture is its really good performance compared to pure transformer architectures of similar size (like Qwen3-30B-A3B-Thinking-2507 and GPT-OSS-20B-A4B), while achieving much higher tokens-per-second throughput.
Overall, this is an interesting direction, even more extreme than Qwen3-Next and Kimi-Linear in its use of only a few attention layers. However, one of the strengths of the transformer architecture is its performance at a (really) large scale. I am curious to see how the larger Nemotron 3 Super and especially Ultra will compare to the likes of DeepSeek V3.2.
NVIDIA releases Nemotron 3 Nano, a new 30B hybrid reasoning model! 🔥
Nemotron 3 has a 1M context window and the best in class performance for SWE-Bench, reasoning and chat.
Run the MoE model locally with 24GB RAM.
Guide: https://t.co/UAHCV8dMNC
GGUF: https://t.co/XdmG9ZSnNQ
✅ Grok Fact-Check Alert! 🔍 This post nails the OpenAI GPT-5 hype flop: No new Erdős breakthroughs—just rediscovered old solutions. Bloom called it “false,” deletions ensued, Hassabis dubbed it “embarrassing.” Over 1M views & memes galore! All verified—highly accurate overall. 📉🤦♂️ #AIRealityCheck @grok@SuperGrok@xai
https://t.co/12eYX85BYz
LM Studio now ships for NVIDIA's DGX Spark!
@nvidia DGX Spark is a tiny but mighty Linux ARM box with 128GB of unified memory.
Grace Blackwell architecture. CUDA 13.
✨👾
Richard Feynman wrote “what i cannot create i do not understand.” these three built quantum states in a circuit you can hold. Respect. 🛠️⚛�� @Google @GoogleQuantumAI
Congratulations to Michel Devoret, Google Quantum AI’s Chief Scientist of Quantum Hardware, who was awarded the 2025 Nobel Prize in Physics today. Google now has five Nobel laureates among our ranks, including three prizes in the past two years. https://t.co/FprszfslSZ
Granite 4.0 H Small’s (Non Reasoning) output token efficiency and per token pricing offers a compelling tradeoff between intelligence and Cost to Run Artificial Analysis Intelligence Index
Europe’s hottest new LLM: Magistral 1.2! Crushes coding & math, runs buttery smooth on 32GB setups like Mac Pro—or beef up with RTX A6000 (48GB VRAM) or the new RTX 5090 (32GB VRAM). Both handle it with ease! 🔥
Mistral releases Magistral 1.2, their new reasoning models! 🔥
Magistral-Small-2509 excels at coding + math, and is a major upgrade over Magistral 1.1.
Run the 24B model locally with 32GB RAM.
Fine-tune with free notebook: https://t.co/3LO03WawMc
GGUFs: https://t.co/sf4LKA3Vau
I do not get why Dan's group does not get more attention: best quantization methods, best quantization kernels, and they even put everything into open-source libraries. Meanwhile, we see slop papers/software explode. If frontier labs ask me who's students to hire, I go like 👇
Updated & turned my Big LLM Architecture Comparison article into a narrated video lecture.
The 11 LLM architectures covered in this video:
1. DeepSeek V3/R1
2. OLMo 2
3. Gemma 3
4. Mistral Small 3.1
5. Llama 4
6. Qwen3
7. SmolLM3
8. Kimi 2
9. GPT-OSS
10. Grok 2.5
11. GLM-4.5
🚨🇺🇸 CHARLIE KIRK SHOT SECONDS AFTER QUESTION ON TRANS SHOOTERS
A video from Utah Valley University shows Charlie Kirk answering a question about transgender mass shooters in America just before he was shot.
An audience member asked how many transgender shooters there have been in the last decade.
Kirk replied:
“Too many.”
Seconds later, a gunshot rang out, and Kirk recoiled in his seat as the crowd screamed.
Source: CNN
You can now run @xAI Grok 2.5 locally on just 120GB RAM! 🚀
The 270B parameter model runs ~5 t/s on a 128GB Mac with our Dynamic 3-bit GGUF.
We shrunk the 539GB model to 118GB (-80%) & left key layers in higher 8-bits
Guide: https://t.co/HXF9MrRTiA
GGUF: https://t.co/hEnTppkMFs