New AI model alert! MiniMax-H3-X2-Detail-VAE is here to supercharge your image-to-video game. It's a specialized VAE for ComfyUI that adds insane detail and x2 upscaling. Perfect for creators who want crisp, high-res videos from a single reference image.
Turn DiffusionGemma into a Jev-like model.
This brings a new level of precision to structured tasks. By seeding a canvas with a response template, vLLM lets you extract confidence scores and probability distributions for yes/no, multiple-choice, and scored questions, all in a single denoising step.
Fast, efficient, and ready to scale!
NVIDIA and Stanford just challenged Jev.
(their new System 1 architecture runs up to 9x faster.)
It is called a Contrastive Language Model, or CLM.
Like Jev, CLM is not designed to generate text. It handles the small, repeated decisions inside AI systems, such as choosing a tool, ranking a patch, routing a request, or selecting the next action.
But CLM reaches those decisions differently.
Instead of generating an answer token by token, it treats decision-making as a retrieval problem.
Here is how it works.
1) Encode the state
CLM takes the current situation, such as an agent’s context or the state of a game, and converts it into a vector.
It uses a frozen Qwen3-8B model with a small trainable state projection head.
2) Encode every possible action
A separate action head converts each candidate into the same vector space.
In the Mario example, the candidates are left, jump, and right run. CLM does not invent a fourth option. It only evaluates the actions supplied by the application.
3) Learn which states and actions belong together
During training, the correct state-action pair is pulled closer while incorrect pairs are pushed apart.
A batch of B examples produces a B × B similarity matrix. The matching pairs sit on the diagonal. Every other pairing becomes a negative example.
This contrastive training uses InfoNCE, the same general mechanism behind systems such as CLIP and dense retrieval.
4) Turn similarity into a decision
At inference, CLM measures the cosine similarity between the state and every candidate action.
A softmax converts those scores into a probability distribution. The application can choose the winner, apply a confidence threshold, or escalate an uncertain result.
The real speed advantage comes from separating states and actions.
Actions can be embedded once and cached. If an agent repeatedly chooses between the same tools, CLM only needs to encode the changing state and compare it with stored action vectors.
That replaces repeated generation with one embedding pass and a set of cheap dot products.
The researchers report that CLM-8B matches Jev across computer-use, gaming, and tool-calling evaluations while reaching up to 9x lower latency. The improvement is largest when actions repeat or the candidate set grows.
CLM still has limits. It cannot generate new actions, its probabilities are relative to the supplied candidates, and its strongest verifier results require task-specific fine-tuning.
But its central idea is powerful.
The entire research is open-source, including the code.
Read more here: https://t.co/I9kPwMPI7B
When software already knows the possible answers, an AI model should score them instead of generating more words.
I also wrote a full breakdown on how system one models like Jev work.
The article is quoted below.
As an outcome of electron hydrodynamics, electrons flowing as a viscous fluid are expected to form vortices. Now, researchers have directly detected electron vortices for the first time using a nanomechanical resonator.
Read the Letter: https://t.co/JUXVM2RWEU
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics.
The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra.
The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
I am excited to share some of the progress we are making towards using Gemini to accelerate scientific discovery in the real-world.
We present an extension of Co-Scientist that transitions from pure in silico ideation to execution-grounded research partner, adapting to multiple levels of autonomy across different scientific domains.
One theorem every ML engineer should know:
The Johnson–Lindenstrauss Lemma.
It states that high-dimensional data can be projected into a much lower-dimensional space while approximately preserving pairwise distances.
Why it matters:
• Explains why random projections work
• Enables scalable learning in high dimensions
• Used in embeddings, compressed learning, and ANN search
• Helps fight the curse of dimensionality
The surprising part:
You can reduce dimensions dramatically without destroying the geometry of the data.
That’s why many ML systems can operate efficiently even with massive feature spaces.
Modern representation learning is deeply connected to this idea:
Good embeddings preserve structure while compressing information.
In ML, compression is often not loss of intelligence —
it’s removal of redundancy.
Last night 2nd stage of @SpaceX Falcon 9 carrying Starlink satellites flew over Finland. My cameras did not captured it but I got some faint northern lights and low altitude fast moving ice clouds. Shot from the top of Majakka residential tower in Kalasatama, Helsinki.
llm.c by hand ✍️ C meets Transformer
This combination is perhaps as low as we can get to explain how the Transformer works.
Special thanks to @Andrej Karpathy
--
Part 1: gpt2_forward
Karpathy's llm.c implements the transformer forward step as the following sequence.
(skip)layernorm_forward
1. matmul_forward
2. attention_forward
3. matmul_forward
(skip) residual_forward
(skip) layernorm_forward
4. matmul_forward
5. gelu_forward
6. matmul_forward
(skip)residual_forward
--
Part 2: matmul_forward
B: Batches (of Tokens)
T: Tokens
C: Input Channels
OC: Output Channels
--
C programming and matrix multiplication are two of the most important topics. But it is often difficult to get people excited by these topics.
I hope this exercise can help people see further into the LLM black box, and appreciate C programming and matrix multiplication by hand. 😀
---
100% original, made by hand ✍️
Join 56K+ readers of my newsletter: https://t.co/fFt8roc8D9
Don’t overthink it.
• Build a Calculator to master logic & loops
• Build a Weather App using live APIs
• Build a CRUD Web App with Flask + DB
• Build a Chatbot UI with Streamlit + GPT
• Build a File Organizer with os & shutil
• Build a Resume Parser using NLP
• Build a Stock Predictor using ML
• Build a Job Tracker that updates Notion
The best way to learn Python?
Projects. Not tutorials.
MVP4D: Multi-View Portrait Video Diffusion for Animatable 4D Avatars
TLDR: A multiview video diffusion model is trained conditioned on a reference image and FLAME mesh renderings. A four mode training conditioned on some viewpoints and/or some timesteps allows for arbitrary generations, this is used to distill a 4D Gaussian Avatar.
📽️ Project Page: https://t.co/Q7wHQs8yKd
📜 Paper: https://t.co/tV1zJ3KlOb
💻 Code coming soon
Excited to release new repo: nanochat!
(it's among the most unhinged I've written).
Unlike my earlier similar repo nanoGPT which only covered pretraining, nanochat is a minimal, from scratch, full-stack training/inference pipeline of a simple ChatGPT clone in a single, dependency-minimal codebase. You boot up a cloud GPU box, run a single script and in as little as 4 hours later you can talk to your own LLM in a ChatGPT-like web UI.
It weighs ~8,000 lines of imo quite clean code to:
- Train the tokenizer using a new Rust implementation
- Pretrain a Transformer LLM on FineWeb, evaluate CORE score across a number of metrics
- Midtrain on user-assistant conversations from SmolTalk, multiple choice questions, tool use.
- SFT, evaluate the chat model on world knowledge multiple choice (ARC-E/C, MMLU), math (GSM8K), code (HumanEval)
- RL the model optionally on GSM8K with "GRPO"
- Efficient inference the model in an Engine with KV cache, simple prefill/decode, tool use (Python interpreter in a lightweight sandbox), talk to it over CLI or ChatGPT-like WebUI.
- Write a single markdown report card, summarizing and gamifying the whole thing.
Even for as low as ~$100 in cost (~4 hours on an 8XH100 node), you can train a little ChatGPT clone that you can kind of talk to, and which can write stories/poems, answer simple questions. About ~12 hours surpasses GPT-2 CORE metric. As you further scale up towards ~$1000 (~41.6 hours of training), it quickly becomes a lot more coherent and can solve simple math/code problems and take multiple choice tests. E.g. a depth 30 model trained for 24 hours (this is about equal to FLOPs of GPT-3 Small 125M and 1/1000th of GPT-3) gets into 40s on MMLU and 70s on ARC-Easy, 20s on GSM8K, etc.
My goal is to get the full "strong baseline" stack into one cohesive, minimal, readable, hackable, maximally forkable repo. nanochat will be the capstone project of LLM101n (which is still being developed). I think it also has potential to grow into a research harness, or a benchmark, similar to nanoGPT before it. It is by no means finished, tuned or optimized (actually I think there's likely quite a bit of low-hanging fruit), but I think it's at a place where the overall skeleton is ok enough that it can go up on GitHub where all the parts of it can be improved.
Link to repo and a detailed walkthrough of the nanochat speedrun is in the reply.
@gabriberton It is a fork of VGGT-long, extended with COLMAP export and bundle adjustment from the VGGT implementation. To enable track prediction for long sequences, I replaced the DINO-ranked query frames with regularly sampled frames and introduced frame chunking in base_track_predictor.py
Applying transformer models like VGGT to 200–300 image benchmarks (NeRF 360, Tanks & Temples) for Gaussian Splatting is challenging due to VRAM limits.
My repo enables chunked VGGT-Long processing to bridge this gap: https://t.co/pBep7DzbLg 1/3
📸 COLMAP vs VGGT-Long example 👇