@oliviakoshy came back to ML@B speak about her work at Hex (@hexanalytics): benchmarking agents for data analytics, why these evals are so challenging and how their benchmark, DataBench, attempts to fix some of these problems!
We're excited to announce the first Berkeley BioML seminar of the semester happening next Monday 10/5! Join us for a talk by Danny Reidenbach from NVIDIA on codesigning protein sequence and structure with test-time search!!
https://t.co/jvo9UbxKnE
There’s lots of buzz around agentic harnesses for robots, particularly for long horizon tasks that require complex reasoning and memory. But what will it really take to turn a reasoning agent into a reactive, reliable and low-latency robot policy?
In our new paper, workspace models, we design a new memory architecture that acts as a latent harness for stronger reasoning models. Here’s why we think this might be the way forward:
(1/9)
Couple weeks ago, @mtvnastya and I gave a talk at UC Berkeley's ML Club @BerkeleyML about decentralized AI. We discussed access to AI compute, verification of work, how Gonka approaches these problems, and open research questions students can contribute to.
https://t.co/jTXMynnx1H
Hardware yearns for block sparse attention, yet it seems largely absent from open weight LLMs. DeepSeek developed NSA, and people speculated DeepSeek v4 would integrate it, yet it was never utilized.
We have a hypothesis as to why.
We found that replacing dense attention with NSA significantly degraded its ability on synthetic retrieval tasks, even when finetuned on them. On our 32k context benchmark, it scored 0.300 compared to dense attention’s 0.904.
We found the reason, and how to fix most of it. Introducing COBS🌽(Cumulant Order Block Sparse Attention), with @AdiGhai18@sanjitneelam@ZVasania@tensorpro:
• Raises NSA’s 0.300 → 0.820, closing ~86% of the gap to dense
• 15.15x less KV-cache read traffic than dense (just 1.21x the NSA baseline)
• Lower position-wise NLL than dense in our comparison
The key insight: block selection is the keystone to block sparse attention, and existing methods are mathematically stuck in storing a first-order approximation of a cumulant generating function. COBS caches a compressed second cumulant and escapes this ceiling.
Details in the paper.
Deciding when to jump in and help someone—and when to hold back and let them work through it—is something humans navigate constantly. How do AI assistants handle this tradeoff?
We introduce Int-Bench, a framework for evaluating interventions during problem-solving tasks.
This week, I wanted to see if we could get the smallest possible Mixture of Experts going that takes a few hours to pretrain on a single GPU, is fast for experiments, but also somehow performs well compared to larger MoEs. For reference, we have Nanochat for transformer experimentation that's 561M params. But Qwen's smallest MoE is still 14 billion total params which is still +++ GPUs. It's been very fun to work on nanochat version of that 🐸⚡️
nanoMoE is a 500M total param MoE that you can train for <5 hours on a single H100 using 3.5-4 less OOMs than the smallest MoEs. And packs a punch for its size (reaches 87% accuracy of OLMoE-1B-7B.). Also tried out the new Quantile Balancing from kimi so no hyperparam sweeping.
More numbers on performance here:
https://t.co/5IE6Gxf1dk
This is Google’s new diffusion LLM, DiffusionGemma’s denoising canvas over time.
Diffusion LLMs can generate tokens in flexible order. But in practice, do they just become autoregressive anyway?
1/ 🧵
What if the best visual reasoning steps are ones humans can’t specify? 🤔
Existing VLM reasoning is often constrained by language, pixels, and human-designed intermediates.
We introduce Latent Implicit Visual Reasoning, where we show that VLMs can discover the best visual reasoning steps by themselves — no bboxes, no intermediate images, no extra supervision.
Presenting this week at CVPR!
(1/n)🧵
our members ran a great reading group on major architecture changes in deepseek v4!!
they covered: hyper connections + manifold hyper connections, KV cache + MQA/GQA intro, Deepseek Sparse attention (prerequisite to understanding the new CSA) & a walk through of CSA and HCA
We'll be closing out this semester's Berkeley BioML Seminar on 4/28 with a talk from @antoinekoehl on PEINT, a powerful deep learning framework for both phylogenetic inference and protein engineering. Sign up below! https://t.co/QM7WXMZTu4
super proud of members Avy Harish (@AvyukthH60737), Sahir Tandon and Chris John for hosting our first physics informed ml seminar!
they covered: physics-informed nns, neural odes, sparse identification of nonlinear dynamics, fourier neural operators & other applications
We're excited to announce the fourth Berkeley BioML seminar of the semester happening next Tuesday 4/7! Join us for a talk by Shreshth (@shreshth_gandhi) from @tahoe_ai on Scaling Perturbation-Trained Single-Cell Foundation Models!
https://t.co/QfYDqvYh1N
Wrote a deep dive on implementing a language model from scratch in JAX and scaling it with distributed training!
If you’re coming from PyTorch and want to see how the same ideas look in JAX, or just want a hands-on intro to distributed training, check out this blog post: https://t.co/nsR3O3Zjxg
Comes with code + an assignment and test cases so you can follow along!
@threebarebears hi daniel! we're berkeley's oldest machine learning club and were wondering if you'd be down to talk to us about the process of making hoppers (: