@pfau Can you please ask for a python program (say "A") that takes another program (say "B") as input and:
- if B stops, A should run forever;
- if B doesn't stop, A outputs "Does not stop".
Need it for a friend.
Their newly released pretraining report reveals that Soofi S 30B-A3B is based on NVIDIA’s open Nemotron 3 Nano reference architecture, the same hybrid Mamba + Transformer MoE design with ~3B active parameters and nearly identical architectural choices
What’s actually new isn’t the architecture, but the training:
-~27 trillion training tokens
-German deliberately up-weighted
-End-to-end training on Deutsche Telekom’s Industrial AI Cloud
-Full training recipe, hyperparameters and evaluation methodology released openly
Ngl, excited to see europe kinda trains their own models.
📄 Releasing the Soofi S pretraining tech report: a sovereign, open foundation model for German and English
Today we’re publishing the full pretraining tech report and project page for Soofi S 30B-A3B — a Mixture-of-Experts hybrid Mamba model trained on ~27 trillion tokens with deliberately up-weighted German.
What’s in the report:
🏆 Strongest fully open model in our evaluations on BOTH the English and German aggregates — ahead of Olmo 3 32B and Apertus 70B (full methodology in the report)
📋 Radical transparency: complete per-source data accounting, all hyperparameters, training + eval code, checkpoints — everything under permissive licenses
🇩🇪 Trained end-to-end on Deutsche Telekom’s Industrial AI Cloud in Munich — sovereign AI infrastructure on German soil
Soofi S combines frontier-level capability with the highest measured aggregate long-context decode TPS, and unlike full-attention dense baselines maintains high throughput as context grows. The figure plots Capability Index versus measured aggregate decode TPS/GPU at 40K context and batch 32. The Capability Index averages five benchmark groups, i.e., Code, GSM8K, GPQA-Diamond, English aggregate, and German aggregate, after normalizing each group to the best plotted model. Aggregate decode TPS/GPU is measured with a TP=1, one-B200 vLLM latency-subtraction protocol.
Really excited about the testing workshop tomorrow! Feel free to stop by my poster "Membership Circuits" to see if we can obtain formal guarantees with learned models, happening at 15:30 🤗
ICML main conference is coming to an end, but the discussions are not over yet.
Join us at the ICML 2026 Workshop on Hypothesis Testing.
📅 July 11, 2026
📍 Room 318, COEX, Seoul, South Korea
Website: https://t.co/ifMZJU6AAs
Accepted papers: https://t.co/eli8iyRVLh
Really excited about the testing workshop tomorrow! Feel free to stop by my poster "Membership Circuits" to see if we can obtain formal guarantees with learned models, happening at 15:30 🤗
ICML main conference is coming to an end, but the discussions are not over yet.
Join us at the ICML 2026 Workshop on Hypothesis Testing.
📅 July 11, 2026
📍 Room 318, COEX, Seoul, South Korea
Website: https://t.co/ifMZJU6AAs
Accepted papers: https://t.co/eli8iyRVLh
Excited that our ProbML symposium has now started (and is online on zoom: link is on https://t.co/861JfFwERF)! We have a full day of talks and posters on probabilistic machine learning.
Excited to share that HaloProbe is accepted at ICML @icmlconf ! 🎉
Bayesian detection + mitigation of hallucinations in LVLMs (with a Simpson’s paradox twist).
arXiv: https://t.co/B0Z9OikDcA
See you in Korea 🇰🇷
#ICML#ICML2026#LVLM_Hallucination
Diffusion models fail at multi-object generation — but why? 🤔
In our #ICML2026 paper, we built MOSAIC, a controlled framework to diagnose these failures.
Spoiler: it's not mainly data imbalance. Scene complexity and missing compositions in training matter much more! ✨
(1/n)
Just arrived in Seoul to join ProbML and ICML. Looking forward to many interesting talks and discussions 😊
(I will also present a poster at the Hypothesis Testing in Machine Learning workshop on saturday, details coming!)
Our position paper “Modular Memory is the Key to Continual Learning Agents” (https://t.co/dv77KFe7ll) has been accepted to #ICML2026@icmlconf as a spotlight! 🎊
Read for a modern perspective on memory, continual learning, and sustainable adaptation at foundation model scale!
MoE routing shouldn’t be a hard top-k guess.
ProbMoE turns TopK routing into probabilistic inference over expert subsets, giving routers more informative gradients, more exploration, and better expert utilization. Because it’s probabilistic, it naturally extends to Dynamic-k, learning both which experts to use and how many each token needs.
Moreover, ProbMoE is broadly applicable to TopK-routed MoEs, while keeping training efficient through sparse expert execution.
Sparse compute. Smarter routing. 🚀
So wichtig 👇Wir brauchen einen positiven Blick auf und einen aktiven Zugang zu #KI Modellen in 🇪🇺. Schon lange warnen viele KI-Wissenschaftler vor ungewollten Abhängigkeiten. Glaubt uns endlich. Bitte ! @HolgerHoos
Excited to share KletterMix 🇩🇪🚀
A ~725B-token German pretraining + annealing corpus.
Proud to have co-led this with @HarleRuben, Sebastian Sztwiertnia, Abbas Goher Khan, Mehdi Ali, @effi288, and @kerstingAIML.
Paper: https://t.co/dzg3YQgUyV
I'm in Denver at #CVPR2026 this week! 🏔️
Together with @WolfStammer, I will be presenting our recent work, Vision-Language Programs, at Poster Session 3 on Saturday ⏰ 11:45 am - 1:45 pm.
Very excited to share our work and discuss the ideas behind it!
Last week, we hosted the Rhine-Main Universities AI & Creativity Symposium in Darmstadt, and the atmosphere was fantastic! 🧠💻
We gathered researchers from many disciplines to explore the big question: What does it truly mean to be creative, for both humans and machines?