Om Namah Shivay 🔱
O
M
N
A
M
A
H
S
H
I
V
A
Y
Om Namah Shivay🪔
O
M
N
A
M
A
H
S
H
I
V
A
Y
Om Namah Shivay🙏
O
M
N
A
M
A
H
S
H
I
V
A
Y
Om Namah Shivay ⚜️
O
M
N
A
M
A
H
S
H
I
V
A
Y
@grok
Devon ke Dev Mahadev ki adbhut Darshan,,🔱🙏
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
♥️ I LOVE ALLAH♥️
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
♥️ I LOVE ALLAH♥️
#خاتم_النبیین_محمدﷺ#درود_وقرآن
Soofi Consortium Releases Soofi S 30B-A3B: An Open 31.6B Model for German and English Hitting 79.1 German Aggregate With Only 3.2B Active Parameters.
Here's how it works. 👇
1. Sparsity in two places at once
52 layers: 23 Mamba-2, 23 granular MoE, 6 Grouped-Query Attention. The MoE router picks 6 of 128 experts per token, plus 2 shared. Mamba-2 carries the sequence mixing with a fixed-size recurrent state, so 46 of 52 layers keep no KV cache at all.
→ 3.2B of 31.6B parameters active per token
2. Reference architecture on purpose
No bespoke backbone. It adopts NVIDIA's Nemotron 3 Nano design without modification — for day-one vLLM kernels, for serving efficiency, and for scientific control. That last one is the real move: Nemotron becomes an architecture-identical baseline, so the data recipe is the only variable left.
3. German as the deliberate variable
Three-phase Warmup–Stable–Decay curriculum. Phase 1 is breadth at a 1e-3 plateau, Phase 2 concentrates high-quality data as the LR decays, Phase 3 stretches context to 1M tokens.
→ ~26.68T consumed tokens
→ German 7.2% → 15.32% of the mixture, vs ~5% for all non-English in the Nemotron reference
→ +4.2 German aggregate, +1.8 English, +6.7 held-out English over Nemotron
4. Where the architecture pays: memory bandwidth
Every decoded token re-reads the weights and, for a Transformer, the attention cache of every sequence in the batch. Six KV layers instead of 52 keeps that per-sequence state small. Measured on one B200, TP=1, vLLM latency-subtraction.
→ 8–9× aggregate decode TPS/GPU vs dense 14–24B models at 40K context, batch 32
→ decode stays flat from 4K to 256K
5. The numbers (base model, lm-evaluation-harness, 16 open baselines)
→ 70.1 English aggregate, +2.8 over Olmo 3 32B
→ 79.1 German aggregate, +6.3 over Apertus 70B
→ 73.8 HumanEval, 84.2 MBPP-DE, 88.8 GLP-DE, 61.2 INCLUDE-DE
Full analysis: https://t.co/c2VlQO8jTH
Paper: https://t.co/GGOUHcVk1s
Technical details: https://t.co/XoIKZuSpQg
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
Allah
♥️ I LOVE ALLAH♥️