We just dropped our research note on Eli (4B).
Trained on 30.2K curated code + CoT samples using custom T4 CUDA kernels for maximum reasoning density on budget compute.
Constraints force clarity.
Read the breakdown: https://t.co/cPTgW4FsYJ
New Research Note: On Eli Training & T4 CUDA Research.
Eli is our 4B model engineered for code intelligence & long-form reasoning—trained on 30.2K curated samples with custom fused CUDA kernels on T4.
Read the full note: https://t.co/u19OM5JFTD
Hello 👋 , after the release of playbox - we aren't stopping. Next up is the model named Gauss named after the prince of mathematics Carl Friedrich Gauss.
This model should be available in the coming months as open weights.
Read more here : https://t.co/IImx595CQZ
So we routed instead — kept the base model untouched, added a separate router for tool calls. Result: 50.3% GSM8K (beats Qwen2.5-3B's 49.0%, at 1/4 the params) and 75.5% on routed OOD arithmetic.
Full writeup + what we tried before this worked: https://t.co/qHYHz2moJT
Today we have @EpochAILab have posted our first model of many called Playbox 🧰.
Here is the break down (🧵thread )
We started with Qwen 3.5 0.8B , It already had great mathematical prowess but lacked the ability to actually be useful in math.
Shipped Playbox .
0.8B params. Built for math reasoning + routed tool use.
50.3% GSM8K
neck-and-neck with Qwen2.5-3B at 49.0%
75.5% on routed OOD arithmetic
Small model. Real benchmark. Close race. 📦