After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution
Today, we’re announcing Ternary Bonsai 2 27B.
Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance.
Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use.
Ternary Bonsai 2 27B is available today under Apache 2.0.
@asheposhtepa1 البته که این مدل ها قابل تحسین هستن چون دارن چیزی به ما تحویل میدن که ما شاید ۲ سال پیش از پرچمدارا می گرفتیم اما با حجم خیلی کمتر مدل مه خیلی مهم هست اما فعلا نمیشه گفت سلام لوکال، فعلا فقط باید از اینا برای گاردریل و تول کالینگ و خلاصه بگم به عنوان بازو های مدل اصلی استفاده کرد
@asheposhtepa1 شما همون بنچمارک ها رو هم اگر نگاه کنی متوجه میشی که مقایسه فقط با مدل های لوکال دیگه انجام شده و اگر همین مدل ها رو با پرچمدارا توی هر بنچمارکی مقایسه کنید متوجه منظور من میشید.
Introducing SubQ - a major breakthrough in LLM intelligence.
It is the first model built on a fully sub-quadratic sparse-attention architecture (SSA),
And the first frontier model with a 12 million token context window which is:
- 52x faster than FlashAttention at 1MM tokens
- Less than 5% the cost of Opus
Transformer-based LLMs waste compute by processing every possible relationship between words (standard attention).
Only a small fraction actually matter.
@subquadratic finds and focuses only on the ones that do.
That's nearly 1,000x less compute and a new way for LLMs to scale.
@TheAhmadOsman Are we also considering the price in this comparision!?
Most expensive Mac Studio brings 80+GPU Nodes to the table and it costs less than $7.5K. Can we say that it is an equal to 80 vRAM?
If yes the price worths it for real.