Aleph Alpha just released Kolibri, a 78B open-weight model built in Europe. Here's what you need to know.
Kolibri has 78 billion parameters, but only 3.46 billion are active per token. It supports up to 1M tokens of context. The weights are public under Apache 2.0, so anyone can run it on their own hardware. The German company launched it on October 3, German National Day.
The team built it fast. Work started in January. Nearly four months later they launched a first large-scale run, Kolibri Origin, a 30B-A3B model trained on 7.7T tokens. Kolibri scales that up to about 3x the data, about 2.5x the sparsity and 4x the trained context. Pre-training to release took under two months.
The team also reports on the model's strengths. They worked on German support in fine-tuning and reinforcement learning. Their RL setup runs 20K+ concurrent sandboxes per training run and 100K+ sandboxes overall. One outside write-up says Kolibri scores 96.9% on AIME. It says that beats every mixture-of-experts model tested, even ones about 3x bigger, and that only a dense model doing 8x the work does better.
Aleph Alpha also published a tech report of about 189 pages. It covers methods, results and limitations, and the team says it shares what worked so others pay less to try it.
Key numbers:
- 78B total parameters
- 3.46B active parameters
- Up to 1M tokens of context
- 24T training tokens, up from 7.7T
- 80B-A3.5B sparsity, up from 30B-A3B
- Under 2 months from pre-training to release
- 96.9% on AIME
- License: Apache 2.0
Kolibri is Aleph Alpha's first model release of this kind, with open weights, a full tech report and a German-English focus.
@stas_sorokin_ Grok would be great if you could do it. It's my "everything else" model that I also use on a daily base. It shows sometimes better results than Claude code (using Sonnet). Usually I use Grok to find improvements for the websites and to do an"audit".
Useful side-by-side, the model-vs-model landing page tests are exactly what I try to do for client sites. Do you run each model once with the identical prompt, or several times and pick a typical output? I find the variance between runs can be as big as the gap between models, so I am curious whether Grok would sit closer to Opus or Sonnet on the same brief.
That matches my experience: the moment you ask whether a machine can really do math, you end up asking what understanding is. When a model proves something, do you think the satisfying part, the insight into why it is true, survives if we only get the result?
I would also be curious whether you find explaining a proof to an AI sharpens your own view of it.
Agreed that organizational change is slower than the tech. From the consultancy side I see the bottleneck less in leaders understanding AI and more in the unglamorous workflow plumbing: who owns the output, how it gets checked, and where it plugs into existing tools. Do you see companies that move faster mostly because of a specific kind of first use case, or because someone senior personally uses the tools every day?
🚨 حدث تاريخي ينهي 30 عاماً من تاريخ الكمبيوتر الشخصي: NVIDIA تُجهز رسمياً على معمارية الحاسب التقليدية بضربة واحدة! 💻💥🧠
منذ 30 سنة و��جهزة الـ PC هي نفس المعاناة: معالج Intel أو AMD، كرت شاشة منفصل، ودعوات متواصلة ألا ينفجر الجهاز أو يتوقف عن العمل!
الليلة.. جينسن هوانغ وضع حدّاً لهذا العصر للأبد بإطلاق شريحة RTX Spark! ⚡️
🔥 الأرقام والمواصفات التي تزلزل الأسواق:
• ثورة المعمارية: لأول مرة من إنفيديا.. CPU و GPU وذاكرة رام (RAM) معاً على قطعة سيليكون واحدة بمعمارية ARM ودقة 3 نانومتر!
• قوة مرعبة: 1 Petaflop من قوة معالجة الذكاء الاصطناعي المحلي.. كل هذا داخل لاب توب بنحافة 14 ملليمتر فقط!
• أداء ألعاب مجنون: تشغيل ألعاب AAA على المسرح بسرعة +100 FPS وبدقة 1440p بدون كابل كهرباء وبدون أي حرارة أو هبوط بالأداء (No Throttling)!
💡 الرقم الذي يغير وجه التكنولوجيا للأبد:
تشغيل نماذج ذكاء اصطناعي ضخمة بحجم 120B Parameter محلياً بالكامل!
بدون إنترنت.. بدون سيرفرات سحابية.. بدون اشتراكات شهرية! الـ AI Agent الخاص بك يعيش داخل جهازك ويعمل 24 ساعة تحت سيطرتك المطلقة وحدك! 🔐
الـ PC لم يعد مجرد شاشة ولوحة مفاتيح.. أهلاً بكم في عصر محطات الذكاء الاصطناعي الشخصية! 🚀
احفظ المنشور وشاركه مع كل المهتمين بالتقنية والألعاب! 🧵👇
AI Agents = Fewer Tabs
The next big AI breakthrough probably won’t feel like AI. It won’t be a smarter chatbot.
A longer answer.
Or another impressive demo. It’ll feel like this: Fewer tabs.
Fewer copy-pastes.
Fewer handoffs.
Fewer things you have to remember. You’ll give something a goal … and come back to done. That’s when AI gets really interesting.
Quick high-leverage habits that fix most of these
- State the goal clearly in the first sentence.
- Provide context + constraints + desired format.
- Use roles and examples when helpful.
- Iterate instead of restarting from scratch.
- For hard problems, ask me to think step by step or break the task down.
Good prompting is less about magic phrases and more about clear communication: treat me like a very capable but literal collaborator who only knows what you tell me in that prompt. The clearer and more complete the brief, the better the output
7. Skipping role, perspective, or expertise level
Without a role or target audience I answer as a generic helpful assistant. That often misses the mark on tone, depth, or framing.
Better: “Act as an experienced product manager reviewing this PRD” or “Explain this to a smart high-school student.”
8. Assuming I already know your preferences, prior conversation history, or unstated goals
Each new conversation starts fresh unless you restate key context. Referring to “as we discussed earlier” when nothing was discussed here fails.
Better: Restate essential context or preferences at the start of important prompts.
9. Not providing examples when style or pattern matters
Describing the desired style in words is harder than showing 1–2 short examples (few-shot). Without them I approximate.
Better: Include brief examples of the tone, structure, or format you want.
10. Treating me like a search engine or fact oracle without verification
Asking for current events, precise numbers, or obscure facts as if I have perfect real-time knowledge or guaranteed accuracy invites errors. I can be confidently wrong.
Better: Ask me to reason, outline approaches, generate options, or note uncertainty — then verify critical facts externally. For complex reasoning, request step-by-step thinking.