Wake up, baby! Germany has joined the race of frontier LLMs with its sovereign model "Kolibri".
And it's Open-Source, too!
Beats Qwen2635B, Nemotron3 Super, and Mistral Small 4 in the Math, Science, and Code arena! 🎉
The underlying architecture: custom MoE Transformer
Next stop for Sovrano: @COLM_conf 2026.
We’re heading to San Francisco to spend a week with researchers and teams working at the frontier of AI.
Find @SovranoAI at our booth, or message us to grab a coffee ;)
Most annotators asked to fix a model's answer rewrite the whole thing. It feels thorough. It is the slow way, and it quietly damages the data.
A rewritten answer is a human's answer. It carries the annotator's phrasing, their structure, their habits. Train on enough of those, and you are teaching the model to imitate your reviewers rather than to correct its own mistakes at the point where it makes them.
A recent paper tested the alternative. The annotator reads the model's response, finds the first token that is wrong, fixes just that, and lets the model continue from the corrected point. Repeat until the answer holds. Median annotation time dropped by half, and the final text stayed close to how the model actually writes, which is what you want if the goal is to correct behaviour rather than replace it.
The interesting part for anyone running expert programmes is what this asks of the expert. Locating the exact word where a legal answer goes wrong is a far more specialised skill than writing a good answer from scratch. It requires knowing the domain well enough to spot the turn, and knowing the model well enough to trust it with the rest.
The expensive judgment is where it broke. Everything after that is typing.
Hiii we are the people behind Sovrano 👋
Turns out getting the whole team in front of a camera at the same time is harder than it looks, but we made it happen!
Follow us & stick around!
What makes a good benchmark actually good?
At @AISummitBCN we put it to the test.
Participants built evaluation rubrics, tested them against model outputs and compared how human judgement changes the way we measure AI performance.
Good benchmarking starts with good evaluation!!
A lot of the data used to teach models what "better" means is wrong, and the reason is usually the form, not the person filling it in.
Preference data is built from pairs. Two answers, pick the stronger one. Training methods then treat every one of those picks as a clean fact. But anyone who has run a rating programme knows what actually happens at the screen. Half the pairs are obvious. A good share are genuinely tied. And a rater who is not allowed to say so will pick one anyway, because the interface demands it and the queue is long.
That forced pick is recorded as a confident preference. Multiply it across a dataset and the model is being trained, with full conviction, on coin flips.
New work on preference optimisation is starting to treat this seriously, catching pairs mid-training and relabelling them as clean, flipped or tied rather than throwing the noisy ones away. It outperforms simply filtering. Which is a technical way of saying something operational: the tie was always real, and pretending otherwise was the error.
If your rating interface has no honest way to express "these are equivalent", you are not collecting preferences. You are collecting the rater's coping strategy.
We have something very special to share today!
We developed a standard to assess benchmark quality, which we plan to apply to the most-used AI capability metrics going forward.
We hope this will raise the bar for designing and interpreting evaluations. Let us know what you think!
Everyone says AI trains itself now.
This week: OpenAI has hundreds of contractors reading real ChatGPT conversations, scoring every reply so the model learns what good sounds like.
That is not automation. That is judgment. It still does not scale without people.
https://t.co/u1Rmwu5NKI
649 papers in. 17 out.
This week: synthetic data finally gets a number, one example breaks a fairness benchmark, and uniform GRPO beats eight cleverer data policies.
The research worth reading →
https://t.co/NheqxvvaI3
Mistral just closed 3 billion euros. Biggest tech round in European history, 21 billion valuation.
Capital isn't Europe's problem anymore.
The expert judgment that actually trains a great model is. That layer gets built person by person, not in one check.
LLMs had the internet. Robotics doesn’t
Every hour of real-world robotics training data has to be created on purpose, with real people, environments and hardware
The race for physical AI is also a race for data ↓ https://t.co/1cJ7zU2Vul
Sovrano on tour 🇮🇹
First stop: Naples for ECML PKDD
AI, great conversations, a questionable amount of pizza and very little sleep... We’re just getting started 👀
Guess our next trip?
OpenAI confirms weeks of AI safety talks with Anthropic and Google DeepMind, as Trump's team dismisses safety concerns and pushes to keep pace with China.
https://t.co/ra2YGQhjt4
AI labs have spent years competing on models and compute
The next competitive advantage is human data
Europe has an enormous, largely untapped supply of real-world data and human expertise, that market will be worth billions
I wrote about why↓ https://t.co/vrSaLWc7bw
‼️ BREAKING: Revolut handed over customers’ passport copies, verification selfies and full transaction histories to a malicious actor.
The actor sent lawful government information-demand emails using a genuine government domain that passed domain authentication.
Revolut later concluded they were not authentic.
Affected customers were notified on Friday. What may have been disclosed ranges from name, date of birth and home address to account statements, withdrawal records and complete Bitcoin transaction history.
The company says it has alerted the agency to the unauthorised mailbox on its domain, blocked the address and begun notifying regulators.
It has not named the agency, explained how someone obtained a mailbox there, or given a number of affected customers.
ZachXBT, who circulated the notices, believes the incident was limited in size and aimed at high-net-worth users.
We ran OpenAI's newest model against the current EuroExec leader on European executive reasoning.
EuroExec is our blind, expert-graded benchmark for how frontier models handle real European executive decisions; the calls a C-suite makes under EU regulation, incomplete facts, and competing stakeholders.
Not trivia. Not code.
We took a sample of 47 expert-authored tasks, had GPT-5.6-sol and Anthropic, Fable 5 (our current leader) each answer them, and graded every response against the domain experts' checklists.
The result:
The two finished within ~1.4 points of each other (under 3%)
GPT-5.6-sol covered ~97% of what Fable 5 did against expert criteria.
A statistical dead heat, with Fable 5 holding a razor-thin edge.
Two takeaways for anyone building or buying frontier models:
1. The frontier is converging here.
2. There's still real headroom.
More from EuroExec soon.
Reward models re-read the entire reasoning trace to score it. Cost grows with length squared.
KV-PRM reads the KV cache built during generation instead. One verify token. Cost grows linearly.
Up to 5,000x fewer FLOPs. Same accuracy on MATH/GSM8K/AIME.
By Peng Kuang et al. → https://t.co/8oI4Hee6oU