2ร the number of models trained โ 2.02ร the parameters in one model.
We fitted a new scaling law with the number of models trained on the x-axis. The exponent came out uncannily close to Chinchillaโs: 0.345 vs. 0.34.
As the population grows, emergent specialization appears and the overall capability of the system improves.
We will provide limited access to the models via API in the upcoming weeks, stay tuned for more.
TLDR:
- Doubling the number of models trained predicts 21.3% lower evaluation loss in our fitted scaling trend. Chinchilla predicts a comparable 21.0% reduction in reducible language-model loss from 2ร the parameters and 2.32ร the training data. Each model in our population stays the same size.
- At 8 models, selective training used 84% less forward/backward compute, scored 10.7 points higher while saves 87.1% in required memory than training 8 models in isolation.
- At 16 models, the best single model scored 56.0% versus 86.2% for the best-per-task population, a 30.2-point specialization gain
- Emergent skills across benchmarks shaped >80% of specialization, while benchmark identity explained only 1โ2%
an AI model's blind spot can be another model's strength.
We trained 16 specialized models as a population, so an answer missed by one can be supplied by another. Just like a human experts team.
Model training today means squeezing as much capability as possible into one single model.
But human intelligence scales in numbers, not just the size of our brain. The smartest entity on earth has never been a single person. It's the civilization.
At Banbury Road, instead of a single model, we're training many models end-to-end, allowing them to self-organize, collaborate, and specialize into emergent roles. Together, even smaller, open-source models can achieve frontier performances after training at a fraction of the cost.
Superintelligence is not monotheistic.