$40,000 worth of GPUs
π§π΅πΆπ πΆπ π π’π₯ππ‘π: the largest African language model ever trained from scratch, at 1.5 billion parameters. Not a fine tune, this was trained from the ground up for Africa. π
Trained on 12 African languages, it beats models from Google, Meta, Alibaba and HuggingFace on African language modelling and translation while being up to 8x smaller.
π π’π₯ππ‘π doesn't belong to us. The weights are open and free to use commercially, so researchers can build on a base that doesn't treat African languages as an afterthought.
Weights in the comments.
Thanks to @UNDP, @AIHub4SD and CINECA for the GPUs and the support. None of this happens without them.
We benchmarked DeepSeek V4.1 Flash by @deepseek_ai .
It reached 98% of GPT-6 Astraβs score at 1.4% of the cost on everyday design tasks based on user requests.
Every model except Astra scored lower AND cost more.
Are open models overtaking closed ones?
Full results below βοΈ