aleph hired many "kids" early in their careers who went on to be highly productive elsewhere.
connor, niklas, marco, hannah, jan, johannes, christian, nikolas, kilian, felix, szymon, vedantโฆ
the massive difference between SF and (especially) Germany is the willingness to pay exceptional young people high salaries.
one team (finally) delivered on some of the promises made years ago. the other looks close to shipping a model that competes at a larger scale. today is not the day to discuss aleph vs. mistral.
rather about how to deploy more capital the right way and move forward.
Sorry, Seb, usually lots of love for your work, but how is Europe now having n=2 players in the LLM game a generational fumble for Mistral? Can we please just celebrate having more teams who try and build models, rather than immediately using this as a reason to hate on either? Europeโs market is big enough for several sovereign players (btw, the world is also a pretty nice TAM) and different models (from my understanding Mistral is working on a much bigger model).
Big congrats to my friends at Aleph. Glad I was still around when this team formed and got to see how much work went into building it end to end.
https://t.co/VVpZLOM4rK
We will offer inference for free for the next ~48 hours.
Hit me up if you are interested in bringing this into production or post-training it.
Big congrats to my friends at Aleph. Glad I was still around when this team formed and got to see how much work went into building it end to end.
https://t.co/VVpZLOM4rK
We will offer inference for free for the next ~48 hours.
Hit me up if you are interested in bringing this into production or post-training it.
Small bird, fast wings, Kolibri is here.
78B parameters. 3.46B active. Up to 1M tokens of context. Built in Europe.
Now the weights are yours. Run it on your own hardware, under Apache 2.0.
We're hosting a Continual Learning Hackathon in Berlin on August 1st.
Together with @raphaelmosaic are bringing together ~20 of Europe's strongest ML researchers and engineers to work on models that keep learning during deployment.
Topics: synthetic RL environments, on-policy distillation, KV-cache compaction โฆ
We will provide dataset and compute together with our friends at Lyceum. Looking forward to share more on what we are up to with everyone there.
You can find the application link in the comments to this post.
We will organise travel stipends for strong student candidates outside of Germany that wouldnโt be able to make it otherwise. If you are working on similar topics reach out directly via PM for those.
cc: @blwiertz
@ivanburazin Building the continuous learning platform for company wide assistants. Just got our MVP done and feeling the pain of maintaining self-hosted Firecracker.
For #ICLR2025 we are unveiling a new, high-quality pretraining dataset for German LLMs. Shared to strengthen the open research community. Shaped by our belief in excellence and transparency.
https://t.co/DxI4kVaosA
Excited for ICLR 2025 in Singapore? Join our BoF Social (24 Apr, 12:30 p.m., Opal 103-104) on tokenizer-free, end-to-end architectures. Ready for insightful discussions and networking? Sign up here https://t.co/x7szBWOCg2
#ICLR2025#AIResearch#EnterpriseAI#Tokenizers