🐻 Democratizing LLM Pretraining
Gumini-1.5B Open-Source Release
What if you could match trillion-token LLMs
with just 3.14B training tokens?
Gumini is a Big signal for what's possible with tiny compute and data.
Report: https://t.co/MzgTfIVs2O
Model: https://t.co/5HRPAd6v4r
@zhengyiluo Hello, I am using Gear Sonic deployed on the G1. Is what is shown in the video the Teleoperation of the upcoming model?
The existing model does not bend that deeply. Also, what new squats will be added?
@SunnySanyal9@ziv_ravid@AlexGDimakis@sujaysanghavi Thank you, I truly appreciate this.
Your Inheritune paper was a huge inspiration, and it directly enabled me to build a strong model with limited compute and data.
Credit really goes to your work for showing what’s possible.
🐻 Quick update on Gumini-1B
A community-contributed GGUF version reached 1,300+ downloads in just 4 days.
Really appreciate the early interest this kind of signal matters a lot for independent research.
#OpenSource#LLM
🐻 Quick update on Gumini-1B
A community-contributed GGUF version reached 1,300+ downloads in just 4 days.
Really appreciate the early interest this kind of signal matters a lot for independent research.
#OpenSource#LLM
🐻 Gumini-1B is now open-source
A Korean–English bilingual pre-trained base LLM,
released a few days ago.
• 1.08B params, 10 layers
• Eval PPL: 14.95
• Inheritune-based design (lazy-layer aware)
https://t.co/4wdvM4AAtq
#LLM#OpenSource#Pretraining
Technical note:
Gumini-1B is inspired by the Inheritune methodology
(Sanyal et al., arXiv:2404.08634, CC BY 4.0).
The focus was on minimizing lazy layers
and preserving early attention representations,
rather than brute-force scaling.
https://t.co/4wdvM4AAtq
🐻 Gumini-1B is now open-source
A Korean–English bilingual pre-trained base LLM,
released a few days ago.
• 1.08B params, 10 layers
• Eval PPL: 14.95
• Inheritune-based design (lazy-layer aware)
https://t.co/4wdvM4AAtq
#LLM#OpenSource#Pretraining