Excited to launch Gemma 4: the best open models in the world for their respective sizes. Available in 4 sizes that can be fine-tuned for your specific task: 31B dense for great raw performance, 26B MoE for low latency, and effective 2B & 4B for edge device use - happy building!
Are small models still undertrained?
We are releasing a 2B model that beats GPT-3.5. The crazy part is that it was distill on only 2T tokens from a small model.
Distillation is the future of LLMs with the growing availability of large and efficient open models!
Gemma 2 is available to researchers & developers. At 27B it delivers best-in-class performance for its size and is competitive even to models over twice its size! Proud to continue our tradition of thoughtfully bringing cutting-edge research to the open models ecosystem.
Gemma 2 is out!
As with our first model, we're super focused on creating models at useful, practical sizes, so that they can be easily deployable... all the while being amazing in quality.
We upgraded our 9B so that it's truly awesome and best in class across many benchmarks. And we're introducing a brand new 27B, also best at size, and actually stronger than some larger models.
Both did real nice on LMSYS.
The 27B Gemma 2 model is designed to run inference efficiently at full precision on a single Google Cloud TPU host, NVIDIA A100 80GB Tensor Core GPU, or NVIDIA H100 Tensor Core GPU.
And of course, this is our open weights model line... enjoy!
https://t.co/TmgaJH52Zi - try it in AI Studio
https://t.co/ypeIKONwSC
More in the tech report =>
https://t.co/2wnb6dIRWH
We’re introducing new additions to Gemma: our family of open models built with the same technology as Gemini.
🔘 PaliGemma: a powerful open vision-language model
🔘 Gemma 2: coming soon in various sizes, including 27 billion parameters
→ https://t.co/KYNJ1O4xfL #GoogleIO
Gemma is expanding.... we just announced CodeGemma, a version of Gemma tuned for code generation. And bonus... Gemma is now bumped to v1.1, addressing lots of feedback we got.
Congrats Gemma team for one more amazing release!
https://t.co/brZgeFtCic
Building Gemma together with an exceptional team has been a delight, and now we're thrilled to share it with the world. A huge congrats to the entire team!
Special thanks to Kathleen & Alek, @triswarkentin, @armandjoulin, @clmt – you are all amazing :)
We have a long history of supporting responsible open source & science, which can drive rapid research progress, so we’re proud to release Gemma: a set of lightweight open models, best-in-class for their size, inspired by the same tech used for Gemini https://t.co/0aVehuXila
Bard is now available in the US and UK, w/more countries to come. It’s great to see early @GoogleAI work reflected in it—advances in sequence learning, large neural nets, Transformers, responsible AI techniques, dialog systems & more.
You can try it at https://t.co/m9D7JYTHvU
Common HTML understanding tasks can be done without custom NN architecture design and with orders of magnitude less data by fine-tuning LLMs. Bidirectional attention appears to be crucial, and context windows remain the bottleneck.
Exciting news: #Parti and #Imagen teamed up to create a hybrid system with Parti creating 256x256 images which then recieve Imagen super resolution to produce 1024x1024 pixels! See the diagram below for how it works.
See thread for more info and new images with this system!
What happens when you combine the best of language models with robots that operate in the real world? Take a look at our new work from @GoogleAI and Everyday Robots!
After 2 years of work by 442 contributors across 132 institutions, I am thrilled to announce that the https://t.co/wezEGzDEHt paper is now live: https://t.co/4Yg36EB9Ru. BIG-bench consists of 204 diverse tasks to measure and extrapolate the capabilities of large language models.
Tool Augmented Language Models. abs: https://t.co/yrK3R1IAVN
Smaller tool augmented models outperform larger non-augmented models in two domains (thus far), and on out-of-distribution examples.
Great collaboration w/@AlwaysParisi and @YaoZhaoAI!
Am so proud of the team’s exceptional research & engineering over the past year+! We are excited to share PaLM 🌴 with the world! The paper is at: https://t.co/G4xOp2kgDL
For a year, the T5 team has collab'd with FLAX and JAX to build a successor to our research library, using it to train models at many scales 📈 on TPU...
...and now you can too!
T5X is still in rapid development, but you can use it or find inspiration at https://t.co/xLHeni3jfB!
My first @GoogleAI residency project was accepted to @emnlpmeeting#EMNLP2021!
Prompt Tuning can condition a frozen T5 XXL model to perform new tasks while only adding 0.003% more parameters and no performance loss.
Camera Ready 📸: https://t.co/UMSpXpidmH
Quick Thread 🧵(1/7)
CALL FOR TASKS CAPTURING LIMITATIONS OF LARGE LANGUAGE MODELS
We are soliciting contributions of tasks to a *collaborative* benchmark designed to measure and extrapolate the capabilities and limitations of large language models. Submit tasks at https://t.co/eJJXFtqPpi
#BIGbench