Excited to see Grok 4.5 ship. I haven't been here long, but the past few months have been quite a ride. The ramp-up wasn't easy and I hit some rough patches, which makes me all the more thankful for everyone who helped me along the way, also learned a ton from the talented people here. Glad I got to chip in across several training stages, and especially to work closely on a full post-training workflow for the first time. Now back to the grind and thrilled for what's coming next.
Announcing Grok 4.5, our first model trained specifically for coding and agents. It was trained with Cursor and offers frontier intelligence at leading speeds and cost efficiency.
https://t.co/i8HpU7w64k
Excited to share that our paper Rope to Nope and Back Again: A New Hybrid Attention Strategy (https://t.co/siCUzQnsLJ) has been accepted at NeurIPS 2025! Thanks to my amazing co-authors @bharatvenki, @DwaraknathG, John Lin, @davidcairuz, Phil_Blunsom & @acyr_l for their hardwork.
I'm excited to the tech report for our @Cohere@CohereForAI Command A and Command R7B models. We highlight our novel approach to model training including the use of self-refinement algorithms and model merging techniques at scale. Command A is an efficient, agent-optimised multilingual model offering best-in-class capabilities. Read more below! ⬇️
🚀 Big news @cohere's latest Command A now climbs to #13 on Arena!
Another organization joining the top-15 club - congrats to the Cohere team!
Highlights:
- open-weight model (111B)
- 256K context window
- $2.5/$10 input/output MTok
More analysis👇
🌐 Join @cohere at #GTC25 to dive deep into scaling models for longer contexts to advance #LLMs.
Explore key challenges, from data prep and model architecture to GPU optimizations and what's next in this cutting-edge #research session. ➡️ https://t.co/vzNjJjbGkH
It is also a larger instance of paper Rope to Nope and Back Again: A New Hybrid Attention Strategy: https://t.co/jV6qQh8bv8 with excellent long-context capabilities.
Stay tuned for our technical report!
Today @cohere is very excited to introduce Command A, our new model succeeding Command R+. Command A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, multilingual, and coding usecases. 🧵
📷 Excited to share our new paper: "Rope to Nope and Back Again: A New Hybrid Attention Strategy" where we propose a novel architecture that outperforms RoPE-NTK-based approaches with full attention span. (1/8)
Our findings align with recent work showing hybrid attention mechanisms often outperform full attention for long contexts. This opens up questions about the fundamental nature of attention mechanisms and how we improve model design accordingly. (7/8)
Canada is a leader in AI because of companies like @Cohere.
We are working with Cohere to build a cutting-edge AI data centre here at home — essential infrastructure for powering AI.
We’re excited to be included in the Canadian government’s critical work bolstering the country’s AI industry. Proud to partner with @JustinTrudeau, @cafreeland, and @FP_Champagne to secure AI infrastructure in Canada and further contribute to Canada’s global AI leadership. 🇨🇦