NEWS: NVIDIA Groq 3 LPX is now in full production.
NVIDIA Vera Rubin NVL72 is the foundation of every AI factory. Paired with Groq 3 LPX, it unlocks faster, smarter agents and breakthrough user experiences.
Through extreme co-design across seven chips and five purpose-built racks, #NVIDIAVeraRubin is the most extensive AI factory platform.
Read the release ⬇️ https://t.co/jxk66QzW0S
Groq has entered into a non-exclusive licensing agreement with Nvidia for Groq’s inference technology.
GroqCloud will continue to operate without interruption.
Learn more here:
https://t.co/yg4TeBpuqa
OpenAI’s open models are live and already running on Groq. Try gpt-oss-20B and gpt-oss-120B today.
Groq delivers 128K context and built-in tools such as code execution and browser search. For the first time, developers and enterprises can deploy open models backed by OpenAI instantly, anywhere, at scale.
Start building now. Links in comments.
🚨 Bell Canada has selected Groq as its exclusive AI inference partner.
This is how sovereignty scales—and Groq is powering it. For countries that want to control their AI future, today is a turning point.
👇 More below
We built the region’s largest inference cluster in Saudi Arabia in 51 days and we just announced a $1.5B agreement for Groq to expand our advanced LPU-based AI inference infrastructure.
Build fast.
Groq is rivaling Google & Amazon in LangChain API usage?! 🤯
@DanielNewmanUV and @PatrickMoorhead caught @GroqInc CEO @JonathanRoss321 in Davos. Get their take on AI's future & why Jonathan believes Groq will DOMINATE inference.
Hint, it has to do with:
🚀 2 MILLION chips this year
🌎 Global network
⚡️ Energy efficiency
🤫 4nm chip coming?!
It's not every day an infra company goes viral, even in AI. Groq is having a moment and it's a case study in direct comms. Here are 4 things they're doing really well and 3 untapped opportunities.
This is where the team really nailed it:
1. MAKING IT VISCERAL
a) Groq coined the perfect new term with LPU (Language Processing Unit), they’re creating a new category around it, and they’ve made fetch happen through sheer repetition.
b) The below visual of their chip next to a GPU is brilliant, specifically because (1) obviously visuals hit in a way that words don’t, (2) images get more engagement and better algo boosts, and (3) while a list of specs might show the GPU winning in many ways, the image highlights Groq’s winning metric: purpose-built simplicity.
c) They use simple numbers (10X, 100X), which trades being precise for being memorable. They also use analogies to make complex ideas clear and easy to repeat (comparing a general purpose chip to an assembly line you have to set up from scratch every time). Much easier to recount over dinner.
2. GOING DIRECT
a) The CEO addresses criticism on Twitter quickly and proactively, while showing his credentials as a technical domain expert. The founder entering the chat to clear up technical details himself can be a KO move.
b) Their team is also super active on Hacker News, responding to comments on a very technical level and sharing details with relevant people (eg potential customers) — while stopping short of compromising trade secrets.
c) Lastly, Groq has been posting content for years now, gradually building up an audience to tap into when the time is right (now)
3. BENEFITS BEFORE BACKEND
a) Many founders get so excited about the technology that they lead with the specs. What people really care about is how it can make their lives easier. The Groq team understands this. Their entire homepage is basically a free demo. Everything else (who we are, all the marketing type stuff) is in the sidebar.
b) “Show, don’t tell” is powerful. The CEO understands this, telling the CNN interviewer that it’s better to show “the magic” before diving into how it works.
c) Jonathan Ross is a gifted communicator and has a way with words. When the interviewer mistakenly asked about Groq’s “models,” he smoothly said, “We don’t make the models, we just make them fast.” To show how engagement drops with lag, he said, “Imagine…if….I…spoke…this…slow…” to make the point. (I think it’s no coincidence he reads a lot of books about comms — see linked post).
4. PERFECT TIMING
a) While they’ve had tech coverage here and there for a while, they timed their mainstream press push perfectly, going on CNN after the biggest wave of chip news in probably the last decade. I often tell founders to concentrate their efforts on a lightning strike moment instead of a low-level drumbeat, and that's what the Groq team pulled off here.
And here are 3 opportunities to get even better (from the uninformed perspective of an outsider):
1. MESSAGING
a) Talk about Groq’s speed advantage not as a marginal difference on a scale (others are only somewhat fast while we’re very fast) but as a 0 to 1 problem: when natural language models actually feel natural, new products and experiences become possible that never were before. (In fact I’d suggest that as a tagline: “Making natural language natural”
b) Target this messaging specifically to potential customers (Groq’s speed isn’t a marginal improvement but a step change), and help developers understand what that speed ENABLES (new interactive experiences like in video games). Groq has undersold this so far and have a chance to own the whole concept of having a true natural language conversation.
c) The debate about pricing vs. NVIDIA is misguided. Even Citrini wrote, “You’d need about 600 of these to have the same performance as a 8x H100 box” — but Groq should emphasize that processing power doesn’t directly represent the speed experienced by the consumer actually running the models. You can’t have a baby in one month with nine women!
d) Similarly, some people will judge Groq on the quality of the output, so the team should keep reminding people that the point isn’t to compare Mixtral to Llama or whatever, the point is that each model runs faster on Groq than on other chips.
2. MORE TARGETING
a) Get more precise about their audience. Eg, sometimes they compare themselves to infra providers, sometimes to NVIDIA. It probably makes most sense to do the infra providers comparison when they’re targeting developers and do the NVIDIA comparison when they’re targeting investors and talent.
3. LEVEL UP WHAT’S WORKING
a) Lean more into LPUs as a new category. Eg, get someone to create a Wikipedia page for “Language Processing Unit” and cite Groq as the inventor of the first LPU and creator of the category.
b) Complement their excellent metrics graphs with even more visceral demonstrations of the product’s capabilities (eg with gifs or more analogies). They're already doing this a lot and it works so well.
c) Use X more effectively. Eg, update the pinned tweet to this one from @BrianRoemmele https://t.co/Hp5p14k5ii.
And consider retiring #groqspeed; hashtags don’t work, nobody else is going to use it, and it makes a post look spammy. I assume more than three hashtags in one post is likely to get treated as spam by the algo too; I think that’s what LinkedIn does.
It's awesome to see more and more founders building their own direct comms channels, and I hope Groq's success on this front is helpful for other founder-led companies in the future.
CEO & Founder, Jonathan Ross, explains how the #Groq#LPU ™ Inference Engine operates, as CNN's Becky Anderson converses with the incredible technology.
https://t.co/va2UwXAzgh
430 tokens/s on Mixtral 8x7B: @GroqInc sets new LLM throughput record
Groq has launched its new Mixtral 8x7B Instruct API, delivering record performance on its custom silicon. Pricing is competitive at $0.27 USD per 1M tokens, amongst the lowest prices on offer for Mixtral 8x7B.
Artificial Analysis has independently benchmarked these results and will continue monitoring performance. See our website for a full breakdown, including throughput, latency and pricing, with analysis of variance, performance over time and comparison to other API providers:
https://t.co/nMl9swpAL0
240 tokens/s achieved by @GroqInc's custom chips on Lama 2 Chat (70B)
Artificial Analysis has independently benchmarked Groq’s API and now showcases Groq’s latency, throughput & pricing on https://t.co/qhhMCIyXHF
This represents a milestone for the application of custom silicon to large language models and AI
Groq are serving a full quality FP16 version of Llama 2 Chat (70B) with the model’s full 4k context window
See full results here: https://t.co/yNuDh8hxl1
There's a first time for everything right? Our first public benchmark results are here. Not only are we a part of the LLMPerf Leaderboard by @anyscalecompute, we're LEADING with up to 18x faster #LLM#inference performance compared to top cloud-based providers: https://t.co/0oBM17J2tI
Get the full story on our LPU™ Inference Engine's performance in our latest blog and thanks to the work of our incredible team and the team over at Anyscale!
I have two suggestions for Elon:
1. Slartibartfast is a far more appropriate name for a snarky chatbot.
2. He should run the #LLM on #Groq ™ so he can provide sarcasm at speed.
Read my recent blog post to #grok why on both fronts: https://t.co/rB2bbsKCNA
#GroqOn