Huge news: Clarifai has agreed to license our AI inference & compute orchestration IP — plus the patent portfolio behind it – to @nebiusai.
Our core team is joining them to keep building.
Our technology becomes a key part of Nebius Token Factory, the inference platform inside their full-stack AI cloud. Faster inference. Bigger scale.
Can't wait to ride this rocket ship and build the foundation for the next decade of AI inference. 🚀
Nebius welcomes @Clarifai's core engineering and research team - led by founder and CEO @mattzeiler - to Nebius.
The company has agreed to license Clarifai’s inference and compute orchestration technology to strengthen Nebius Token Factory as a full-stack inference platform.
Read more: https://t.co/GDUudQ3i8m
Clarifai featured on Ramp’s Top Software Vendors for February 🚀
Ramp’s ranking is based on real spend and new customer adoption across thousands of companies.
At Clarifai, we’ve been focused on the foundations behind that adoption: compute orchestration that lets teams run models at scale across any infrastructure, pipelines for long-running, multi-step AI workflows, and the Clarifai Reasoning Engine built for agentic and reasoning workloads.
This is the infrastructure teams rely on when AI moves from experimentation into production.
More details in Ramp’s post: https://t.co/WacHZpBp8W
Clarifai now uses a single Pay-As-You-Go credit system for self-serve usage.
One balance across training, inference, workflows, and GPU-backed Compute Orchestration. No monthly minimums. Fewer feature gates.
Every verified user gets a one-time $5 credit to try the platform.
Read more here: https://t.co/e8z4Ilfxk2
A lot of teams are realizing that their biggest agent bottlenecks come from the execution layer.
Slow token throughput, fragmented runtimes, and inconsistent environments make agents unreliable even when the models are strong.
At NeurIPS today, our CEO and Founder Matthew Zeiler walking through how a unified compute layer removes these performance cliffs and keeps agents fast, stable, and cost-efficient across cloud, on-prem, and edge setups.
If you’re at NeurIPS, join us:
📅 Dec 4, 2025
⏰ 4:00 PM
📍 San Diego Convention Center, Exhibit Hall A,B
I joined Craig on Eye On AI Podcast to discuss why I believe compute orchestration is the missing layer in the current stack.
Give it a listen: https://t.co/cCtjpXgQea
Inference is becoming the biggest cost in the AI stack.
Our Founder & CEO @mattzeiler joined @craigss on EyeOn AI Podcast to share why inference can’t be an afterthought and how unified compute orchestration cuts GPU spend for complex agents and long-context models.
Check it out👉 https://t.co/qQdsT4O8xm
We’ve partnered with @arcee_ai to bring the Trinity family of powerful, end-to-end American-built open-weight models to the edge. 🚀
Trinity Mini and Nano are now available on Clarifai, proving that frontier models can be built with efficient capital and high performance.
- Trinity Nano (6B/1B active): Optimized for real-time, on-device and embedded use cases where efficiency is critical.
- Trinity Mini (26B/3B active): Tuned for multi-step agent workflows and function calling.
Why Trinity on Clarifai?
This launch is a win for developers seeking ownership and transparency.
By running on Clarifai, these models demonstrate our core value proposition: Twice the performance at half the price.
Key Differentiators:
• Competitive Benchmarks: Trinity Mini achieves 84.95% on MMLU (Zero Shot) and 92.10% on Math-500, demonstrating competitive capability in a compact MoE footprint (3B active parameters).
• Open Provenance: This is an end-to-end entirely American model, providing the legal certainty and control developers demand.
Ready to Build?
The Trinity family is available today with full Playground access and drop-in compatibility with OpenAI-spec clients and agent frameworks.
👉 Try the Trinity family on Clarifai: https://t.co/H7pwmYOafI
👉 Read the press release: https://t.co/5jma7xGtRN
Reasoning workloads demand speed. Agentic AI burns through tokens fast, and if systems can’t keep up, the use cases break.
At 544 tokens/sec on GPT-OSS 120B (benchmarked by Artificial Analysis), Clarifai Reasoning Engine proves GPUs can deliver the scale agentic AI needs.
Throughput is the real unlock for Agentic AI!
When reasoning models burn through tokens, speed decides whether they’re practical.
Clarifai Reasoning Engine, tested on GPT-OSS 120B by Artificial Analysis, delivered 544 tokens/sec on standard GPUs, the fastest GPU-based inference ever recorded.
For the first time, GPUs match (and even beat) specialty non-GPU accelerators — proving that high-throughput reasoning doesn’t have to mean hardware lock-in.
Try out GPT-OSS-120B model here: https://t.co/73X0qxbyog
👀🏆 @ArtificialAnlys just benchmarked GPT-OSS-120B on Clarifai!
500+ tokens/sec throughput, one of the fastest GPU-based inference speeds ever recorded without sacrificing latency, quality, or price.
I’m excited to join Comcast Connect 2025, where I’ll be speaking on the AI and Connectivity – Building Intelligent, Adaptive Systems panel.
We’ll explore how AI is transforming network infrastructure, enabling predictive operations, and creating seamless digital experiences.
Clarifai is excited to join Comcast Connect 2025!
Our CEO @mattzeiler will speak on the AI + Connectivity panel, joined by CPTO Alfredo Ramos and Sr AE Rachael Kloek.
Let’s connect at the event.
Book a slot with the team here: https://t.co/3XudPM2xKp
Proud of the team!
@ArtificialAnlys benchmarks confirm what we’ve been building for over a decade: AI infrastructure that delivers speed, flexibility, efficiency, and reliability — all in one platform. 🚀
We’re excited to share that @ArtificialAnlys has validated Clarifai's Compute Orchestration with GPT‑OSS‑120B as one of the top‑performing GPU inference platforms globally. 🚀
What the data shows:
• 0.27s Time‑to‑First‑Token (TTFT) → tokens start streaming almost instantly.
• 313 output tokens/sec → among the very fastest measured.
• $0.16 per 1M tokens (blended) → this undercuts the $0.26–$0.28 cluster while maintaining top‑tier performance.
Placed in the benchmark’s “most attractive quadrant” for high speed + low price.
This is the payoff from years of building and operating inference infrastructure:
👉 Press Release here: https://t.co/nSkshnTeAR
Excited to be at #CitiTMT in NYC!
I’ll join the Enterprise Applications in AI panel to discuss how @clarifai is driving the next wave of agentic AI.
📅 Sept 4, 8:10–8:45 AM EDT.
If you’re here, come meet me. I’d love to connect!
Agentic AI is reshaping how we measure intelligence!
It’s not just scores, but how models perform and feel in real use.
Two models can score the same yet deliver different experiences.
The future lies in measuring alignment, usefulness, and impact.
Final Call! 🚀 Tomorrow: Learn how to build & deploy enterprise-grade agentic AI.
Join @RelevanceAI’s Ahmed Alassafi & Clarifai’s Minh Tran for a live webinar.
📅 Aug 20 | ⏰ 4 PM ET / 1 PM PT
👉 https://t.co/Stam8yRMXm
#AgenticAI#EnterpriseAI
Join us today for a Live Technical Deep Dive into GPT-OSS!
We’ll explore GPT-OSS-120B and GPT-OSS-20B models, diving into their architectures, performance, and accessibility.
Join us today for a Live Technical Deep Dive into GPT-OSS!
We’ll explore GPT-OSS-120B and GPT-OSS-20B models, diving into their architectures, performance, and accessibility.
Join us today for a Live Technical Deep Dive into GPT-OSS with the Clarifai team!
We’ll explore OpenAI’s GPT-OSS-120B and GPT-OSS-20B models, diving into their architectures, performance, and accessibility. This free, open-weight milestone is set to shape the future of AI deployment.
📅 Date: Friday, 8th Aug 2025 (Today)
🕒 Time: 1 PM EDT (just a few hours away!)
Watch live here: https://t.co/9JlYldML9I
Day 0 support for GPT‑5 on Clarifai! 🔥
OpenAI has introduced GPT‑5, its most advanced reasoning and generative model yet. It offers performance leaps in accuracy, speed, context handling, structured thinking, and problem solving.
There are three variants: 🧵👇