GLM-5.3 at 310 tok/s!
Databricks inference is #1 in both speed and latency, again.
On our internal Databricks coding benchmark, GLM-5.3 is the strongest OSS model for coding, competitive with Fable 5 and Opus 4.8.
GLM-5.3-Flash at 270 tok/s!
Databricks inference team is on fire. So proud of the team!
On our OfficeQA Pro v2 benchmark, GLM-5.3-Flash delivers 10% higher quality than GLM-5.2 at just 1/10 the cost.
Fast, cheap, and good.
I am excited to announce that I am joining @perplexity_ai as research lead! We will be doing ambitious paradigm shifting work, advancing the frontiers in the open. If you want to join us in re-imagining continual learning, agent collaboration, and beyond, please reach out!
Your gaming PC can now serve frontier models at interactive speed using official checkpoints without extreme quantization!
Qwen3.6 35B → 8GB RTX 4060 laptop @ 39 tok/s
DeepSeek-V4-Flash 284B → RTX 5090 desktop @ 22-25 tok/s
GLM-5.2 753B → RTX PRO 6000 workstation @ 15 tok/s
Run your claude code or codex now with frontier model for $0
Meet FreeToken 🧵
Documents are the primitive of knowledge work and Databricks puts them to work 😤
We pair custom trained models with an agentic harness to reach frontier accuracy on the toughest extraction tasks:
- Long documents with thousands of pages.
- Large extraction outputs such as an invoice with thousands of line items.
- Complex schemas that require cross-page reasoning and computations such as a contract value that applies listed discounts across every recorded price.
Across 9,000 complex documents, our AI Extract Precision Mode reaches 94.7% accuracy!
An extremely important functionality for agents is to simply extract fields out of PDFs. This turns out to be harder than people think because LLMs are primarily trained on predicting the next tokens. This leads them to "autocorrect" things that they shouldn't autocorrect. We launched an AI Extract capability that just excels at doing just this task with very high accuracy (95% vs 87% for others) and extremely low cost. Check out this blog on how we did it. The function can of course be called directly from SQL and be used throughout the platform.
https://t.co/60XC2mIZ2E
We are excited to share @databricks 's new AI Extract achieves the new frontier at complex document processing tasks! It addresses a few key customers' pain points:
- Process documents with 500+ pages that exceeds 1M tokens
- Large, nested schema with 1k+ of objects
- Complex schema that requires frontier reasoning
Here is how we achieved it:
- In-house custom model for document extraction that is trained for complex real-world examples and optimized for serving large volume of workloads
- Custom agent harness that is designed for long document extraction. It decomposes large extraction jobs, executes smaller tasks in parallel, and reconciles them into one final structured output.
This is the methodology behind many of our work: identify key large workloads of enterprise customers, build in-house models + harness to address it, and ship to the customers through a close collaboration between the research + product!
Retrieve more, rerank, get better results. That's what most of us expect from search, but it turns out to be wrong! 🤯
I'm SUPER EXCITED to publish the 141st episode of the Weaviate Podcast with Mathew Jacob (@mat_jacob1002)! Mathew led the work behind "Drowning in Documents" during his time at Databricks and is now a Ph.D. student at the University of Washington working on ML systems!
This episode dives deep into "Drowning in Documents". This has been one of the most influential papers for us @weaviate_io as we are exploring scaling reranked retrieval. I think it is a must read for those working in Search and Information Retrieval. 🙌
We begin with an overview of the paper, and then dive into full scoring with cross encoders and phantom hits. We then cover Listwise Rerankers, what next generation cross encoders might look like, and ranking cascades.
On the topic of ranking cascades, I loved learning about Mathew’s work with Melissa Pan (@melissapan), Negar Arabzadeh (@NegarEmpr) and collaborators on “Natural Language Query to Configuration for Retrieval Agents”. Per-query “effort” prediction is certainly going to be a huge component on the future of these search systems! (Congratulations to OpenRouter 😆)
We then dove into Mathew’s work on TraceLab, a super exciting effort to understand coding agents such as Claude Code and Codex. ⌨️
This was a super fun conversation, and I really hope you find it useful!
YouTube: https://t.co/3lPicYhk5R
Spotify: https://t.co/sx0o06kilI
VLDB 2026 (https://t.co/gqsGojEhRE) will include a memorial session for Stan Zdonik on Wednesday, September 2, from 3:45–5:15 PM, at the Westin Boston Seaport District. It's open to all, even if you aren't registered for VLDB. Please register here: https://t.co/0iRLZ7Oe1v.
Introducing Toast 1, our first specialised search agent.
Toast 1 sets a new Pareto frontier for agentic search models.
Frontier search quality, across all domains, 12x faster, at 1/10th of the price.
We’re excited to announce that @ElectricSQL, the maker of PGlite, is joining Databricks to bring WASM Postgres to AI agent sandboxes.
Together, we’re excited to extend Databricks' Postgres capabilities from the lakehouse to the edge.
Electric is leading the way in pioneering backend technology purpose-built for agents. With PGlite, every agent gets its own WASM Postgres build right inside the sandbox where it runs, providing ultra low latency access to local context. And Electric’s real-time sync engine synchronizes distributed state back to a central Lakebase, enabling teams of agents to collaborate without losing track of shared context.
Please join us in welcoming the Electric team! https://t.co/U1pMDP5UF8
Kimi K3 at 239 tokens/s!
Databricks is now #1 for Kimi K3 inference speed and latency on Artificial Analysis.
A huge 2.8T parameters model, it’s the largest oss model we’ve ever served. We make sure the GPUs go brrr at Databricks.
🎉 AI Classify + vector search beats frontier models on accuracy and cost!
At Databricks, we see customers with large taxonomy classification problems like vendor name normalization and biomedical entity linking. Typically they reach for regex, trained classifiers, or direct LLM calls, but these methods struggle on cost, maintenance, and context limits.
To address this, we experiment with 3 classification methods:
1. Vector search
2. AI Classify Workflow: shortlist label candidates with vector search, then classify with the Databricks AI Classify function
3. Direct frontier model calls (GPT-5.4 mini, GPT-5.6 Luna, Sonnet 5, Gemini 3.5 Flash) with prompt caching.
We find that the AI Classify workflow yields the best overall accuracy at ~100x cheaper than the next best direct frontier model.
Shoutout to my teammates @arnav_thebigman@ivanzhouyq@nihit_desai !
We're raising funding at $188 billion valuation to double down on our AI strategy focused on three priorities:
1️⃣ Unity AI Gateway - our multi-AI governance solution that helps control costs.
2️⃣ Genie - our AI coworkers that actually understand your business data.
3️⃣ Lakebase - our serverless Postgres database specifically for AI agents.
https://t.co/Bnc89qxjLG
You may have heard that GLM-5.2 at 328 token/s is cool,
How about 392?
Databricks is now #1 in inference speed for GLM-5.2 on Artificial Analysis. It's a great model, and we did a lot of optimizations.
Really excited to open source a new project: Omnigent, a meta-harness for AI agents.
It lets you build multi-agent coding and custom agents, sitting above Claude Code, Codex, Pi, and agent SDKs to let you compose them. It also adds live collaboration and rich control policies.
New Product Update: We trained a retrieval-specialized model for Knowledge Assistant. It matches Claude Sonnet 4.5 retrieval quality at substantially lower latency.
Introducing Instructed-Retriever-1.
Genie has transformed how Databricks users work with data, with 3x the accuracy of generic agents. We're sharing some of the research behind it and what makes building data agents challenging. Super proud of our research team's impact with this! https://t.co/eLB2ElVo8S
Life update: I joined Databricks this week!
I thought I’d do another startup after Hyperbolic, but I was surprised by how startup-y Databricks AI is.
@alighodsi, @pwendell, @matei_zaharia are in full founder mode. They’re the best founders I’ve met. I like working with people who aren’t “normal” and they definitely aren’t. For example, they invited me to an all-hands before I joined.
I’m also impressed by how many former founders are here. @akhilgupta and @hanlintang are incredible leaders.
A big bonus: I finally have unlimited Claude Code & Codex tokens!
AI adoption on the Databricks AI team is insanely high. Every engineer I’ve met uses AI heavily and shares their own ways to drive agents. Many talented people here.
I’m super pumped for this new journey!