I now constantly get questions about the SAAS meltdown, role of AI, system of records etc. I don't have an answer to all these.
But I do know that we saw an acceleration in our business in Q2, Q3, and now finished the year with accelerating Q4.
The question is, why?
Short answer: AI. But the underlying reason is subtle. We are growing fast because we are finally removing the biggest bottleneck in data: the technical barrier to entry.
For years, if you didnβt know SQL, Python, you were locked out of the value chain. That has changed fundamentally with the πππ§π’π πππ¦π’π₯π², and it is the "secret sauce" behind our recent momentum:
β’ πππ§π’π: Analysts can query data without any SQL. I use this every day myself.
β’ ππππ πππ’ππ§ππ πππ§π’π: Builds end-to-end AI models for you, similar to Cursor for ML on your data.
β’ ππππ ππ§π π’π§πππ« πππ§π’π: Write Spark pipelines, does plumbing, troubleshooting.
We've been talking about DATA + AI democratization, but generative AI finally enabled it in a way that wasn't possible before. That's why we're seeing a market response.
Take πππ€ππππ¬π ππ¨π¬ππ π«ππ¬. We launched this serverless engine for agents and apps recently. At 8 months into its journey, its revenue is already 2x what our Data Warehouse product was at the same stage.
All this taken together, we ended up with the following stats for Q4:
π $5.4B Revenue Run-Rate, growing >65% YoY
π $1.4B AI Revenue Run-Rate
π FCF Positive for the year
π NRR >>140%
https://t.co/yq3riYyr8r
Attackers have frontier AI. Defenders need a frontier AI ecosystemβthe best open and closed models, force-multiplied by a global community.
During the Hugging Face incident, closed AI blocked essential forensics. An open-weight frontier model helped contain the intrusion.
Thatβs why we created the Open Secure AI Alliance.
This is one of those unintuitive things. Agents that cook longer are often worse. Genie just gets to the results faster. Ontology will be key to getting these agents the context they need to get the answers right quickly.
We found that improving data agent quality also improves efficiency. Sometimes more is less: agents that take long, exploratory random walks are often less likely to arrive at the right answer.
We put Genie Code head-to-head against three leading general-purpose coding agents on 400+ real user data tasks.
Result below π
Next up on @AltimeterCap First Pass: A conversation with @matei_zaharia from @databricks on Omnigent
Agents are in their infancy, and so are the ways we develop and build them. Omnigent is a meta-harness aimed at solving many of the gaps in agent development today
We're raising funding at $188 billion valuation to double down on our AI strategy focused on three priorities:
1οΈβ£ Unity AI Gateway - our multi-AI governance solution that helps control costs.
2οΈβ£ Genie - our AI coworkers that actually understand your business data.
3οΈβ£ Lakebase - our serverless Postgres database specifically for AI agents.
https://t.co/Bnc89qxjLG
Excited to be releasing FrontierFinance, the largest and most challenging open benchmark for evaluating AI agents across the full investment workflow!
FrontierFinance is substantially harder than current finance benchmarks: Existing benchmarks like FinanceBench and Finance Agent focus almost entirely on data extraction.
FrontierFinance spans diverse use cases across the full investment process: Screening & Discovery, Company Research, Sector/Industry/Macro, Earnings & Events, and Coverage & Catalyst Monitoring.
Created for ambiguous, long-horizon agents: 220 examples paired with 11,543 expert-crafted rubrics, following Samaya's Criteria Eval methodology. The rubrics are what let us evaluate the reasoning and steps behind a true expert-level output, not just a plausible-looking one.
Evaluations: We evaluated Claude Fable 5, Claude Opus 4.8, GPT 5.5, Gemini, open-source models including GLM and DeepSeek, and others. We used the same public rubric and a standard harness for financial tasks. Samaya's AI system reached state-of-the-art accuracy at 50.8%, at 4x lower inference cost than Fable 5. Next best was Fable 5 (49.2%), then Opus 4.8 (45%) and GPT 5.5 (43.5%).
We're releasing the benchmark, methodology, and full evaluation results - see link in comments.
Future releases: FrontierFinance was curated from Samaya's larger internal set of ~5,000 examples, and we plan to release subsequent, harder benchmarks as well as a more detailed technical report!
Gartnerβs Magic Quadrant for Analytics and Business Intelligence (BI) is out, and Databricks was named a Visionary in our first appearance, the highest debut for any vendor in this MQβs 20+ year history.
BI has already changed. Anyone can drill from a signal down to the truth behind it and agentic loops make sure every answer has the full story. Thatβs where we're ahead with Genie and AI/BI. I use it every day to see the key signals, understand what changed and why, and decide what to do about it.
Huge congratulations to the teams, and thank you to our customers.
At 11k employees, our AI costs are going up. Which model & harness should we use to lower cost but also retain great quality?
We didn't want to blindly trust public benchmarks. So we ran a comprehensive evaluation on our tasks, code base, infra. It's been produced by more than 3,000 software engineers, spans 3 hyperscalar clouds and many languages and tasks.
The results are surprising. We find that for the SAME mdoel, the choice of harness can significantly save costs (~2x). We also find that GLM 5.2 performs extremely well. We run Omnigent in front of these and can easily multiplex different harnesses and models for different tasks.
Check it out:
https://t.co/hiLtLZn1cK
My co-founder @rxin personally wrote this really good paper that explains the main idea behind postgres Lakebase as well as LTAP. It almost serves as a primer on how transactional databases are built and how Lakebase and LTAP work. Maybe more importantly, what are the tradeoffs, and what are you giving up by adopting this new approach. Highly recommended reading:
https://t.co/d5V12x1Fqr
I've said something similar many times...the usage you give an API gives hints on what you're solving. A trusted middleman is really the only solution here...@databricks allows one to consume a model in a safer way, which is what we do at @unconvAI. But still, farming out your intelligence will have many implications for control and ownership of IP. A world where intelligence is cheap but value is created by many is preferable IMO to a world where control of intelligence and value is the hands of a few players.
Databricks ranks #1 on NVIDIAβs SOL-ExecBench kernel leaderboard, in the L1 single operation track, powered by KDA (Kernel Design Agents) π
Whatβs crazy is: we 100% leveraged AI agents to beat the competition.
This is a sneak peek at recursive self-improvement. The core frameworks we used were KDA, Humanize, and Omnigent: Claude writes code, Codex reviews. Together, they enabled agents to run autonomously for as long as possible. The key is setting up the right framework to let the agents cook.
This work was driven by @leshenj15 at Databricks, in collaboration with NVIDIA and MIT HAN Labβs @LigengZhu and @DongyunZou03 .
Databricks AI is like a neolab. Join us if youβre cracked!
Good lessons on managing AI spend. We support a lot of this natively on Databricks with Unity AI Gateway, which makes it easy to analyze and control usage in one place, and itβs also easy to set these up as policies with the open source https://t.co/zh1P01h1B5 framework.
Update from my post from 2 months ago when Genie Code just passed the threshold of being half the code generated on @databricks. Attached the new update from today, AI is now 3x the code written by humans on the platform!
Why the Frontier Ecosystem must be Open β Matei Zaharia and Reynold Xin, Databricks https://t.co/lSnxhhgmbc
@databricks cofounders @matei_zaharia and @rxin explain why Databricks is moving into the infrastructure layer for enterprise agents, how Omnigent creates a shared harness for coding agents and custom agents, why LTAP and Lakebase rethink the split between operational and analytical databases, why agent security needs contextual policies and spend controls, and why the future of software may be as simple as getting the right data in place and putting agents on top.
@heng_yan Oh you're right! Key point being that open source GLM is over 300 tokens per second. This matters for agents (we all hate waiting 10 minutes for responses). Proprietary frontier models are at best at 100 tokens per second. So 3x speedup really matters for agentic workloads.