Today we’re releasing OfficeQA — a new benchmark for end-to-end grounded reasoning that reflects the real work enterprises need AI agents to do.
More details below 👇
Automated prompt optimization (GEPA) can push open-source models beyond frontier performance on enterprise tasks — at a fraction of the cost!
🔑 Key results from our research @DbrxMosaicAI:
1⃣ gpt-oss-120b + GEPA beats Claude Opus 4.1 on Information Extraction (+2.2 points) — while being 90× cheaper to serve.
2⃣ The same technique also lifts frontier models (Claude Sonnet 4, Opus 4.1), pushing them to new SOTA benchmarks.
3⃣Versus Supervised Fine-Tuning (SFT): GEPA delivers equal or better performance at 20% lower serving cost. Even better → GEPA + SFT together gives the highest gains.
4⃣Lifetime cost analysis shows GEPA + gpt-oss is orders of magnitude cheaper overall. At scale, the one-time optimization overhead fades away — making optimized agents highly practical for real-world deployment.
#dspy #gepa #promptoptimization #airesearch
I hadn't noticed until seeing Noah's tweet that DSPy crossed 2M downloads/month apparently.
Yesterday was the highest day ever at 101,000 downloads in a single day.
@jxnlco Yeah even the component for text to GraphQL is super finicky to get right - few foundation models do this well OOTB. What's crazy is customers will come to you even without having a Knowledge Graph asking for KG-RAG. Blog here: https://t.co/8wGf06Zsd3
We built a thing! The Databricks Reranker is now in Public Preview. It's as easy as changing the arguments to your vector search call, and doesn't require any additional setup.
Read more: https://t.co/irmvWnY44O
Well, in the most lowkey way possible, and on a random Tue afternoon, @DSPyOSS 3.0 is out of beta.
pip install -U dspy
So many amazing people contributed to this. Thank you all!
(Release notes below. Stay tuned for a ton of stuff over the next few weeks!!)
Ever wonder what it'd look like if an LLM Judge and a Reward Model had a baby? So did we, which is why we created PGRM -- the Prompt-Guided Reward Model.
TLDR: You get the instructability of an LLM judge + the calibration of an RM in a single speedy package (1/n)
Really excited about ALHF, new work from our research team that lets users give natural language feedback to agents and optimizes them for it. It sort of upends the traditional supervision paradigm where you get a scalar reward, and it makes AI more customizable for non-experts.
Yes, this is a description of how the dspy.SIMBA optimizer works.
> a review/reflect stage along the lines of "what went well? what didn't go so well? what should I try next time?" etc. and the lessons from this stage feel explicit, like a new string to be added to the system prompt for the future, optionally to be distilled into weights (/intuition) later a bit like sleep.
Distilling through the weights is the dspy.BetterTogether strategy.
I'm at ICML 🇨🇦 and I'm hiring at @databricks. Visit our booth if you're interested. My scientific focus: It's 1972 in AI, there's an AI crisis, Dijkstra isn't here to save us, and maybe RL can. Why Databricks? The long road to AGI is being paved here and we have the real evals 🧵
We just published a Databricks App template that shows how to:
- Deploy a LangGraph agent asa Databricks app with a chat UI
- Automatically monitor MLflow 3.0 traces on Databricks (including syncing to delta tables, with Unity Catalog governance of traces)
I've also embedded my favorite OSS stack in this repo, uv, bun, shadcn, etc. Poke around :)
Now this is June 2025 so this repo has an opinionated Claude memory file where I have some of my favorite tricks:
3 self-verificatication loops:
1) running the devserver in the background to see python logs
2) opening playwright so that claude can take screenshots of the UI as you develop
3) deploying the app to databricks and then checking the service logs for any installation issues.
The verification loops are the most important way to get claude to turn that automation slider up.
@umichvoter Best EDs by vote-share though for Zohran were 56-70 at the North BedStuy/Bushwick border and Cypress Hills 54-02 in Southeast Cypress Hills!
Our talk at Databricks Data + AI summit is now live!
This talk showcases the MLflow 3.0 APIs, and how you create a self-reinforcing data feedback loop for GenAI systems on Databricks, featuring eval, monitoring, and subject matter collaboration
https://t.co/JBT93Mdz9v
I talked a bunch about the reinventing of the agentic wheel in my talk at Data+AI Summit! s/o Sutton and Barto and @tom_doerr for the pre-GenAI agent diagrams :) https://t.co/CqAJff1JT6
Here's the write up of my Data+AI Summit talk on the perils of prompts in code and how to mitigate them with DSPy.
As prompts grow in complexity, they begin to resemble programming. Don't program your prompts. Program your program. https://t.co/ptbkL76Zfd
Excited to finally announce what we've been working on for a year - MLflow 3.0: Unified AI Experimentation, Observability, and Governance 🚀
MLflow 3.0 enables continuous improvement of your AI systems through data, allowing you to trace, evaluation, and monitor your AI systems.
We've built MLflow 3.0 with enterprise in mind:
- human collaboration with subject matter experts
- data governance & security
- integration with the rest of the data ecosystem at Databricks
We brought the ideas and tools we built in Lilac to the world of GenAI observability, with a LOT more to come this year.
You try MLflow on Databricks for free today too! Link to blog & signup in the threads.
Thread of the new features in MLflow 3.0 for GenAI 🧵