🤔 How many examples does an LLM need to master competition-level math?
Conventional wisdom says 100,000+ examples.
Our finding? Just 817 carefully selected ones 🤩
With pure SFT, LIMO achieves:
📌 57.1% on AIME
📌 94.8% on MATH
LIMO: Less is More for Reasoning 📝
Anthropic CEO Dario Amodei says AI safety evaluations conducted on DeepSeek showed that it was the worst-performing model they had ever tested at generating potentially dangerous information
PDF parsing at scale is essentially a solved problem now.
With Gemini 2 Flash offering $0.40 per million tokens and a 1M token context, you can parse 6,000-page PDFs with near-perfect accuracy for just $1.
Every engineer is doing vibe coding with.
But that’s just the beginning. Soon, all knowledge work will be based on vibes—have an idea, add some rough context, and transform it into an essay, email, or PowerPoint with a deep researcher agent.
The top performers will be those with the best vibes and the fastest lowercase typing skills.
Could this handle more structured data, like finance or CRM datasets, in addition to contracts? Seems like the architecture could generalize well to those use cases too.
We built a knowledge agent that can do automated contract review against any knowledge base in minutes 📜🔎
Think: matching contracts against your company policies, compliance rules, negotiation playbooks, past agreements - it takes a legal/ops team a few hours to do this manually.
I’ve heard this use case in 2-3 calls already this week. You can do this by interleaving @llama_index agentic workflows with the right architecture for parsing, indexing, and retrieving your data (LlamaCloud).
Notebook: https://t.co/0xZEs97w8T
If you’re interested in this use case come talk to us! https://t.co/Ht5jwxSrQB
ColBERT got ridiculously fast in 2022 with PLAID. I thought that was as fast as it could get.
But Luca Scheerer taught us that you can make it 3x faster: a single CPU core can encode the query *and* search hundreds of millions of tokens in 100ms.
WARP—worth a thread tmrw?
We've been building LOTUS at Stanford and Berkeley to make LLM-powered data processing fast, easy and declarative.
LOTUS is an open-source query engine that makes programming as easy as writing Pandas and optimizes your programs for up to 400x speedups.
To celebrate the holidays, we're excited to share our release of LOTUS 1.0.0 with a batch of new updates that make reasoning over your data faster, easier and better than ever!
Code: https://t.co/qQVJ6Vg6fi
🧵👇
🚀 Excited to announce the release of Unity Catalog 0.2.1! 🎉
This update brings:
🐍 New unitycatalog-client Python library
📋 MODIFY permissions for tables
📈 Enhanced model APIs/UI
❄️ Iceberg Catalog APIs updates
Read the release notes 👉 https://t.co/JCemJSMpfV
#opensource
Today's AI landscape is reminiscent of the early automotive and aviation industries. Although we have seen remarkable demonstrations and early successes, the full transformative impact and proliferation of LLM systems are bottlenecked by robustness and reliability challenges.
Building on the analogy, massive leaps were needed to progress from the Wright Brothers' initial Kitty Hawk breakthrough to the contemporary aviation industry, where over 2M humans fly daily. Notably, the gap from Kitty Hawk to what is considered the dawn of commercial aviation with Jannus was ~10 years.
In this paper, Ion Stoica, along with collaborators @matei_zaharia, @joseph_gonzalez , @Ken_Goldberg, @haozhangml , @ml_angelopoulos, @shishirpatil_, @ChenLingjiao, @infwinston, and I, surveys the landscape and lays out a vision for advancing today’s LLM systems design into a mature engineering discipline with even broader deployed impact. This paper begins to address how we can reconcile the tensions arising from the value of these systems partially being their stochasticity and “creativity” (hallucination) and the engineering imperative to build robust, reliable 'compound AI' systems out of these noisy components.
SPECIFICATIONS: THE MISSING LINK TO MAKING THE DEVELOPMENT OF LLM SYSTEMS AN ENGINEERING DISCIPLINE
https://t.co/xnoLAFRjLy
Gen AI is still very much in the phase of fast price reduction/quality increase! Mosaic's Law (4x increase perf/$ yoy) in full effect.
https://t.co/uRzjDJwS97
Introducing Salesforce Connectors for Lakehouse Federation and LakeFlow Connect!
These connectors provide seamless access to Salesforce CRM and Data Cloud data in @databricks Unity Catalog for comprehensive data lifecycle management. https://t.co/yeUnjvJo4C
We’re excited to announce that Databricks SQL Serverless is now GA on Google Cloud Platform!
Optimize your data for BI use cases with:
▪️ Instant and elastic compute
▪️ Lower infrastructure costs
▪️ No management overhead
▪️ Improved performance
https://t.co/DvSYFG8CXc