We congratulate Levent Alpöge and Tristan Buckmaster on their remarkable mathematical work.
We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem.
While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.
However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs. unforced).
@WestJet Wow had the worst experience with an airline ever!! Stranded in calgary for 24hrs. The connecting flight to iceland got cancelled 😭😭😭 The staff refuse to give a hotel voucher since it's not due to a delay and is due to a cancellation.
Further, if you ever shared online a Claude Code/Codex session with encrypted reasoning blobs, they can be decoded and leak your personal data.
We did a preliminary scan of ~7,000 public traces and found 62 unique API keys, 33 email addresses, 33 passwords, and other sensitive data.
OpenAI’s o4 just showed that multi-turn tool use is a huge deal for AI agents.
Today, we show how to do the same with your own agents, using RL and open-source models.
We used GRPO on only 100 high quality questions from the BFCL benchmark, and post-trained a 7B Qwen model to orchestrate multiple tools without demonstrations from humans or teacher models. We measured a 23% increase in model performance on our eval set (1/n) 🧵
OpenAI’s o4 just showed that multi-turn tool use is a huge deal for AI agents.
Today, we show how to do the same with your own agents, using RL and open-source models.
We used GRPO on only 100 high quality questions from the BFCL benchmark, and post-trained a 7B Qwen model to orchestrate multiple tools without demonstrations from humans or teacher models. We measured a 23% increase in model performance on our eval set (1/n) 🧵
To interpret AI benchmarks, we need to look at the data.
Top-level numbers don't mean what you think: there may be broken tasks, unexpected behaviors, or near-misses.
We're introducing Docent to accelerate analysis of AI agent transcripts. It can spot surprises in seconds. 🧵👇
Introducing Autellix: An agentic AI system that accelerates agentic applications.
Run Deep Researcher, Google Co-Scientist, OAI Operator, or any program 💻—@langchain, @pyautogen, @crewAIInc, or just Python 🐍—4-15x faster than vLLM or SGLang⚡.
Paper: https://t.co/LvLq2B8P4v
People have been wondering which model will be the first to reach a 1400 Elo score; no one believed it would be Grok one year ago, but @xAI made it happen!
Large Language Diffusion Models
Introduces LLaDA-8B, a large language diffusion model that pretrained on 2.3 trillion tokens using 0.13 million H800 GPU hours, followed by SFT on 4.5 million pairs. LLaDA 8B surpasses Llama-2 7B on nearly all 15 standard zero/few-shot learning tasks while performing on par with Llama-3 8B.
We’re Scaled Cognition, developing the first ever models trained specifically for agentic applications:
1. Our first system, APT-1, is now #1 on agentic benchmarks.
2. It was developed by a US team for a total cost of less than $11M.
3. Khosla Ventures led our seed round ($21M closed in 2023), and Vinod Khosla joined our board.
4. We use a fully synthetic, RL-based agentic data pipeline, no human-labeled data.
5. APT-1 is now available for early access.
This is wild - UC Berkeley shows that a tiny 1.5B model beats o1-preview on math by RL!
They applied simple RL to Deepseek-R1-Distilled-Qwen-1.5B on 40K math problems, trained at 8K context, then scaled to 16K & 24K.
3,800 A100 hours ($4,500) to beat o1-preview in math!
Best thing is they open-sourced everything: the model, the training code (based on ByteDance verl library), and the dataset.
With the success of LLM agents like OpenAI Operator, we are entering a new scaling era, but how do we train these agent models?
We present InSTA, the largest training environment for LLM agents, containing live web navigation tasks for 150k diverse websites in multiple languages.
Website - https://t.co/nGIViX9NGc
Environment - https://t.co/nIs3ZiVfE1
🧵Thread below.
1/6
#AgenticAI #LLMs #OpenAl
With the success of LLM agents like OpenAI Operator, we are entering a new scaling era, but how do we train these agent models?
We present InSTA, the largest training environment for LLM agents, containing live web navigation tasks for 150k diverse websites in multiple languages.
Website - https://t.co/nGIViX9NGc
Environment - https://t.co/nIs3ZiVfE1
🧵Thread below.
1/6
#AgenticAI #LLMs #OpenAl
@AlexKrentsel@berkeley_ai 100% @AlexKrentsel It's such a privilege working on a team that fully appreciates the challenges and insights of both systems and AI !!
Thanks AK for sharing our work! 🙏
We replicate Deepseek's success under $5000 and share our training recipe on:
- Twitter: https://t.co/Sd2mWyvCDA
- Blogpost https://t.co/eHqApwRfnH
Come check it out!