Alibaba's PoS: Beyond memory for long-horizon agents
An inference-time framework that maintains explicit belief states, detects when agents stall (Belief Trapping), and recovers with targeted constraints. Achieves SOTA across 4 benchmarks, with gains up to 37.89%.
Can agents improve their own skills without generating any new rollouts?
This paper introduces SkillRefiner, which learns entirely from historical agent traces. It turns past mistake patterns into targeted edits to the agent’s existing skill, by basically treat deployment history like a bug report database.
Additionally, repeated successful behaviors get reinforced, while repeated failures become new guardrails, with an extra evidence check to avoid learning the wrong lesson.
Across all eight settings, it improved the original skill while using 1.4-13x fewer refinement tokens than the "extracting lessons from each past run individually" (Trace2Skill) method, and 2.4-42x fewer than the "generating new runs just to test whether each skill rewrite works" (GEPA) method.
https://t.co/YhWHRRSe3E
After tinkering with Canva Code, we built a few simple tools you can use for your workday. Prompts in the thread! 🔧👇🏻
First up: a Pomodoro timer with the intervals and design we wanted, added directly to the file we’re working on. #MadeWithCanvaAI
Plan with one model
Implement with another
now live in Command Code
Your planning model explores the codebase and writes the plan. Approve it, and your implementation model takes over mid-run.
Set both in /config → Feature models
The same model, with the same weights, scores 62% in one agent harness and 33% in another.
@adithya_s_k and the @huggingface team just released the ultimate guide to multi-harness RL, and it's one of the most practical RL write-ups this year, and everything open!
The trick is simple. Don't touch the harness. Point it at a proxy instead of the model. The proxy speaks all four API formats coding agents use (OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, Gemini). It records the exact token ids and logprobs vLLM sampled, and you train on that. You don't change a single line of Claude Code, Codex or OpenCode.
Results:
🔹 Trained across 4 harnesses at once, LFM2.5-2.6B by @liquidai went from 42% to 54%
🔹 31% fewer tool calls, thanks to a small bonus for solving tasks in fewer steps
🔹 Training in OpenCode alone took OpenCode from 34% to 58%, but the multi-harness model improved everywhere
They also tried the shortcut everyone reaches for: fine-tune on 3,189 successful rollouts from Qwen3.8-27B. Imitation plateaued at 47.5%, below both RL runs. Copying a bigger model doesn't get you there. Practice does.
The best part is that everything is open: the capture proxy in OpenEnv, the trainer in TRL, the tasks, the SFT data, the training code and all seven trained models.
Agents will run in dozens of harnesses.
Now open models can be trained for each of them, by anyone.
Read it here 👇
https://t.co/s2I8GI9SLS
Validate agent fixes before shipping them with LangSmith Engine v2.
Engine now tests and validates fixes automatically:
✅ Replicates the issue with the same deployment environment and same inputs
✅ Builds a fix, tests it, and iterates on improving it until it has a satisfactory solution
✅ You review the fix and deploy it with a ready-made PR
How to design rewards for post-training frontier image models?
Our research suggests human preference reward is necessary, but insufficient: A preference model may still reward outputs that look appealing but miss details, introduce unrequested content, or exhibit other forms of reward-hacking.
We therefore optimize towards a composite reward:
- Bradley-Terry reward model trained on ~5.6M pairwise human votes
- Faithfulness reward from auto-generated prompt checklists evaluated by a vision-language model
- Constraint reward covering explicit and implicit user intent
- Anti-reward-hacking rubric rewards targeting failures such as garbled text and photorealism drift
This post-training recipe improves two already-strong open image models:
- Post-trained FLUX.2-dev gains 69 Elo points on our live T2I leaderboard, scoring 1202
- Post-trained Ideogram 4 gains 20 Elo points reaching a score of 1224 and surpassing all publicly listed open models (as of Sep 04, 2026).
Offline ablations with Gemini 3.5 Flash as the judge, show that these reward components are complementary: win rate against the base model increases as we add faithfulness and then constraint rewards on top of preference-only training, reaching 64.2%.
Finally, we ensemble policies trained with and without the anti-reward-hacking objective directly in weight space, further increasing win rate to 66.0%.
We put @OpenAI dots to work on the final checks that could stop a $420 million acquisition from closing.
Using 22 deal documents in Box, the dot uncovers a $3.85 million funding gap at the planned closing date and a missing customer consent and saves a source-cited readiness brief back to Box.
Then we add a signed equity amendment and ask it to update the brief. The funding gap is resolved. The missing consent still blocks closing.
Connect Box to ChatGPT to put your business documents to work.
Cloudflare AI Gateway now supports native web search API integration in partnership with Ceramic[.ai], Exa, and Linkup. Developers can now inject real-time web context into model inference calls via AI Gateway, REST APIs, or Workers bindings. https://t.co/0Q1bhexFGX #BirthdayWeek
AnythingLLM 1.17.0 is out. We overhauled the Meeting Assistant to run @NVIDIAAI Nemotron 3 Diarization + @moondreamai Parakeet Redux, on your device.
3x smaller download
70-200x real-time on your hardware (GPU or CPU)
Integrated into your knowledgebase.
https://t.co/YK7BU4g8Ts
Introducing AstaBrief 8B, an open model that turns complex research questions + literature excerpts into cited reports.
Run it locally on your own hardware, with open weights + training data you can inspect & build on. 🧵
🤗 Download: https://t.co/KvTMbe1e9j
Banger paper from NVIDIA on long running agents.
A model can accept 128K tokens of context and still make more mistakes the longer it works through a task.
If your agent loses its place partway through a long table or ledger, this work measures what causes it.
The setup:
Long-Transduction asks a model to keep reading, updating and outputting state-dependent results over thousands of outputs, and varies three factors separately.
Results:
Across seven open-weight models, accuracy drops 62.8% when context grows from 4K to 128K, 36.5% when only the input format changes, and 39.9% when the per-step operation gets harder.
Paper: https://t.co/kjWSHzoegq
The Comfy Dev Platform Challenge is open!
Partners: @ltx_io & @bfl_ai
For two weeks (10/5–10/19), we're challenging the community to open-source a paid AI feature.
🏆 $10K grand prize
🖥️ 3x NVIDIA RTX 5090s
🎁 Bonus prizes from LTX & Flux
⚡ Free credits for the first 100 sign-ups
Register and learn more 👇
Announcing our new and improved Analytics, giving engineering leaders transparency into consumption, model efficiency, and adoption, broken down by model and by user across every session in the organization.
You can see Jev Router's thinking process in OpenRouter Chat.
The routing insights panel shows the model it picked for each turn and the scores behind it, like task, difficulty, precision, and larger model benefit.
Test Jev Router here: https://t.co/kVWPuaDkgr
NUS researchers release APPL on Hugging Face
APPL uses structural priors as the interface between skill learning and composition.
It boosts OOD generalization from few demos: 89.6% skill success with 2 demos,
50% on shifted objects vs 10% for DP.
Who wants more speech training data?
YODAS v3 is now available @huggingface
It’s 1.1M hours - the biggest audio dataset ever. And the first at this scale with stereo audio at 48kHz, with timestamped transcripts and translations. 100+ langs
https://t.co/KLS44NvDC5
CC-BY-3.0
Our Neuroevolution textbook is finally in print!
Free online edition: https://t.co/3IDTb1remp
Pre-order: https://t.co/SUpqBz1HK8
I am incredibly grateful to my co-authors Sebastian Risi, Yujin Tang, and Risto Miikkulainen for making this happen. Neuroevolution is a subject very dear to my heart. It is the field that convinced me that nature has already figured out how to build intelligence: through evolution, collective behavior, and adaptation under constraints.
This idea, that intelligence emerges from evolution operating under constraints rather than unlimited resources, is what eventually led me to founding Sakana AI here in Japan.
The concepts in this book about open-ended creativity and self-organizing systems are exactly what we build at @SakanaAILabs. Our name and logo are inspired by schools of fish moving together, adapting as one. It is the core philosophy of what we build. This book captures the theoretical foundations of that belief.