@thsottiaux Please offer a $500+ subscription tier! The current limits aren’t enough—GPT-5.6 Sol burns through the entire seven-day allowance in just 1–2 days.
🚀Hy3 is here.
295B MoE. Best in its size class. Rivals trillion-scale flagships.
Reliable and affordable for most agentic usecases.
Apache 2.0. Friendly for commercial use.
FREE API for 2 weeks → https://t.co/EyURKwTdgi
🤗 https://t.co/twqJpqb2SL
📖 https://t.co/4uEkIU1cW4
@SethiLiam Haha, definitely hoping for good AC. I’m most excited about Knapsack RL. It’s about RL scaling and compute allocation for better LLM exploration!
I’ll be heading to ICML 2026 in Seoul and presenting several posters on LLMs, RL, Agents, and Optimization. If you’re also attending, I’d be happy to connect and chat. See you in Seoul! #ICML2026
Can imitation learning avoid compounding errors without adversarial training?
Yes. Our ICML paper introduces Dual Q-DM, the first non-adversarial IL method with a guaranteed O(H) bound, matching adversarial IL.
The key mechanism: value flow
#ICML2026
How can we train phone-use agents that actually work on real phones?
Real apps are faithful, but slow, stateful, and hard to reset.
Mock apps are scalable and verifiable, but may not transfer.
Announcing PhoneBuddy-4B + our 5-paper phone-agent stack. 🧵 (1/n)
This is inspired by a growing line of work on game-dev agents, coding-agent evals, and interactive world building.
Excited to see more people pushing this direction — from small games to richer playable worlds, and maybe broader creative tools.
cc @iamwaynechi@sethkarten@YixiongFang@valeriechen_@OpenHandsDev
We wanted to ask a simple question: can a coding agent actually build a game you can play?
GameCraft-Bench: 140 Godot tasks across 15 families. Agents ship a complete project and are judged by replayed gameplay.
Best score: 41.5%.
Far from solved.
#LLM#CodingAgent#Game
https://t.co/agvjo62kJU
Why games?
They sit between two futures.
One path goes toward interactive world-building: small games → bigger worlds → long-horizon production.
The other touches creative pipelines: scenes, assets, timing, feedback, and eventually film-like production tools.
Why can adversarial imitation learning succeed with so few expert demonstrations �� sometimes even just one trajectory — and still work over long horizons?
Excited to share that our work tackling this question has been accepted by IEEE TPAMI! 🎉
We develop a stage-coupled analysis for TV-AIL and show a horizon-free imitation gap, offering a theoretical explanation for the strong empirical performance of AIL in small-sample regimes.
🥳Excited to share our work on imitation learning theory, accepted at IEEE TPAMI! Joint work with @ZiniuLi, Yang Yu & Zhi-Quan Luo.
paper link: https://t.co/BC0FUg4WnD
👋Hi /haɪ/, we're the Tencent Hy /haɪ/ team🐧
Today, we open source Hy3 preview (295B A21B), a leading reasoning and agent model in its size, with great cost efficiency.
Give us feedback to help improve Hy3 official!
🤗 https://t.co/jc10JODXJ8
📖 https://t.co/VIRoNnwng0
I strongly suspect that Claude Mythos is a looped language model, as described in the paper "Scaling Latent Reasoning via Looped Language Models" from ByteDance
The authors of that paper called out graph search as one of the areas where looping provides a huge theoretical advantage over standard RLVR. And look at where Mythos blows out its competitors the most
[1/8] Excited to share our new paper:
“Understanding by Reconstruction: Reversing the Software Development Process for LLM Pretraining”
👉 https://t.co/1DSokbPksD
We turn static repos into long‑horizon agentic trajectories! 🧵👇
Dollars Matter, Scores Flatter. 💰
Can LLM agents really do high-valued experts work end-to-end — not just “solve a question”?
Introducing $OneMillion-Bench benchmarks real expert work, priced in dollars 💵