@dronathon@Mei_keyuu@chuanyang_jin Thank you!
I think better control over condition boundaries and specificity is needed. SMC reduces noise and captures stable modes, but the hypothesis space does not explicitly represent where it applies or at what granularity. That boundary still depends on the base model.
Can your agent smoothly adapt to your preferences across sessions?
Meet HyperTrace, a training-free LLM personalization framework. It tracks hypotheses about what you want now and what tends to last, rather than reducing you to a fixed profile or a list of memories. 🧵
Can your agent smoothly adapt to your preferences across sessions?
Meet HyperTrace, a training-free LLM personalization framework. It tracks hypotheses about what you want now and what tends to last, rather than reducing you to a fixed profile or a list of memories. 🧵
Paper: https://t.co/e9yFklwi5o
Code: https://t.co/M7Em1pwyox
Huge thanks to @Mei_keyuu, Minghao Shao, @chuanyang_jin, Yusong Wang and Ailiang Lin for your amazing contribution and support!
On PRISM, HyperTrace demonstrates steady online adaptation. It also leads with 80% of feedback withheld and under frequent topic shifts, remaining robust in realistic settings where feedback is sparse and session switches rapidly.
Embodied Recursive Self-Improvement (RSI) needs more than a model—it needs an environment that learns where the agent fails.
🛵 We introduce DeliveryGym—an adaptive RL environment built in Unreal Engine 5, where embodied agents learn to navigate Paris, deliver food, and earn money.
The key idea is simple: earnings are the reward, and failures shape the curriculum. DeliveryGym automatically finds where the agent struggles and generates harder tasks targeting those weaknesses.
📈 With this adaptive RL loop, Qwen3-VL-4B improves net income by 54.3% in a shift.
✨ A step toward embodied agents that fail, adapt, and improve through an environment that evolves with them.
A coding agent can build an entire harbor and still forget that objects need something underneath them.
We gave GPT-6 and Fable 5.1 the same prompt and asset pack in Code4Scene. 👇
Curious how other frontier models would handle the same scene. Who should we test next?
GPT-6 coded an entire world: a busy harbor you can step inside.
From a single prompt and an asset pack, it assembled a container port with towering cranes, docked ships, rows of containers, and forklifts across the yard, all within an explorable 3D scene.
A first look at our upcoming Code-for-3D benchmark, featuring the text-to-scene setting.
How to teach AI assistants to infer human minds and proactively provide online assistance?
Sharing our ICML 2026 paper 🧠 𝙈𝙞𝙣𝙙𝙕𝙚𝙧𝙤: Learning Online Mental Reasoning with Zero Annotations!
🤗 Code and weights are fully open source
🌐 https://t.co/xC2yZ1WxWy
What are users thinking during their interactions with LLMs?
We introduce ThoughtTrace — the first large-scale dataset that captures what users think during real-world human–AI conversations, not just what they type.
→ 10,174 thought annotations
→ 2,155 multi-turn conversations, 17,058 turns
→ 1,058 users
→ 20 LLMs
These thoughts improve user behavior prediction (+41.7%) and model alignment (+25.6%).
This opens a new paradigm of user-centric LLM research. Full information in the thread 🧶
Read our paper: https://t.co/lRYJvGJ7bb
Check our project website: https://t.co/AupCn1YQOk
Environment generation is the missing scaling axis for embodied AI.
Introducing SimWorld Studio: a self-evolving factory for endless interactive 3D env where agents act, fail & learn.
Env-agent co-evolvution improves navigation success 50% → 90%.
From a prompt, our SimCoder writes code to automatically build an interactive world. Agents train inside it. And their performance shapes the next world.
Agents can now live like humans in a virtual city built with Unreal Engine: exercising outdoors, strolling through parks, or even playing chess with friends!
🌍 Check out SimWorld: an open-ended simulator for LLM agents to act, perceive, and plan in richly embodied, endlessly diverse virtual worlds.
SimWorld now supports plug-and-play integration with arbitrary UE environments. Deploy agents in your own scenes and build fully customizable simulations for autonomous driving, collaborative games, or social experiments!
SimWorld is a strong testbed for evaluating embodied agents, as it provides:
- ⚙️Physical + social simulation: realistic physics (collision/friction/inertia) and social dynamics like traffic rules/flows.
- 🤖 LLM/VLM-friendly interfaces: a Gym-style API with multimodal observations and grounded language actions.
- 🤔 Long-horizon tasks + reasoning: diverse physical/social tasks for systematic training and evaluation.
1/
#LLM #Agent #UnrealEngine #simulation