🎮 Qwen3.8-27B + RSIGame > GPT-5.5 one-shot in game development. How?
We introduce RSIGame, an autonomous framework that turns game generation into a long-horizon RSI process: agents explore, diagnose, improve, verify, and keep evolving the games they create—while monitoring global progress and preserving the best checkpoint.
Fully open-source:
🎮 https://t.co/dytYjcuiaR
📄 https://t.co/zamAtXwizT
💻 https://t.co/xx3zSVwVGm
🚨 VLAs struggle on long, context-dependent tasks, so dense progress, a score at every step, matters for them.
🤔 Can progress reward models label dense progress for these context-dependent tasks?
🧭 Our work shows why they fail: not because they are blind, but because they get lost in the context. Given the right context, the same progress reward models are reliable again.
🧵 So we propose ProgressCompass, a training-free agentic system: off-the-shelf progress models + VLMs, guiding progress estimation in long tasks.
📄 ProgressCompass: Embodied Progress Reward Models Are Lost Without the Right Context
🌐 https://t.co/VLHUZEZ9lZ
🤖 Reusing skills is one possible way for embodied agents to generalize. Manipulation varies but shares a few skills: wiping a table or a window is one skill. Many works give agents skills to plan with. But where do these skills, and their data, come from?
📼 Today, VLMs mostly label videos one at a time, and nothing carries over. People don't learn that way. As experience streams in, we spot skills we know, add new ones, and reuse them later. This streaming setting matters, but it has received little attention.
🧠 VLMs may become the brains of embodied agents. So we ask: from streaming experience, can they find reusable skills and keep one consistent skill library?
🧵 Video2Skill: From Streaming Experience to Reusable Embodied Skills
Why it matters:
📈 It can label much more data for training agents.
🔁 It is planning in reverse. If a model can't find skills in what it has seen, how can it plan with them for something new?
📄 Paper: https://t.co/hrKsEI1WZE
🌐 Project: https://t.co/1QsXmGeTrb
🤗 Data: https://t.co/vlKWhQflC5
4/4 Results: +7.4% over uniform sampling on Video-MME at 8 frames, 67.1% at 32 frames — best among all sampling baselines, and ~4× faster than LLM-agent / video-RAG pipelines.
Code is out — try it on your own videos 👇
💻 https://t.co/wHYDwReLGh
1/4 🎉 LENS is accepted to #ECCV2026!
Given a fixed frame budget for long-video QA, should you spend it on fine-grained detail or on broad temporal coverage? Our answer: let the model decide, per question.
📄 https://t.co/Szp4mtzMsV
🌐 https://t.co/Yc3rzBh4vM
3/4 For the temporal branch, frames form a video graph via pairwise SSIM, and manifold ranking propagates relevance along it — suppressing isolated noisy peaks that fool per-frame retrieval scorers. No training, no extra models, works with Qwen2-VL/2.5-VL, LLaVA-OneVision, GPT series.
pySpatial is accetped by ICLR 26!
We introduce pySpatial, a visual programming framework that equips MLLM agents to interface with spatial tools via Python code generation.
• 🧭 3D Reconstruction – Feedforward 3D reconstruction with novel view synthesis turns images into explorable 3D scenes, improving spatial reasoning.
• 🔍 Interpretability – Produces grounded, interpretable, and executable 3D visual programs with verifiable intermediate results.
• 🤖 Real-World Applications – Deployed on a quadrupedal robot for navigation in complex environments using purely VLM-based reasoning.
project page: https://t.co/ise82hoknk
code: https://t.co/od8cWXXOs3
It’s been an exciting journey to see this come to life: BFM-Zero🤖 A behavioral foundation model capable of zero-shot goal reaching, tracking, and reward optimization — all from a single latent space. Check out Yitang's post for more details!
🚀 We are thrilled to release a new open-source Deep Research Agent, Cognitive Kernel-Pro, from Tencent AI Lab! We focus on building a fully open-source agent with (to the maximum extent) free tools, showcasing impressive performance on GAIA with Claude-3.7-sonnet and surpass the counterpart, SmolAgents by a large margin.
In addition, we study the training recipe for an open-source Deep Research Agent Foundation Model. We curate high-quality training data (queries, trajectories, and verifiable answers across web, file, code, and reasoning domains). Our finetuned Qwen3-8B (CK-Pro-8B) surpasses WebDancer and WebSailor with the similar model size on the text-only subset of GAIA.
📜 Paper: https://t.co/CHLs2U0JLy
🔧 Code: https://t.co/8dxtnjdsxn
🤗 Data & Model:
https://t.co/jzaV1LywnK
https://t.co/KD8k9WqGM9
This work builds on the previous efforts of Tencent AI Lab (Fig. 2). Be sure to check them out if you're interested!
🚀 Check out our paper: WebEvolver: Enhancing Web Agent Self-Improvement with Coevolving World Model, from Tencent AI Lab!.
We present a world model-driven framework for self-improving web agents, addressing critical challenges in self-training—such as limited exploration and performance plateaus.
🔍 Key Innovations:
- Co-Evolving World Model:
The world model is implemented as an LLM that predicts the next webpage state (observation) given the current state and a planned action.
In addition to fine-tuning the agent's policy model using self-generated trajectories, the same data is repurposed to train the world model.
- World Model as Web Server
Training Phase: sample pseudo trajectories by replacing the real web server with the world model.
Inference Phase: simulate the outcome of candidate actions with 1-2 step look-ahead planning, to help better select actions.
📊 Results on Real-World Tasks (Mind2Web-Live & WebVoyager):
✅ ~10% higher success rate vs. pure self-training.
✅ significantly fewer environment interactions—efficient yet powerful!
arxiv: https://t.co/7tMVy4iI0J
code: https://t.co/wzutAx8zNR
🚀 Introducing WebCoT! By capturing and verbalizing web agent cognitive patterns (reflection, branching, rollback) as chain-of-thought, WebCoT outperforms rejection sampling distillation by 10% on WebVoyager, Mind2web-live, and SimpleQA. https://t.co/oCruyDg3UL
🚀 Introducing VScan! A two-stage visual token reduction framework for efficient multimodal LLMs, enabling up to 2.91× faster inference and 10× fewer FLOPs -- while keeping 95.4% of original performance. Check it out at https://t.co/xCxmY2PhZo!