Mirroring reality into a world model that can be continuously verified and improved.
Building an OS where the world model and agent co-evolve through interaction and feedback.
From mirroring reality to pushing beyond the known, MirroS is transforming the future of intelligence.
At MirroS, weโre using mechanistic interpretability to understand which ingredients of physical intelligence emerge with scale, and which need new training signals or system designs. S-Space helps separate failures of perception from failures of computation, pointing toward systems that connect learned spatial maps with reliable reasoning and richer interaction with the physical world.
New MirroS research explained: S-Space, a spatial workspace in multimodal models. ๐งญ
When you look around a room, you keep track of where things are and can imagine the same scene from another viewpoint. We found evidence of a similar internal map in multimodal AI, which we can read, manipulate, and track as itchanges during reasoning. ๐ง
This gives us a new lens into spatial intelligence inside: how such representations emerge through training, how reliably models use them during inference, and how reasoning reshapes them at test time. ๐ฌ
Explore S-Space:
โข๐ Interactive blog: https://t.co/7fa7J7qxvv
โข๐ Technical report: https://t.co/spxsPySJpg
โข๐ป Code for ALL experiments: https://t.co/wnkUJ5BM1e
Models also struggle to manipulate the map accurately. During chain-of-thought, we can watch S-Space gradually rotate toward an imagined viewpoint, but the transformation remains imperfect. On SpinBench perspective-taking tasks, Qwen3.6-27B reaches 83.56% accuracy with chain-of-thought. Reading its S-Space coordinates and rotating them with code reaches 95.21%. Its internal representation supports better answers than its own reasoning produces.
What does it mean for an AI to understand space? ๐
Not just to recognize what is in front of it, but to build an internal world: where things are, how they relate, and how that world changes when the viewpoint changes.
We found evidence that multimodal models have already begun to build such a world inside themselves.
We call it S-Space ๐ง โ an internal spatial workspace that can be read, manipulated, and transformed during reasoning.
A glimpse into how AI may learn not just to see the physical world, but to reconstruct and reason within it. โจ
Explore S-Space, new research from MirroS ๐ https://t.co/7fa7J7qxvv
HarnessEval-W Public Set is now live ๐
Weโve released the Public Set, updated the leaderboard with the latest model results, and opened community contributions โ including new evaluation cases & taxonomies, evaluator skills, evaluation models, and more.
๐ Project Page: https://t.co/iGvETeP4Yy
๐ป GitHub: https://t.co/1u7Cps5CFF
๐ค Contribute: https://t.co/4IhyoDVviW
Help us build a continuously evolving benchmark for evaluating visual worlds!
#HarnessEvalW #WorldModels #Agents #Evaluation
Today, we introduce HarnessEval, bringing the power of agentic workflows to the benchmarking community! ๐
With Harness, a benchmark is no longer a static rubric. It becomes an intelligent agent that proactively interprets context, decomposes high-level evaluation problems into sub-problems, assembles the right tools, and spawns sub-agents to uncover exactly why a model fails. By mimicking human evaluation workflows and breaking down complex questions into solvable sub-tasks, HarnessEval moves beyond the traditional Q&A probing in existing benchmarks, and provides fully verifiable reasoning traces for every final score. It evolves evaluation into a dynamic, executable agentic system.
To realize this vision, we are open-sourcing our first agentic benchmark for visual generative world models: HarnessEval-W. All the harness, skill libraries, evaluation cases, and results are publicly available. We invite the broader community to contribute to this agentic benchmark workflow together! ๐ ๐ค
๐ป Code: https://t.co/lTf6rbGcvS
๐ Leaderboard: https://t.co/iGvETeP4Yy
๐ Blog: https://t.co/gb7YzjWl6G
๐ Technical Report:https://t.co/wHwQUJX5K6
We introduce Code-as-World ๐, a new paradigm for representing the physical world as executable code!
Pixels capture how the world appearsโbut not what exists within it, how it evolves, or the mechanisms that govern its behavior. Code-as-World reconstructs visual observations as executable code ๐ป that describes a worldโs composition, dynamics, and appearance.
Through an agentic discovery loop ๐ , an agent proposes a hypothesis, simulates and renders it, compares the result against observed evidence, and iteratively revises its code. World modeling thus becomes a process of active discovery and verificationโnot one-shot generation.
The resulting executable worlds provide scalable physical supervision, enabling state-of-the-art performance in quantitative physical reasoning ๐ and opening new paths toward physical intelligence.
We are releasing the technical report, code, models, project page, and blog. This is one step toward MirroS's broader goal: expressing the physical world in a language that agents can understand, simulate, and verify. We will keep exploring how code and language can provide a shared foundation for physical RSI. ๐
๐ Blog: https://t.co/3anShWqx9A
๐ Technical Report: https://t.co/4sqVBYd4TZ
๐ Project Page: https://t.co/JlZe70Myyw
๐ป Code & Models: https://t.co/YJzz0FFYUh
Today, we introduce HarnessEval, bringing the power of agentic workflows to the benchmarking community! ๐
With Harness, a benchmark is no longer a static rubric. It becomes an intelligent agent that proactively interprets context, decomposes high-level evaluation problems into sub-problems, assembles the right tools, and spawns sub-agents to uncover exactly why a model fails. By mimicking human evaluation workflows and breaking down complex questions into solvable sub-tasks, HarnessEval moves beyond the traditional Q&A probing in existing benchmarks, and provides fully verifiable reasoning traces for every final score. It evolves evaluation into a dynamic, executable agentic system.
To realize this vision, we are open-sourcing our first agentic benchmark for visual generative world models: HarnessEval-W. All the harness, skill libraries, evaluation cases, and results are publicly available. We invite the broader community to contribute to this agentic benchmark workflow together! ๐ ๐ค
๐ป Code: https://t.co/lTf6rbGcvS
๐ Leaderboard: https://t.co/iGvETeP4Yy
๐ Blog: https://t.co/gb7YzjWl6G
๐ Technical Report:https://t.co/wHwQUJX5K6
Building Physical RSI Beyond the Known World
Intelligence is not about optimizing within a closed worldโit is about transforming every surprise into the surge of its next evolution.