Over the past few weeks, GPT-6 Astra, Claude Code with Opus 5.5, and other recent foundation models have surprised many of us in graphics, CAD, and robotics. They can now work directly with Blender, CAD kernels, game engines, physics simulators, and real robots very well. New examples appear almost daily, faster than the publication cycle can capture.
📝To keep track, we maintain a Living Survey of frontier AI for Design & 3D Modeling & Robotics:
🌐 https://t.co/irJ9uUvdh2
Currently, it organizes 243 cases from 345+ public showcases and developer reports across 3D modeling, CAD and industrial design, animation, simulation, and robot control, each linked to its original source.
(If you are working on related projects, it's better to check this out: )
Different from a traditional survey, three ideas guide it:
1. Horizon scanning beyond publication lag
In an era where frontier foundation model capabilities evolve rapidly, traditional academic publishing cycles could lag behind public community developments and empirical findings. This platform establishes a centralized, high-velocity empirical synthesis repository to provide researchers and engineers with timely visibility into ongoing developments across 3D generation, parametric CAD, and embodied robotics—anchoring these observations in objective evaluations of strategic opportunities and critical safety boundaries.
2. Demonstrations as a distributed record of use
We analyze the corpus of over 345 publicly documented showcases and developer reports as an extensive, distributed "crowdsourced user study." This framing captures how models operate when prompted across diverse geometry kernels (CGM, Open CASCADE), DCC software (Blender), physics simulators (Isaac Sim, MuJoCo, Genesis), and physical robot hardware—revealing real-world workflow friction, prompt overhead, and boundary failures that static benchmarks miss.
3. Ranking by open verifiability
A central challenge is balancing the timely collection of rapidly emerging results with the need for high-quality, evidence-based evaluation. Our approach is to rank cases by open verifiability—the amount of evidence that others can directly inspect.
Equally impressive demos can carry very different evidence, so every case is ranked by what others can inspect:
- Rank 1 · Demo + Implementation Code: public code, scripts or CAD/robot harnesses (availability/verifiability)
- Rank 2 · Demo + Interactive Web Link: a live web app, 3D viewer or cloud CAD project
- Rank 3 · Demonstration Media Only: video or screenshots only
Note: Rank 1 means the evidence is inspectable, not necessarily reproducible from a clean setup.
Gallery: https://t.co/Tgo48HU4dH
Contributions are very welcome via pull requests on GitHub. We especially welcome Rank 1 contributions (demos with code) and Rank 2 contributions (demos with interactive webpages).
Many thanks to our collaborators: Jamison Meindl(@meindljamie), Akihisa Watanabe(@Akihisa_Wat ), Anna Deng, Tianyu Huang(@tyanyuy3125), Igor Sadalski(@igorsadalski ), Harrison Liang, Minghao Guo(@GuoMh14 ), Benjamin Tod Jones, Wojciech Matusik(@wojmatusik ).
Can an LM, starting from random init (!!), learn to generate all of its pretraining data?
Introducing Self-Play Pretraining with Zero Data. Two models start from random initialization: a generator proposes programs for a universal Turing machine and a learner trains on their outputs. We never train on any real data, but see predictable scaling on natural datasets: zero-shot val loss on images, text, audio, and melodies decreases predictably with self-play compute. And the learner develops in-context learning capabilities.
A fun proof-of-concept, co-led with @AdityaCowsik and @KfirDolev and co-authors @gbruno_dl, @ANourya@noahdgoodman, and @YoavLevine.
Introducing Contrastive Language Model (CLM): an ultra-fast System One Model trained with a contrastive learning objective that connects states and actions.
CLM-8B is pre-trained on internet-scale data and delivers up to 9× faster inference than Jev ⚡ while achieving comparable performance across computer-use, gaming, and tool-calling tasks.
With lightweight fine-tuning, CLM-8B sets a new SOTA on challenging agentic coding benchmarks, such as DeepSWE (81.6%) and Terminal-Bench 2.1 (87.6%). In contrast, Jev fails to serve as an effective verifier for these long-horizon tasks.
We also build an efficient training and serving infra for CLMs by disaggregating states and actions, allowing their embeddings to be cached and reused independently. This substantially reduces inference latency in settings where the state evolves continuously while the action set remains fixed.
Finally, we establish scaling laws for CLMs and show that the test contrastive loss decreases predictably as a power law in training compute, model size, and dataset size.
📄 Blog: https://t.co/zwi9JOHKGx
💻 Code: https://t.co/rsHRYCGR8I
🗣️ Discord: https://t.co/Uqtdefvo3J
🤗 Data & Models: https://t.co/wdSWGGO3hu
More details on CLM’s architecture, data recipe, and scaling laws in the thread below 🧵
We built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, and then help us decide what to try next.
Using only 1,300 H200s, plus months of our experimental data, we mid-trained and RL’d an open-source model to surpass GPT-6 Astra on our analysis benchmark. We call it Neon.
This is real footage from our lab. We’re focusing first on hard problems in materials science, including superconductors, magnets, and semiconductor materials.
Read our blog posts below.
New paper: Understanding Reasoning from Pretraining to Post-Training!
We study the full LLM training pipeline from pretraining to post-training, find a joint scaling law, figure out how the compute should be allocated, and study what RL is doing to the policy.
🧵
The first experimental evidence of recursive self-improvement (RSI).
Autoresearching the autoresearch agent for eight days.
The result beats the harness we hand-tuned for two years, on held-out benchmarks: 🧵(1/7)