You can now generate an entire 3blue1brown style video from any research paper with Opus 5.5.
Here’s a 8min video summary of “Regularized Recursive Self Improvement of Agent Harnesses”.
The 90%ile educational YouTuber is fully automated.
How a robot arm is controlled, explained for ML people new to robot learning. First of the explainers from my own speedrun.
Next one will be on ACT. Follow me to catch it. https://t.co/g25ayJDVhe
Astra tip: Ask codex/ ChatGPT work to verify its work as it goes!
Specially for visual tasks like 3D, Front End & Video editing tasks it materially boosts the end artefact quality
Add this to your project agents md
I kept seeing #GPT-6 Astra modelling results on here and got curious enough to build the thing myself: an agentic pipeline that turns a plain RGB video into 3D assets. The living room for robots @MIT_CSAIL.
Input: one handheld phone walkthrough of our lab kitchen (~20s).
Nothing else — no depth sensor, no CAD, no asset library. ~1 day (I actually slept overnight), including human-in-the-loop.
Live: https://t.co/VVon1iWyPf
It is still not perfect — thin and shiny things are still weak, a few objects are drafts, the room shell needs another pass.
The loop, roughly:
- Monocular video → metric scan (ViPE): camera poses + depth. This is the measuring instrument, not the output.
- GPT-6 lists what should exist as separate objects, then open-vocabulary detection + tracking gives per-object masks; each object is fused and measured in metres.
- For every asset, GPT-6 works in its own sandbox with tools: it writes the object as a program in a small Blender DSL (closed solids, PBR colours, hinges/drawers), builds it, renders it over the original video frames, compares it with the scan points in 3D, and iterates.
- A separate GPT-6 session is the verifier — it can render the model on any frame it wants and must justify every complaint with a frame. The modeller never grades its own work.
- A completeness pass renders the whole modelled scene from the video's own cameras, puts it next to the real frames, and says what's still missing. What the detector keeps missing (a row of identical cabinets) gets placed geometrically instead.
- Out comes MJCF/URDF with joints.
Now, adding simulation to this env.
Can a robot acquire real-world dexterous manipulation skills from just human videos?
Meet Video2Sim2Real: full-stack autonomous dexterous skill acquisition from a single RGB-D human manipulation video — without robot data or expert intervention.
Project: https://t.co/gegklCcs5A
Paper: https://t.co/5AcK8NFQAM
这是中文版本,希望相关受害者帮忙宣传。
GPT-6 Astra 模型发布后,大量用户遇到:
Selected model is at capacity. Please try a different model.
Stream disconnected before completion: Our servers are currently overloaded. Please try again later.
还有,GPT-6 Astra模型出现严重降智的问题,从知识库日期以及鹈鹕骑自行车测试(如图)可以直观看出来。
OpenAI对于用户的问题完全置之不理,无论是面对用户发送的海量反馈邮件,还是X上铺天盖地的讨论热度,相关负责人 @thsottiaux@sama 显然都看在眼里,却依然选择保持沉默。同时我们观察到出现问题的大部分是中文用户,那么,OpenAI是否默认会选择性的牺牲中文用户的��益来减少算力支出?即使付费了同样的订阅价格,也并不保证这部分用户的正常使用?
最可恶的是,没有任何邮件提醒和警告,也没有任何账号违规提示。OpenAI并不主动说明降级规则,只是静默降级路由,使用户完全无法正常使用。在模型严重降智之后,使用额度仍然保持锐减。这一系列操作简直就是诈骗!
OpenAI,如果你不想服务这群用户,为什么不直接给别人退款呢?像Anthropic一样。看来OpenAI不仅想赚钱,而且还想减少支出,两头吃。真是比Anthropic的直接种族歧视更恶心🤢。 @OpenAI @ChatGPT
Following the release of the GPT-6 Astra model, a massive number of users have been encountering errors:
Selected model is at capacity. Please try a different model.
Stream disconnected before completion: Our servers are currently overloaded. Please try again later.
On top of that, GPT-6 Astra has suffered severe capability degradation—the model has clearly been dumbed down, as evidenced by its outdated knowledge cutoff and the classic "pelican riding a bicycle" test (see attached image).
OpenAI has completely ignored user feedback. Whether it is the flood of complaint emails or the massive traction on X, those in charge @thsottiaux@sama have undoubtedly seen the outrage, yet they consciously choose to stay silent. Meanwhile, we have noticed that the vast majority of affected users are Chinese. Is OpenAI quietly choosing to sacrifice the interests of Chinese users to cut compute costs? Why are users who pay the exact same subscription fee denied normal access and basic service guarantees?
What makes this even worse is the complete lack of any prior notice, warning, or account violation alert. OpenAI never explicitly discloses its throttling policies; instead, it silently reroutes and downgrades accounts, leaving the service practically unusable. Worse still, even after the model is severely crippled, usage limits continue to plummet. This sequence of actions is nothing short of a scam!
OpenAI, if you no longer wish to serve this user base, why not just issue full refunds like Anthropic does? It seems OpenAI wants to have its cake and eat it too—pocketing subscription revenue while cutting back on service costs. Honestly, this underhanded double-dipping is even more repulsive than Anthropic's blatant discrimination.🤢 @OpenAI
@Michael_J_Black I've paused part of my group's research to evaluate Astra head-to-head on existing benchmarks. This moment demands a rethink: where we keep re-building instead of measuring, and which long-assumed immovable North Star problems have already been shaken loose.
This Fall at CMU we're teaching a new course on AI Agents!
The goal is that you learn how to create a scaffold, build evals, and train an agentic LLM using RL.
We'll try to balance theory and practice, and introduce modern frameworks and best practices.
Gave Astra a 3D scan of my parents old home to make it an editable scene. Then I turned down gravity and opened a portal.
Turns out multimodal LLMs might have crossed the chasm on inverse graphics. Tremendous implications for everything from VFX and AR/VR to robotics.
Process: ~400 2D photos I took in 2019 went through RealityScan and became a photogrammetry mesh. That’s the classical reconstruction step.
Then I dropped that mesh into Blender and used Astra to turn it into a fully editable, cleaned up 3D scene. Asked it to use textures from my photo scan and create procedural 3D shaders for the rest.
It's good but it’s definitely not perfect -- funny thing is some of the mistakes are inherited from the reference mesh RealityScan produced i.e. the bird statue’s legs, contours of the furniture etc.
This is where humans still have much better “spatial autocomplete” and can take an output like this all the way.
My next step will be to see if I can let it use CUA with the photoscan software to interactively reference the source pixel imagery and refine its work in Blender and get even closer to reality.