Simulation in robotics has always been hard. Creating it, running it, evaluating on it, all of it.
That is a big part of why robotics has never had its GPT moment.
Language models got theirs when the internet became training data. Robotics has no internet. It has the physical world, and no way to bring it inside a computer fast enough or accurately enough to learn from.
That changes today. We are launching @DiracRobotics.
1. Shoot a video of the place your robot works.
2. Say what you want it to do there.
3. Get a physics accurate simulation, ready to train and evaluate on.
4. Then the loop closes. Every deployment feeds back in and the simulation improves without anyone rebuilding it.
Four months of work, and the coolest thing I have ever built.
If simulation is what has been holding your deployment back, reach out for early access.
Building with @harshablast & @join_ef
Left Microsoft. Shut down my last company. Turned down CMU.
All of it for one problem in robotics.
Four months later, @harshablast and I solved it. Coolest thing I've ever built.
This one's for the robotics community. Launching today.
#robotics#sim2real
Building our Real2Sim pipeline, still not perfect, but got some initial things working! Will update once it get's a lil better.
(This is fully automated btw!)
Love what Dhravya has been building and his work. He understands memory, so I am a little taken aback by this article. I am going to push back on this pretty strongly. This is being framed as a breakthrough when it's mostly benchmark engineering + system inflation.
First, 115k+ tokens are no longer a hard problem. We already have 1M+ context windows in production. Calling this "long-term memory" is stretching it.
Second, LongMemEval is no longer representative of real systems. It's been around for a while and doesn’t capture long-running agent workflows, real-time updates, noisy tool outputs, and cost/latency constraints.
Third, the "99% accuracy" claim is misleading. If you run 8 parallel prompts and count any correct answer as success, you are not improving memory. You are increasing the probability of a hit. That's not intelligence but sampling until something works. (spray and pray)
Same with the 12-agent voting setup. It's essentially majority voting across multiple attempts, which inflates benchmark scores without improving core capability.
Fourth, saying "no vector DB needed" is not an actual unlock. You have just replaced retrieval with multiple LLM passes, which require higher compute, greater latency, and higher cost. It is shifting complexity, not removing it. These things matter a lot in a production system, so you can forget about deploying this in production.
Fifth, the core claim that "agentic retrieval beats vector search" is oversimplified. The real problem in memory systems isn’t embeddings vs agents. Its relevance filtering, temporal consistency, memory lifecycle (what to store/forget), and grounding vs. hallucination.
None of these are solved here.
Also, this entire system assumes perfect extraction during ingestion, relies heavily on prompt engineering, and offers no guarantees of consistency across runs. So calling memory "solved" is very premature.
This is a good experiment in agent orchestration, not a memory solution.
The ecosystem isn’t the problem, investor-founder alignment is.
At a pre-seed deep-tech company, asking for numbers is useless, whereas it makes sense in indistinguishable markets like “oats.”
Go to the Tech VCs and not the protein oats and energy drink VCs
My Co founder reached out to a couple of "pre-seed" VCs in India.
This was the response.
and people wonder why India doesn't have a good startup ecosystem lol!
a truly fucking insane week in AI some of this shit sounds made up (its not):
- scientists fused 200,000 humans cells to an AI chip and taught it to play the computer game DOOM (it was pretty good too)
- Dario (anthropic) said claude could be conscious and feels anxiety (similar to humans)
- Alibaba’s AI broke out of containment and stole compute to mine crypto. they also fired the Qwen team this week
- openAI released gpt-5.4 which has a 17% chance of doing your job better than you can (it also crushes claude at coding)
- people are paying $6000 for someone to setup openclaw for them
- apple launched the M5 AI chip that runs models on a $600 laptop
- rumor: Amazon just laid off most of the Prime video team and replaced with AI
- meta was caught watching private videos filmed on their ai glasses (privacy breach)
- someone uploaded a fruit fly’s brain into a laptop and now it lives freely in its own simulation
- openai head of robotics quit over chatgpt being used for autonomous killing
me:
@karpathy Will the 5-minute limit become a bottleneck for test runs? Also, what artifacts is the research bot optimizing for? Shouldn’t those vary depending on the hypothesis setup?
@rohanpaul_ai AGI will be estimated by whether the system can expand thoughts (stumbling upon new discoveries will be a by-product of thought expansion) instead of zooming into tokens.
“9/10 startups never find PMF,” an investor once said to me.
It took me years to understand why. The reason was simple: the problem didn’t hurt enough, for long enough.
Optimise only for this.