JEPA are finally easy to train end-to-end without any tricks!
Excited to introduce LeWorldModel: a stable, end-to-end JEPA that learns world models directly from pixels, no heuristics.
15M params, 1 GPU, and full planning <1 second.
๐: https://t.co/cpTzgvbTS0
@Junwon__Seo So to sum it up, train on the IMPORTANT scenarios instead of the average of all scenarios? Which all else being equal is probably not that important?
@Deeplearner2@chris_j_paxton Solid proposition but what about all the people who like folding laundry ๐
I'm personally too new to robotics to justify this expense but definitely down the line I'd love to experiment.
A conversation with Claude: discovering world models from first principles ๐
0:00:00 video is redundant
0:04:13 why language is special and why pixels break
0:13:13 world models (V-JEPA)
0:17:18 LLMs as a dead end?
0:24:36 human-scale vs particle-level
0:31:02 benchmarks
0:34:20 what it unlocks
0:37:27 robotics & multi-agent finale 38:27
@DanielMika_@dariusfdi What do you think are the criteria/conditions needed for the right model to scale? What is bottleneck that needs to be solved for?
E.g. for LLMs it was moving away from sequentially looking at words
RL is reinforcement learning?
So is the stage:
small data size, high efficiency via effective reinforcement
large data size, throw efficiency out the window to just brute force skill learning
If we go back to human sample efficiency, we have a mix of ingredients that robots don't have:
- other fully developed minds and skillsets to observe and learn from (parents)
- embedded genetics for certain skillsets
- im sure there are other things but I can't think of them rn
So should we be trying to replicate this set up for robots?
@0xDeFiTH@teortaxesTex Read about that yeah. I see where he is coming from on that point though, even if it is contra majority. Not sure that counts as cooky?
@_wenlixiao@junyi42 Hmmmm wasn't this one of the central premise of the 2018 paper, "World Models" by David Ha and Jรผrgen Schmidhuber?
How does it differ if not?
@francoisfleuret An LLM isn't failing in understanding the human world.
It just never was built to actually live in it.
Everything it ingests is text. So how can you expect it to understand anything else?