Introducing TabFM, a foundation model designed specifically for tabular data classification & regression. This approach allows generation of high-quality predictions on previously unseen tables in a single forward pass.
Learn more and try out the model →https://t.co/OTbVQ8oUQs
Here’s an early sneak peak of OpenResearch, our brand new feature for reproducing and experimenting on top of papers
We put together a template so you can train VPO on ToolRL in one click on a single gpu
Vector Policy Optimization trains models to generate diverse answer sets by randomly weighting reward dimensions so each answer specializes in a different tradeoff
The result is better test-time search as the sample budget grows. Check it out below!
📢(1/11)Diffusion LMs are fast and controllable at inference time! But why restrict such benefits for processing text data?
We are excited to announce LaViDa, one of the first and fastest large diffusion LM for vision-language understanding!!
We present MMaDA, first diffusion that unifies text reasoning, multimodal understanding, and image generation through Mixed Long-CoT, and unified RL - UniGRPO.
📚 Paper: https://t.co/O9i9Exsgdt
💻 Code: https://t.co/hltAWYaw2d
📦 Model: https://t.co/EFNwgnBRBL
Excited to announce GR00T N1, the world’s first open foundation model for humanoid robots! We are on a mission to democratize Physical AI. The power of general robot brain, in the palm of your hand - with only 2B parameters, N1 learns from the most diverse physical action dataset ever compiled and punches above its weight:
- Real humanoid teleoperation data.
- Large-scale simulation data: we are open-sourcing 300K+ trajectories!
- Neural trajectories: we apply SOTA video generation models to “hallucinate” new synthetic data that features accurate physics in pixels. Using Jensen’s words, “systematically infinite data”!
- Latent actions: we develop novel algorithms to extract action tokens from in-the-wild human videos and neural generated videos.
GR00T N1 is a single end-to-end neural net, from photons to actions:
- Vision-Language Model (System 2) that interprets the physical world through vision and language instructions, enabling robots to reason about their environment and instructions, and plan the right actions.
- Diffusion Transformer (System 1) that “renders” smooth and precise motor actions at 120 Hz, executing the latent plan made by System 2.
We deploy N1 on GR1 robot, 1X Neo robot, and a large collection of simulation benchmarks. N1 achieves up to +30% boost in diverse manipulation tasks for household and industrial settings.
While humanoid robots are the main focus of N1, our model also supports cross-embodiment. We finetune it to work on the $110 HuggingFace LeRobot SO100 robot arm! Open robot brain runs on open hardware. Sounds just right.
Let’s solve robotics, together, one token at a time.
Links to our Whitepaper, Github repo, HuggingFace model, and open dataset page in the thread: 🧵