⚡ A lot of representation models still start from a massive generative VLM.
We trained an encoder from scratch instead: native on text and images, multilingual, and built to be fast from the beginning.
I’m super proud to share NeoMME with my friend @tonywu_71. 260M and 800M, no vision tower. We also fine-tuned a visual document retriever on top: NeoMME-Retriever.
🚀 LightOnOCR-3 is out! 🦉
This time we go beyond OCR: text transcription, document layout, image descriptions and chart extraction, all in one model.
We release 3 sizes: 0.8B, 1B and 4B, all under Apache 2.0!
We built 10 live-camera demos for d1-3B, from gesture-controlled games to content moderation, with one forward pass per frame.
In collaboration with @NVIDIARobotics, we also show d1-3B navigating an environment in Isaac Sim, with the model served on a Jetson in a hardware-in-the-loop setup.
5/
Tiny decision models you can run everywhere:
> d1-3B has impeccable performance in text and vision
> d1-omni-600M supports text, vision, and audio!
Try our @huggingface space with interactive games today
New company, new team, same frontier results & same open license
Happy to announce the release of our new embedding model family
More than just an upgrade in performance, the models are now multimodal, multi-vectors, and use a shared embedding space to allow cross-model querying
Most decisions don't need a big model. They need a fast one.
Today we're open sourcing d1-3B, a vision-language decision model, and d1-omni-600M, which takes text, images and audio.
Today we release Open d1: two open-weight multimodal models in our d1 decision model family.
> d1-3B: text + vision
> d1-omni-600M: text + image or text + audio
> Real-time decision making anywhere, from data centers such as @nvidia DGX to RTX workstations to Jetson at the edge.
1/
here you go! our first series of open-weight multimodal decision models in collaboration with @nvidia 🤝✨
I am obsessed with our 600m omni model that truly brings in new possibilities to the field processing images, audio and text.
Fine-tune these bad-boys for your system one applications! 🚀
masterful work by our post-training and applied AI work.
Models: https://t.co/jJnmLOPivB and https://t.co/tKye3GXtvc
Blog post: https://t.co/c1piP7xWVr
Arcade: https://t.co/Qtev6HRel9
I want to see what you build with it!