Check out our demos using LFM2.5-VL-3B, our latest lightweight, vision-language model that reads screens, documents, and the physical world.
First up: LFM2.5-VL-3B running fully on-device in the browser with WebGPU to understand a document page. The model parses the entire layout in one pass and returns regions and labels that the interface renders as an overlay. The demo highlights OCR and layout understanding for visually structured content such as forms, reports, receipts, and other documents.
🧵
really stoked about the new/experimental document parsing with layout: you get bounding boxes and semantic classes for each OCR'ed snippet of text. Shoutout to @rshubert98 for going above and beyond to bake this capability in during midtraining
LFM2.5-VL-3B reading a handwritten Hideki Yukawa manuscript and a letter from Santiago Ramón y Cajal.
Historical manuscripts are a tough test for vision models, especially with handwriting, equations and old document layouts.
Pretty damn good for a 3B model.
Solid release from @liquidai
really stoked about the new/experimental document parsing with layout: you get bounding boxes and semantic classes for each OCR'ed snippet of text. Shoutout to @rshubert98 for going above and beyond to bake this capability in during midtraining
Today, we release LFM2.5-VL-3B, a lightweight vision-language model that reads screens, documents, and the physical world. It handles digital screens across mobile, web, and desktop, grounds objects to coordinates, reads text and charts, and calls tools from either text or image input.
Built on LFM2.5-2.6B base, with a SigLIP2 400M NaFlex vision encoder
> Pre-trained on ~34T tokens
> Vocab size: 128K
Comparable or better scores compared to models up to 2.6x its size:
> ScreenSpot-v2 80.7, ahead of Gemma-4-E4B at 51.2
> RealWorldQA 73.1, ahead of InternVL-3.5-4B at 67.7
> TextVQA 84.3, ahead of Qwen3.5-4B at 81.2
> RefCOCO-avg 87.9, up from 57.1 on LFM2-VL-3B
> ToolSandbox 59.5, up from 26.4 on LFM2-VL-3B
🧵
In a new article published today in Science Robotics @SciRobotics, we showed a class of neural nets can understand the task they are given (flight navigation)!
These are AI systems that WE understand what they do and THEY understand what they do! (1/n)
https://t.co/TWMvDHHTmR
Future automated driving tech—hands-free, 'eyes-off' highway driving—has potential to redefine our relationship with our vehicles. Imagine getting hours of your life back to be more productive or reduce stress. Life changing! We've established Latitude AI to help us get there.
World models, intuitive physics, planning, problem solving, discrete search for solutions, continuous optimization of control parameters...
Dogs manage to do all this with 2 billion neurons.
Why debate human-level AI when we can't approach dog-level intelligence yet?
1/N
As @stephenfry has very eloquently noted, the most damaging and horrible human emotion is self-pity. It's a veritable black hole that is extremely hard to pull oneself out of and that consumes nearly all resources poured into it. It cannot be satisfied.
Statements like just highlight how naive the “woke” crowd tends to be. You have indeed identified a problem in society; medicines are often inaccesable due to lofty prices. However, your only “proof” that this is a symptom of capitalism is based on the reaction of a hive-
Carl Sagan: "extraordinary claims require extraordinary evidence."
No, NASA didn't find evidence of a parallel universe where time runs backward https://t.co/AtJ8pWZepm via @CNET (a good overview of the science and the @newscientist piece that started the frenzy)