There's a lot of excitement around OCR capabilities in multimodal models. DeepSeek-OCR, for example has been doing great - but I think most developers might not know the short-comings of such models.
If you threw in a cropped image of the US constitution (Article I) - it will confidently hallucinate some passages from either sections (not visible in the image) or other articles.
For documents as canonical as the US Constitution, the model has an extremely strong next-token prior. If the vision encoder’s signal is weak/ambiguous (blur, glare, small font, compression), the decoder can minimize loss by producing the “most likely” continuation of the Constitution rather than the literal text on the page.
Try It Yourself → https://t.co/cOhWr3rQJh
BREAKING DeepSeek just let the world know they make $200M/yr at 500%+ profit margin.
Revenue (/day): $562k
Cost (/day): $87k
Revenue (/yr): ~$205M
This is all while charging $2.19/M tokens on R1, ~25x less than OpenAI o1.
If this was in the US, this would be a >$10B company.