If what you say is true, the VISION and the LANGUAGE on-device could be dangerous
First came AIs that could see. Like the AI in your self-driving-ish car and your phone’s face ID. But these vision-only models had no ability to use language or think.
Language and “thinking” ability came with the LLM explosion. But, at first, language models couldn’t see.
Then vision and language merged. Your AI has been able to see, think about what it sees, AND use language for at least a couple of years.
These “vision-language models” are typically huge models in the cloud running on massive GPUs.
But what about the many cases when you can’t, or don’t want to, use AI in the cloud?
At @SuperfocusAI, we’re building on-device AI systems with VLMs for these situations.
Our demo video below uses the case of AI on a drone to show what is possible.
The challenge of on-device AI is that the chips used in most devices have only a tiny fraction of the firepower that cloud GPUs have. So you can only use a very small VLM. But you still have to pack all of the intelligence that you need into this small VLM.
Imagine a scenario where the drone has to deliver supplies to a “medic truck”.
The video shows what a vision-only model (top) would do in this scenario versus a VLM (bottom).
Both models are running together live on a chip that is commonly used in drones. We can also do this on even cheaper, lower firepower chips.
The vision-only model does a good job of identifying trucks. But it can’t think about what it is seeing, so it can’t distinguish between the trucks without further training. It is impractical to train in advance for every nuance that might arise in the moment.
In comparison, the VLM just needs to be told that the medic truck is the one with the red cross on it. With that information alone, the VLM reasons about what it is seeing and correctly picks out the medic truck from the group of similar looking trucks.
We’re using vision and language together in on-device AI systems to add contextual intelligence like this in all sorts of use cases. More examples of what we’ve built to come in future posts…
En garde!
@hsu_steve@reindsummit
Two lesser-noted points in the Jan '25 DeepSeek paper caught our attn at Superfocus dot ai.
1) They produced AI models small enough to run directly on consumer electronics devices but with surprisingly strong reasoning ability. 2) They did it with relatively fewer data samples.
We wanted to know if these smarter “on-device” size models are actually good enough to be useful? And how cheap of a chip could we run these models on?
Why does this matter? There are lots of situations where you can’t, or wouldn’t want to, use AI in the cloud via the internet.
Personal & family privacy (ex: your home). Child safety. Security. Data/IP privacy (factories). No or unreliable signal (commercial/defense drones). Low latency (robot/vehicle safety). No or outdated physical/cloud infra (industrial, warehouses).
Plus, you don’t have to pay compute costs for on-device AI!
We’re sussing out many of these use cases, but because of our shared interest in education, my co-Founder @hsu_steve & I immediately thought about kids & learning.
Below is a video of my daughter demo’ing our prototype AI reading buddy. We’ve rigged our little AI puck up to a pair of smart glasses from our homies at @MentraGlass . (h/t @caydengineer )
This may not be a big deal for huge models in the cloud running on GPUs. What’s significant is that we’re doing it with small models and completely on-device. No internet connection.
The AI sees what my daughter sees. It listens to her and follows along as she reads. And helps her when she needs it by speaking to her. All from small AI models that we finetuned, engineered and got running fast and accurately on a chip that costs less than $25.
It’s not quite a reading tutor yet, but perhaps you can see how it could become one? Or how on-device AI like this could solve other problems?
Would love to know what you think, good, bad, or other!
It ain’t no fun if the Americans can’t make none.
There is not a single low-cost American AI chip for running genAI models on-device.
That’s one of the eye-opening things we’ve learned at Superfocus dot ai investigating on-device AI, and a point I made in my @EmbVisionSummit talk last week in the Valley.
These are the chips in phones, tablets, smart glasses, wearables, drones, robots etc.
The cheapest American on-device AI chips cost $100+ at volume, likely more. The lowest NVIDIA GPU chip and the top Qualcomm smartphone chips.
A $100 chip means a ~$500 product. At least.
Chinese and Taiwanese on-device AI chips that can do the exact same thing cost $25, maybe less, at volume.
A $25 chip means a ~$125 product.
There are lots of use cases where you can’t, or wouldn’t want to, use AI in the cloud.
For those situations, why isn’t there a low-cost American on-device AI chip?
Does it matter? The cheaper options from China and Taiwan are good (but not perfect) chips from solid companies.
@edgeaivision@phil_lapsley@JeffBierBDTI@Rakshit_Ag@vghadiok@daveselinger@hsu_steve #EVS26