Build and ship agentic devices with one AI stack.
Hojo gives edge AI developers everything from Models and Memory to Agent Runtime and Device Access.
Make real devices listen, remember, reason and act.
GitHub: https://t.co/q8H1krjyVr
Hugging Face: https://t.co/g38ngfwz1n
Clean audio is easy. Real conversations aren’t.
People switch languages halfway through a sentence. They change their minds, speak over traffic noise, use bad microphones — and still expect the model to keep up.
So we want to test the messy stuff.
Send us a short, non-sensitive audio sample or tell us about a difficult voice scenario. We’ll test it with Hojo-ASR-Multi-V1 and share what works, what doesn’t, and where the model still needs improvement.
What should we test first?
Model: https://t.co/hDRzYy7Uo5
@KtAIFeed It really is. “Noise” is also broader than it sounds—a steady fan, traffic, music, and another person speaking can affect ASR very differently. We’re thinking of testing the same sentence against several noise types so the results are easier to compare.
Exactly. And we don’t want to collapse all of those into one generic “real-world test.” Code-switching, background noise, and poor microphones can cause very different kinds of errors. We’re planning to test them separately and share the raw outputs, including the cases that don’t work well.
Small voice models get interesting when they leave the demo.
Hojo-TTS-Light-40M runs on CPU and supports both Chinese and English — making local voice practical for the products people actually use.
An e-reader?
A game character?
A local assistant?
A creator tool?
Where would you put a voice?
https://t.co/S81hk1VkXI
We’re introducing Gemini 3.8 Live and 3.8 Live Extended Thinking – our best conversational AI.
The models talk, think, and handle tasks in the background without breaking your flow. 🧵
Deep Research 🤝 Gemini Live.
Use your voice to explore a topic — in depth — with Gemini Live's Deep Research integration.
1. Just ask the @GeminiApp to run a Deep Research report on a topic, then feel free to close the chat, lock your screen, or keep chatting about other things.
2. Gemini will work asynchronously in the background and send you a notification when your full research report is ready.
3. Once it is, you can chat through the results, ask follow-up questions, or refine the details.
Can your ASR keep up with a real conversation?
Real conversations don’t stay in one language — or one version of the sentence.
“Let’s meet at 七点半… actually, make it eight.”
Can an ASR model keep the language switch, the time, and the correction — without cleaning up what you actually said?
We’re putting Hojo-ASR through the kind of code-switching people use every day. Drop us a sentence you want tested.
AI hardware needs better ears.
AI is moving closer to the body — into earbuds, glasses, and other always-on devices.
But ambient AI only works when it can hear real life: traffic, accents, interruptions, and mid-sentence corrections.
AI hardware needs better ears.
That’s where Hojo-ASR comes in — helping intelligent hardware turn speech into reliable input in the situations people actually use it.
What voice scenario should we test next?
📷
Hearing clearly is only the start.
You can hear someone perfectly clearly and still get stuck on a local expression. Turning up the volume won’t help.
The Sichuan dialect test shared by @MinLiBuilds brings up a similar question for ASR: how much can you work out from the sound alone?
Regional speech comes with familiar shortcuts, local vocabulary, and pronunciations that don’t neatly match standard Mandarin. People who speak it every day barely notice. A speech recognition system has to work through all of it.
Sometimes, the rest of the sentence is what helps you work out an unclear word. Knowing how people actually use the language matters.
Hojo-ASR-V1 combines acoustic features with a Qwen3 LLM decoder, bringing language knowledge into the recognition process itself. The audio provides evidence of what was said; linguistic context helps the model make sense of that evidence when the sounds are ambiguous.
The transcript still needs to stay faithful to the recording, including the speaker’s dialect expressions and corrections.
That’s why Sichuan dialect is an interesting test. It asks how well a system handles the way people speak at home, with friends, or halfway through a thought.
For someone building a voice-enabled device, that’s a practical concern. Users shouldn’t need to rehearse a sentence in standard Mandarin just to get it transcribed.
Try Hojo-ASR-V1 with the way people actually speak around you.
https://t.co/uAJeKbsCRJ
A smart device should be easy to talk to.
Your hands are full of groceries. You tell your earbuds, “Remind me to leave at seven thirty tomorrow. Actually, make that eight.”
You keep looking for your keys. You expect the reminder to be set for eight.
That’s what people want from AI hardware: to say something naturally and have the device keep up.
With glasses, it might be asking for directions while walking. With earbuds, setting a timer over the kitchen fan. With a robot, asking it to pick something up, then adding, “Wait, leave it there.”
All of these interactions need a reliable way to get spoken words into the system. ASR provides that input. If it misses a correction or a device name, the model handling the request has less to work with.
We see Hojo-ASR as the ears of AI hardware: a voice input layer that developers can build into glasses, earbuds, and robots. Hojo-ASR-Multi-V1 gives teams a starting point for testing those interactions across languages.
Making the whole experience work also takes good microphones, audio processing, and timely responses. Speech recognition belongs in that conversation from the start of product design.
Because once you’ve got both hands full, “Please repeat that” gets old pretty quickly.
https://t.co/hDRzYy7Uo5