This robot works in real time locally running multiple AI models on apple MLX. Whisper for listening, smolvlm2 for image description (you can use Gemma for this but it adds so much latency), Gemma 3 4b for logic and autonomous action! Let me know if you want to check out the repo
🪄 Bring your 3D characters to life with words!
Introducing UniMate: one unified model for text-driven animation across diverse skeletons—from humans and animals to articulated objects.
🎬 A rigged asset + a text prompt → motion.
No per-skeleton retraining needed!
🌐 https://t.co/4dkns5mr88
Announcing d1, our first decision model.
It's the first model to outperform Jev on @huggingface's Decision Index.
> wins on multilingual evals
> more robust against prompt injection
> handles longer inputs more effectively
> built for fast, structured decision-making in software environments
> Liquid API: https://t.co/Lp1Qf15Qcc
> Available on OpenRouter soon
Holy crap, this is wild. Just found RelateAnything, a new open-vocabulary computer vision model that figures out exactly how objects in an image interact like holding, wearing, or sitting on using just raw pixels and bounding boxes.
The coolest part is that it is insanely fast. The entire model is only 53 million parameters and clocks in at around 20ms per frame, making it completely viable for real-time video streaming or robotics. It already understands over 19,000 relation words right out of the box with zero retraining required. Plus, it looks purely at the visual context and geometry instead of cheating by guessing relationships based on object names.
Spatial relations are still a bit of a weak spot for it right now, but the potential here for live video feeds is massive.
I think this is the most interesting new AI release in a while and I hope to see more models like this.
Lightning fast, reliable, smart, and cheap. (The output tokens are FREE)
Within a day of its release, I've seen crazy demos of what this model is capable of
- The fastest AI browser use I've ever seen
- A virtual car driving itself in realtime
- This doom demo
Its a completely different type of model than an LLM, though it's still promptable. This model, Jev, ONLY responds in simple narrowly defined outputs and confidence scores, compared to the chat models we're all used to. Yet, its able to power all these cool use cases because computers work best with narrowly defined outputs.
So do much of our AI automation pipelines. Let me give you a familiar example, RAG. Even though a human is in the loop here, there are many places you can give tokens to a model like Jev where it would improve the process' cost and speed.
In this example, we have a retrieval that could use some work, we get results that don't match the question, that need follow up processes for certain industries and for our output different types of approaches for translating languages. We can define how it answers.
Is this relevant to the asked question? (Yes) or No
Is this about x y or z industry? x or (y) or z or No
Language? French or German or (English) or Spanish
... and you can continue to define the data received as long as you want with minimal costs.
Very powerful! But it doesn't replace the LLM's we use today. You'd still need the LLM to relay the information after its been massaged, but we've relieved that model of all the definitions above - saving tokens and time.
I hope to see many more models built like this, and can see many use cases in the future for this including robotics. Would love to see these multimodal. I'm keeping an eye on the Typesafe team.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution
A fun experiment I had playing with the cutting edge in AI 3d model generation.
In Amsterdam, I saw this beautiful decorative bowl in the Rijksmuseum that was once owned by a countess. I wanted it. I thought once the technology was good enough, it would be mine, from a photo.
I was surprised to find out the technology is kinda already that good!
Using meshy ai (use this link if you want to try pro for just 1$) I used the photo as reference and got a near perfect 3d model of the bowl. I since then made some alterations so it would be easier to print, the wings became structural, and the 4 legs were removed as plastic is weaker than silver and gold.
I'm super happy with the results and even using this kinda feels like magic. I've played with image to 3d in the past and the results were cool but mostly funny. I attached a photo of a result I got a year ago with a tencent ai model for comparison.
You could vibe code in group calls! Plug this into a super fast reliable model like GPT 5.6 Luna and you can make cool demos in a call with your friends
Muse Voice Transcribe is MSL's first real-time audio perception model -- rolling out today. SOTA in streaming speech-to-text, it handles speaker diarization, and endpointing natively in a single model.
Muse Voice Transcribe is MSL's first real-time audio perception model -- rolling out today. SOTA in streaming speech-to-text, it handles speaker diarization, and endpointing natively in a single model.
There is a child with a rare disease who is currently suffering and struggling to manage his symptoms. Rare as this is, you can directly help him.
Today we are launching the "Rare Disease, Real Kid" Hackathon, and there are $50,000 in prizes from @AnthropicAI and @awscloud.
We (@huggingface & @Sagebio) are helping this child open his genome and clinical data to the community, so that we can find what's caused his disease and what currently-approved drugs could help him.
I doubt I need to motivate this much further or explain how rare it is for a family to share their child's genome and clinical data, but if you're not sure, consider this:
Until very recently, it wasn't feasible for patients like this to get treatment because their disease was so rare that the economics could never justify the investment. Now, as we've seen, people with rare diseases are starting to be able to find the answers themselves (with the help of AI tools, cheaper sequencing, etc). This kid is not able to do that for himself and neither are his parents, so we're asking you for help. Both for this kid and to prove that it's possible for everyone else suffering from a rare disease.
More details in 🧵.
https://t.co/hjjagALtMu
This is a super cool website.
You can scroll around and look at any country's population pyramid, fertility rate, and mortality rate by year.
Check this out:
Flock takes their privacy very seriously.
But the ethical gray area in surveillance technology works both ways.
Under no circumstances should elected officials or public servants enjoy an expectation of privacy when spending tax dollars to attend a surveillance lobbying event.