@dhh The other model needs the original request too. Otherwise it can approve perfectly tidy code that solves the wrong problem. I'd let it reach its own verdict before showing it the writer's explanation.
@DrJimFan Cleaning creates new problems as you go: a missed grasp can turn a drink into a spill. I'd love to see more demos where the robot recovers and finishes the job without a human resetting the scene.
@GoogleLabs I'd try two versions of the same maze: one with a timer, one where you can only see a few steps ahead. A playable prototype lets you compare the tension each rule creates before committing to the rest of the game.
Scale's AgentEnv lets you build a simulated workplace, give an agent tasks and check what it actually did. You can swap agents while keeping the tools and data fixed, making it much easier to compare approaches before connecting one to a real workflow.
PhotoCraft is an open-source image editor with layers, masks and PSD support. Its editing tools also work through an API, so an agent can edit a file and you can keep working on it by hand. Still early alpha, but that shared workflow is worth watching.
@udiWertheimer I'd rerun a task that used to need lots of hand-holding with just the goal and the same tests. If it works, that tells you which parts of the old workflow you can drop. Otherwise it's hard to tell what the new model changed.
@lagz152507 Edge Impulse has a nice demo of a tiny model spotting unusual motor vibrations on a microcontroller. You could send alerts instead of streaming raw readings all day, while detection still works offline.
https://t.co/LROF8xKdYu
Google’s new model puts text, images, audio and video in one embedding space. Once indexed, you can search across them with words or an image. This could make apps that search your photos, recordings and documents together much easier to build
Introducing EmbeddingGemma 2! 🚀
Our lightweight, multimodal embedding model maps text, code, images, video, and audio into a single, unified embedding space. Optimized for on-device use cases, it features:
- 740M parameter form factor with modular encoders
- Flexible dimension sizes (768dim-128dim) via Matryoshka Representation Learning (MRL)
- 8K context window (4x larger than text-only EmbeddingGemma)
- A commercially permissive Apache 2.0 license
Whenever I’m working on something, I try to ask:
What kind of return does this actually have?
If the upside is capped, do less of it.
If it’s linear, find someone else to do it.
If it compounds, that’s probably where I should spend my time.
@googlegemma Nice detail in the docs: you can index media with the full model, then use just the 270M text setup for typed searches. They share one embedding space, so a local search app doesn't need the audio and vision encoders sitting in memory whenever you search.
@jans I like that the work comes back to the task. “Reply to this email” can become a draft waiting for review, without another chat to keep track of. For busy people, being able to leave and come back to something reviewable feels like the real win.
@ElevenLabsDevs For a voice tutor, I'd stream the LLM's text straight into this. The student can hear the first part while the rest is still being generated, so you don't pay the full text-generation wait and then the speech-generation wait on every turn.
Qwen Image-2.1
Official repo: 90K downloads
Uncensored version: 1.6M downloads…
Anyone here actually running the uncensored one?
Drop some of your “research results”
I just want to know what the hell all 1.6 million of you are working on 😆
I’ll stay up all night working on the product.
Growth? I’ll put it off for three days.
Doesn’t matter how good the product is if nobody uses it.
I know. I just don’t want to do the growth part.
@ATinyGreenCell For lab fixtures, keeping tube diameter and spacing as parameters is the useful part. A new tube size becomes a small edit to a working design, rather than another CAD project. That makes one-off tools much easier to justify.
Dust turns each token into a tiny training experiment by nudging its activations. One forward pass tests thousands of nudges together. Where you put the noise changes how much search you can afford. Still far costlier than backprop.
https://t.co/tx1iCpxV6Z
@sharonal_lee The plug example makes this click. If it can slide through the socket wall in sim, the robot can get rewarded for a move that fails in reality. The collision model is part of what you're teaching it.
@Charles_Y_Wu The clever bit is giving different parts of a prediction their own space, so they don't all learn the same easy pattern. For a world model predicting many steps ahead, losing a small signal early can throw off the whole rollout.
AMADEUS uses isolated animals in the video to generate synthetic overlaps for training. The easy-to-track moments help teach it the crowded ones. For animal-behavior researchers, that could cut the manual labeling needed to study social interactions.
Cool direction for hardware builders: explore packaging while you're still changing the product. This prototype keeps cavities, clearances and sealing areas editable around the object. Linking those decisions early could save a second redesign when it's time to ship.
Day 31 building the Figma and Amazon for hardware until somebody invests in it.
[packaging - blister]
So I’m testing the same hardware kit across vacuum skin, blister and pouch packaging. The film, cavity, clearances, sealing area and materials all adapt around the object, while staying editable.
I’m also experimenting with the forming process itself, so you can move from the flat film to the final package and understand how the object changes the structure around it.
The goal is that eventually you don’t stop at designing the product. You can keep going into how it’s protected, presented and prepared for the physical world.