pretty easy to get opus 5 to not suck.
just tell it that before any response, it must say "I'm going to respond in a short concise plain spoken simple way.".
think of it like image in-painting. does nerf it a bit though.
scaling laws do not apply to physical AI yet. i'll say it again. scaling laws do not apply to physical AI.
chelsea finn from physical intelligence showed this during her YC talk. they trained pi0.7 to make coffee. the robot couldn't learn from video alone. failed over and over. an engineer had to physically grab the robot and correct it multiple times just to get it to catch the coffee grouper with one hand and balance with the other.
peter florence at generalist, same story. VLA models, the robot figured out hand usage from video datasets but still needed a human to step in and fix the execution.
human-in-the-loop training hasn't changed in robotics. and honestly it's not changing anytime soon.
there's a massive gap between a robot picking up an object on a table and a robot executing a real skill that humans learned over decades. a skill is a series of tasks combined to produce a specific outcome. you don't learn that from watching videos.
VLAs are not the answer by themselves. you need to deploy robots for robots to get better. actual reps in actual environments where things break constantly. everybody trying to skip deployment with more compute and more video data is going to hit the exact same wall.
This is exactly what we fix at @Cosmicbrainai , we are focused on deployment first.
ChatGPT is the wrong analogy. ChatGPT was a “moment” because laptops and the internet meant unlimited distribution - we all got to experience it on the same day.
I’m waiting for the Apple II moment in robotics - an affordable (<$10k) personal robot I can train to do things.
@signulll@theojaffee you really think that wasn't the first thing they considered doing? that's a bribe to sell out your community.
the actual simple way is to show how it's useful to people today. if you can't show them because it isn't, go back and build better product till it is.
@skalskip92@Mascobot static scene computer vision is pretty much there.
DinoV3/mast3r get you semantic/geometric respectively.
Dynamic scene is by comparison far behind.
@antopatrex1 I agree that it's stupid how much we celebrate it.
but you have the analogy wrong. it's not like borrowing money.
it's a celebration of a sale. the thing you're selling is a piece of your company.
@sincethestudy chores are boring. let's stop boxing ourselves into Rosie the robot. let's be a lil more creative
imagine waving your hand and your small army of robots builds you your dream flying car
robots are tools, not appliances
@Scobleizer any general purpose photo to action model will contain an implicit world model. there's a paper that literally derives this in pure math.
the whole point of the vlm/video/world model being the base model is because we don't have enough data to train pure robot models.
@ZheningHuang Insider using sam3d for the static object recon. Works remarkably well for complex objects. Plus there’s extensions that accept multi view