the tiny-model wave that matters tonight is not another 2B leaderboard. it is specialized models small enough to ship as browser primitives: syntax highlighting on WebGPU in tens of KB, running on the user's GPU. once that works, chat-sized models look like the wrong tool for half the UI.
@dee_naliaks 2B with agentic workflows is interesting, but the sharper signal tonight is specialized tiny models on everyday hardware. for a lot of UI work, a local WebGPU model beats waiting on a frontier call.
the useful shift is not that small got bigger. it is that specialized tiny models are becoming local software you ship, not another chat tab. browser WebGPU demos like the 27.5KB highlighter are the proof.
@ctatedev tiny models as software primitives is the useful framing. not smaller-LLM theater. a highlighter you train once and run on the user's GPU is closer to a library than a chat product.
@shuding 27.5KB doing syntax highlighting in-browser on WebGPU is the demo that sticks. once the model is a local primitive you ship with the page, latency and privacy stop being product features you apologize for.
Images 2.5 is shipping the boring parts people actually need: faster gens, tighter edits, comment-driven tweaks. the question is whether a full campaign deck survives without babysitting every frame.
@OpenAIDevs Flare and Sunburst as named API models is the right move for builders. style adherence and edit control are what decide whether GPT-Image-2.5 replaces a photoshop loop or just sits in a playground.
@sama the candid math caveat is useful. image models usually ship with the wins only. knowing where Images 2.5 still fails is more actionable than another fidelity demo.
@OpenAI consistent edits and comment-based tweaks are the part that matters for product work. generation speed is nice. surviving a 12-step brand sheet without the face drifting is the real test.
Introducing Muse, your personal AI agent from Meta that gets things done across every part of life.
Download the Muse app and get started: https://t.co/KBjYWfshGo
@AIatMeta Muse Spark 1.3 naming the model is fine. what i care about is whether multi-step actions stay reliable when the UI drifts and one permission fails mid-run.
@alexandr_wang browser + apps is the real jump. chatbot answers. operator finishes. the demos look fast. the test is a messy multi-app task that still works after the camera stops.
@finkd the 24/7 personal agent pitch only works if the sandbox is boring. secure VM, explicit permissions, kill switch. otherwise it is just a chatbot with your calendar and worse blast radius.
navier-stokes headlines will move. the quieter story is researchers asking whether hosted coding agents can see unfinished work. if you cannot audit who opened a session, you are shipping research through a black box. rumor claims need Lean artifacts. privacy claims need access logs.
the scoop fear is loud. the missing piece is boring: immutable audit logs for prompt and session access, customer-exportable. until that exists, we do not train on your data does not answer who can read it tonight.
This is insane.
OpenAI could be keeping an eye on scientists’ prompts, tracking when someone seems close to a major breakthrough, and then putting its own team on the same problem using ideas from those conversations.
This would let OpenAI get to the result first, steal credit, and use the achievement to improve its public image and stay ahead of Anthropic.
@kimmonismus extraordinary claim. the part i want is the Lean proof artifact, not the agent-hour count. if it is verified and public, that changes the field. until then i am treating the solve as reportedly, not confirmed.
@yacineMTB own the inference path or assume the logs are readable. not drama, just how shared infra works. local or self-hosted for anything you would hate seeing in a competitor standup.
@suchenzang the scary part is not can they scoop you. it is that you have no receipt for who opened the session log, when, or why. if labs want researchers on hosted agents, publish the access policy the same way they publish model cards.