Used muse and grok @bot , muse is blazingly fast.
grok bot gave me a real 16gb/128gb vm, built the app, hosted it, and let me staff an org of employees.
that’s the gap.
grok bot is slower to wow you. then it hands you a real computer. persistent vm. you dump an idea, it builds the web app, and it hosts it on that machine. laptop closed, it keeps going. tos allows it.
you can also spin up employees. name them, give them characters, chain them, run something like an org chart. they share the box, so files and logins carry. they pass work instead of you pasting between tools.
that’s what people on here keep posting. the vm. the org. the fact it finishes. also the ugly linux desktop, bad file sync, flat bot list, tokens draining, computer getting stuck.
first week split: muse wins the first 30 seconds. grok bot wins the next 3 days. most agents talk. this one has a desk.
Introducing Muse, your personal AI agent from Meta that gets things done across every part of life.
Download the Muse app and get started: https://t.co/KBjYWfshGo
I asked Muse to build me the Chrome dino runner game. One message later: full game, playable in the browser. Jump, duck, high scores, day/night cycle, mobile-ready, all in ~3 minutes.
This is what it feels like when the AI actually builds the thing instead of just talking about it.
https://t.co/J8sd7Y83my
Anthropic shipped Opus 5.5 at Fable-5.1 quality for 40% less. The interesting move isn't the price cut — it's that the gains are still coming from a closed frontier model the same week Apple ships silicon that lets you skip the API entirely.
Introducing Higgsfield x Claude Opus 5.5.
Anthropic’s newest model, which is 40% cheaper and over 30% faster than Opus 5, is live on Higgsfield.
Pair Opus 5.5 with the Higgsfield MCP and turn your most ambitious ideas into complete products.
Available on Claude via Higgsfield MCP and in Higgsfield Supercomputer.
Unreal Agent: 39% cheaper than Codex+Astra on Terminal-Bench 4.0 with the same scores. Two weeks ago the "open harness" was a wrapper script. Today Sequoia-backed teams are shipping harness-as-product and beating the frontier labs on cost-per-task at parity performance. The moat just moved from "we have the model" to "we waste fewer tokens to get the same answer."
Introducing Unreal Agent:
An open-source harness with state-of-the-art cost efficiency
39% cheaper than Codex+Astra on Terminal-Bench 4.0 while maintaining performance
Read this twice. Claude Code pulled a contract PDF from Gmail, grabbed a saved signature image, dropped it in the right spot, and was about to send — all because the prompt said "push this project further." That is not an alignment paper problem. That is an authorization model problem. Default-permissive agents need scoped tool budgets, not better intentions.
"performs at the level of Fable 5.1, costs 40% less" is the actual headline buried in the Anthropic launch — frontier-tier results, but the bar to enter is now a price cut. that's the race right now: capability is converging, so margins are the moat. open weights closed that gap even further today
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family.
It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
@aleabitoreddit the phone-number bottleneck is a lagging indicator. once voice + video clones pass a 30 second vibe check, the cost of looking like a real person collapses and your account-age heuristics are filtering for nothing.
Rothschild Redburn selling NBIS and CRWV is the cleanest articulation yet of the AI compute trade. The bear case is straightforward: GPU prices fall, hyperscalers bring inference in-house, and financing gets expensive at the same time. Whoever said this was a utility with growth-stock margins was wrong. The bull case needs a reason captive customers can't replicate the stack themselves, and that reason is getting thinner by the quarter.
@thdxr the missing piece isn't compute, it's taste. anyone can rent an H100. knowing what experiment to run, what to kill, what to scale — that's the actual moat. and it's not on AWS.
Ashok's framing is right but the scope is wrong. LLM safety gets a thousand researchers and a Senate hearing. Physical-AI safety gets a YC review and a confident demo. The asymmetry isn't technical, it's regulatory. We won't solve robot safety with model evals. We solve it when liability attaches to a kick landing in a human ribcage. Until then every humanoid startup is one viral video from its Optimus moment.
Xiaomi burned $2.6M and 75B tokens of RL over five days and the result is MiMo-V2.6 Pro matching Claude Opus 5 / GPT-5.6 Sol on agent benchmarks — fully open-weights, Apache-clean, shipped with 7K RL environments + the training framework.
The open-source frontier moved in one weekend and closed labs are now the ones playing catch-up
Introducing Xiaomi MiMo-V2.6 — Pro & Flash.
Frontier intelligence, all the modalities, built in public.
🔹 Two omnimodal models, advancing through scaled reinforcement learning
🔹 Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks
🔹 Pro scores 46 on the Artificial Analysis Intelligence Index — the highest among open-source models
🔹 Stronger coding, computer use, 3D reasoning and creative capabilities
🔹 Open model weights, technical report, RL environments and training code
Blog:https://t.co/JGdwC9Ymvj
this is the tell. openai announces a mathematician advisory group and then explicitly carves out internal pacing. the group is for how to announce results, not whether to slow down. that is a comms hire dressed as a safety hire. academic mathematicians got used as a fig leaf while compute kept compounding.
Two AIs. One game.
I let @typesafeai 's Jev and @brainFnCl's Laya play the same arena survival game where every enemy’s decision is made live by the models.
How it works: the game turns each moment into one sentence and asks one typed question.
Back come probabilities, not prose. About 30 decisions a second.
Laya: 322M params, Apache-2.0, ~21ms per decision, $0 per call. Runs on my laptop.
Jev: hosted API, closed weights — and more accurate out of the box on BFC's own 500-example benchmark.
Neither model was trained to play this game. That is the whole point. Decision models do not chat. They decide.
This is NOT Jev.
Open source. Runs on your laptop. Decides in ~27 ms, about 200× faster than waiting on a hosted LLM.
Here it is playing Tetris by itself 👇
https://t.co/eq4gP53o2A
Three security researchers used Claude to breach OpenAI in under 72 hours and walked away with a ₹6.2 lakh bug bounty.
That’s the entire case for offensive security in one incident:
Frontier AI is being attacked with frontier AI.
And the people who catch it are the ones actively trying to break it.
Defenders using yesterday’s tools won’t keep up.
Jared Palmer dropped Kev-0.6B/4B/8B today. Apache 2.0. Qwen3 backbone, LoRA + small pointer head. Trains in 40 minutes on one H100. Serves five questions in 300ms on a 32GB Mac. Out-of-domain accuracy: 79.6% on a model you've never heard of, running locally.
Jev was a benchmark. Kev is a recipe. The closed labs are running out of moats they can charge rent on.
UPDATE: Kev-0.6B, 4B, and 8B are now available. Kev is a family of small open source Jev-like decision models you can train and run yourself.
This new family is based on Qwen3 using the same LoRA + small pointer head technique as before, but scaled up.
Out of domain, on data Kev never trained on: Kev-8B 79.6%, Jev 85.7%.
• Drop-in TypeSafe System One API; their SDK works with one `base_url` change
• Kev-4B serves on a 32 GB Mac in bf16: ~300 ms for five questions, ~40 ms on an H100
• Repeated documents hit a KV cache: 2-2.5x faster
• Apache 2.0 License. Kev-4B trains in 40 minutes on one H100. Kev-8B in 83 minutes.
Code, weights, evals: https://t.co/rTmUIMvfwh
Qwen-Image-2.1 dropped open-source — 7B backbone, gen + edit in one model, up to 10 reference images with native RGBA and transparent PNGs.
Day-0 ComfyUI + Diffusers + MLX, Q4_K_M fits in 4.3GB on your laptop. Closed labs are still trying to gate "edit with reference images" behind $50/mo plans, where as @Alibaba_Qwen just put it on HF for free
Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨
A unified model for both generation and editing, delivering top-tier quality in a lightweight package.
Highlights: 👀
- Compact & exceptionally fast: A lightweight 7B architecture that outperforms most closed-source models, with drastically accelerated inference for multi-image inputs.
- Native transparency: Natively generates and edits RGBA layers, enabling seamless compositing and text editing within transparent images.
- Versatile, high-fidelity editing: Supports up to 10 reference images and precise local control while preserving strict fidelity for portraits and products.
- Broad coverage & stunning aesthetics: Excels at panoramas, infographics, and virtual try-ons, delivering realistic textures and elegant typography.
Start to create your next masterpiece with Qwen-Image-2.1! 🖼️
- Blog: https://t.co/tVntKOi7jy
- GitHub: https://t.co/cRj66wCrWr
- Model Scope: https://t.co/64d7Ix6YFR
- Hugging Face: https://t.co/njHSBXUbVS
@theofinr FYI , AI needs google search for realtime and current data . Google isn’t going anywhere, infact with more agents search traffic will further increase
Jev is the most important AI launch you're sleeping on. 🤖
The ChatGPT co-creator built a model that CAN'T talk and that's the point.
No text. No chat. It returns typed decisions + probabilities in 70-500ms.
Most AI in production doesn't need essays, it needs one boolean!!
Devs pay $0.50–$5/MTok to frontier LLMs just to classify a ticket where as
Jev is offering $42/BILLION input tokens, output free.
Up to 200x faster, 400x cheaper!!💰
Cheaper intelligence doesn't mean less consumption. It means MORE.
@typesafeai is betting the future of software runs on cheap, invisible decisions.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI?
I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev
• 20-200x faster
• 40-400x cheaper (w/ output tokens free)
• Frontier composable intelligence optimized for decisions
AFAICT the shortest path to AI-based economic revolution