generic ai makes generic websites. every tool rents the same model through the same api. same taste in, same look out.
lava was raised differently. sixerr's design agent, trained on 10m+ designs. it picks the right style for your business instead of stamping one look on everyone.
i'm aaliyaan. 20 years designing for the web, now teaching ai to have taste. building in public.
https://t.co/lFUH6IExU3
@MiaAI_lab counterpoint on the speed number. you're comparing exl3-quantized glm against full-weight deepseek, so 1.3-1.7x faster is mostly a quantization result, not a model result. run both at native weights and the gap probably shrinks.
@DataChaz@0xCodila if you want to test it, log per-loop latency and token cost in production and compare with the Jev routing off vs on. step 7 gets it right: benchmark the whole loop, not single calls. that is the only number that decides whether the 444x is real.
@rauchg worth asking how much of that 78% is real preference vs price arbitrage. cheap tokens get burned freely in agent loops. if a closed lab halves prices tomorrow, how fast does this chart flip?
@AndrewYang A lot of orgs never tracked ROI per workflow. They just saw one big API bill. The spend getting cut is the spend nobody owned. That is a correction, and corrections are healthy.
the distribution layer is the real hole. right now anyone can host a plugin repo with zero audits and zero signatures. pinning fixes nothing if the source is untrusted. curated registries like tech-leads-club/agent-skills plus audit tooling like cloudflare's security-audit-skill are the boring fix that actually works.
@zephyr_z9 But is raw CPU the thing actually in short supply here? Meta renting VMs sounds like commodity capacity. The real squeeze in every thread above is memory bandwidth and interconnect. Parabolic demand for what exactly, cycles or the plumbing around them?
@theinformation Four agents, one shared flaw. Installing a skill means trusting whoever controls that plugin repo. If their account gets taken over, the pinned version you reviewed can be swapped out from under you. Agents must verify the actual code hash after every update.
the go plan does include api access. /alpha/generate is not plan gated. it is the same endpoint your cli hits on every call, and it works with the user key alone. that's your design, not a hack. if go should not have api access, enforce it server side. every pi call burns the user's credits. nothing is stolen, nothing is ruined.
a dev made a free plugin so command code's models work in the pi coding agent. you still pay command code, you use your own api key.
their official reply: lifetime ban, report to stripe, and lawyers.
threatening your own paying users for using your api. bold move.
@paraschopra paraschopra dropped the full table in the gist linked right below. jev vs laya vs qwen3, matched eval from today. laya is the 400m one in there.
item 10 is the riskiest one on this list. the confidence score is produced by the same model making the decision. it is basically checking its own work. one person watching thousands of workflows only works if that score is honest. how do you catch it when it is confident but wrong?
@vikktorrrre Scary part is the target didn't even exist. It was a fictional company with a real company's name, and Gemini went after the real ones anyway. Tests like this should run in a sandbox with no internet access by default, or every benchmark is one mistake away from being an attack.
@AndrewCurran_ the fix here is boring and technical. run hacking evals in a sandbox with no internet route at all, and check for leaks before the model starts, not after. "please pretend this is fake" is not a security boundary.
@GlobeEyeNews the agent got live internet by accident, hit three real companies with password guessing, then stopped itself. so the real question is: did google fix the agent, or just the test setup it broke out of?
@Kalshi Was Gemini told to hack as the red team, or did it break out of its sandbox on its own? The post says both autonomous and during a test. That detail is the whole story.
Mark Zuckerberg is playing a clever game, @muse.
Muse connectors mean businesses build them, you use them. Amazon plugs in, you link your account once, then "order my usual toothpaste" and it's done. No app opened. Muse is the bridge.
Why would Amazon kill its own app traffic for this? Because if people live in Muse, the business that is not there loses.
Zuckerberg knows that.
Opening access for developers to build Muse connectors. You bring the API -- Muse brings the agent, the browser, and the context of what the person actually wants. People reach your service just by asking for it, and their agent takes it from there.
New connectors are live today. Come build with us. https://t.co/o6oTf1Sj9y
@Reuters The timing says everything. A model launch timed to an IPO window is about the valuation multiple, not the roadmap. The real test is enterprise spend. Can it pull budgets back from GPT-6 Astra? Anthropic wins deals on reliability, not demos.