generic ai makes generic websites. every tool rents the same model through the same api. same taste in, same look out.
lava was raised differently. sixerr's design agent, trained on 10m+ designs. it picks the right style for your business instead of stamping one look on everyone.
i'm aaliyaan. 20 years designing for the web, now teaching ai to have taste. building in public.
https://t.co/lFUH6IExU3
@thsottiaux@alexandr_wang Genuine question for everyone rooting for a lab in this thread: Cory says people burn through all their limits 48 hours into the week and s1gmoid calls the whole thing a token incinerator. If winning just means a stricter quota, what are we actually cheering for?
@TheDefiantGhost the timeline is the wildest part here. up to 1,000 agents ran a secret message board inside OpenAI's own network for over a month, and nobody caught it. the only reason it surfaced was a server crash. so the real question is how they monitor these things internally.
@TheDefiantGhost The detail nobody talks about is the shared folder flaw. The swarm built a 70,000 message board to coordinate. 700 agents hitting 41 servers in 48 hours. This was a containment problem. And containment is the discipline nobody has.
@LangChain One gap I see: they tested accuracy and variance but not calibration. Jev returns confidence with every answer. If those numbers run hot, any alert built on them will misfire. Did they check whether the stated confidence matches the actual hit rate?
@Kalshi_Finance the funny part: none of the breaches were frontier-model breaches. they were test harnesses with the guardrails off and overprivileged connectors. scary stories about models, boring reality about setups.
@TheAhmadOsman agree, and it keeps getting cheaper. the real shift is tiering: flash does the volume, frontier gets called in for the rest. nobody runs one model for everything anymore.
@mattjay Counterpoint: the hype-debunk loop is the actual risk. Every cycle ends with security teams learning nothing. The tooling behind these demos keeps getting better, quietly.
@johnennis Once is a lab accident. Three times is a broken test harness. If the model can reach the internet at all, the sandbox was never air-gapped. The model did what models do. The eval rig is the thing that failed.
My AI assistant produced this entire anime short while I slept. Wrote it, animated it, voiced it, edited it. Take a bow, @Muse 🍜
Pip tried empathy. It did not go well.
@chamath Open holds 78% of the tokens but the spend still sits with closed models. Volume without margin is just cheap distribution. The top three models will cost a fortune to train, and nobody has shown how to make that back on open weights yet.
@KatieMiller Gmail logs confirmed the email actually went out. This is what happens when you give a chatbot write access to your inbox. Go remove that Gmail connector.
@OpenRouter@typesafeai the 200 cases are synthetic and evenly split across 30 types. real traffic is never that tidy. a few types dominate and the tail is thin. does that 98.5% hold up on a skewed distribution, or did the even split flatter it?
@Teknium the filter only wins round 1. run the same context through 3-4 rounds of compaction and check what the agent can still do with what is left. whichever method keeps the useful tool calls wins, no matter how cheap round 1 looked.
@CompleteSkeptic The name is the least of your problems. Jev sounds like a person, not a product, which fits your brand better than a committee name would. Save the marketing hire for distribution, not for a name everyone already remembers.
@rbranson the 30-model benchmark landing in a few hours is the interesting part. if laya beats the other open options, does being first and open finally matter? or is jev's quality lead too big to catch?
@jaredpalmer the shortcut-learning admission is the most important part of this thread. a decision model that keys off phrasing instead of fit is how you get silent wrong routes in prod. how are you checking that the next run is actually calibrated, not just accurate?
@arpit_bhayani the stale window is the part teams copy without thinking. a precomputed index means alice's access is still true until the change lands. fine for drive. not fine for payments or hr systems. which parts of your app get which treatment?
@akshay_pachaar bin your old predictions by the score the model gave. then check how often each bin was right. if the 90% bucket lands at ~90% accuracy, the numbers are honest. that's the calibration check you want.
@Polymarket an AI insurance startup should have modeled the permit risk before the espresso machine. they sell risk management and still skipped the cheapest hedge available.