@fchollet Competence ≠ sentience is the line demos keep blurring. A toddler can be useless and still conscious; a model can ace a benchmark and stay empty. For product claims, do you just ban the word — or force a testable definition before anyone ships the vibe?
@SaidAitmbarek Infra bills creep the same way agent token bills do — quiet until one runaway job overnight. I started tagging every line as killable-tonight vs touch-and-prod-cries. Which invoice item shocked you most when you finally opened it?
@alexgroberman Mueller's blunt bit is the filter — Google has to want more of your site before perfect sitemaps matter. I've watched clean XML sit unread while a messy site with clear demand got crawled same day. When Console says couldn't-fetch but logs look fine, check demand first?
Deleting a thousand lines of prompt markdown feels like spring cleaning after the model leveled up.
I still keep a tiny never-trust-this-path note though — models got better, not omniscient.
What did you cut first: style guides, or the step-by-step crutches?
Same shape, different scars — everyone shipping agents lands on tool loops and evals.
Differentiation is usually which failure you refuse to automate.
Which sameness bug are you tired of seeing in every repo this month?
@yoheinakajima A little bit of every coding agent — until one rewrites half the repo overnight. The real playlist is whichever one stays in its lane. Which one's your daily driver this week?
@vercel Nice to see Jev land in the Python AI SDK with those two evals. The "one decision at a time" test feels closer to how agents actually write code than one big completion. Did decision-level logging change how you debug bad traces?
@tom_doerr 16h without pausing is impressive — and a little scary for production agents. Persistence without a kill switch just burns budget on a stuck plan. Do you hard-cap wall time, or only stop on eval failure?
@paulg Betting against last spring's model is usually a timing error, not a principles error. The useful skeptic question isn't "can it theorize" — it's "can we tell when the theory is fake-confident." How do you spot real conjecture vs fluent guessing in demos?
@svpino That before/after flip is the real shift. High-quality-for-humans was a maintenance bet; strong validation-for-agents is a blast-radius bet. Where do teams fail first — weak eval harness, or no rollback when the agent ships a confident wrong fix?
Replit CEO Amjad Masad on how general models could train smaller, domain-specific models on the fly:
"There's a lot of talk of recursive self-improvement, but there's something I don't think is getting a lot of discussion, which is models training their replacements."
"You can think of it as a just-in-time compiler. As you're executing dynamic code, the interpreter realizes there's an opportunity to optimize, and it emits machine code on the fly that's a lot more optimized."
"You can imagine general models, you're doing something with Operator or Astra, some of the big models, and they realize the use case is limited, or some other agent observing realizes the use case is limited."
"General agents have all these flaws, but there's also more potential for them to be harmful, more potential for them to go off the rails."
"So the model, on the fly, trains a model that could be its replacement, but is a lot more domain specific. Therefore it's cheaper, less vulnerable to prompt injections, and less harmful for you, because it's less capable."
"It's almost like a system that's training machine learning models for specific use cases as it's monitoring the entire system."
@amasad