So an agent can build a web app, start its server, create an artifact, enable live preview, and bind it to that process.
Then I can just access the artifact from the webapp.
@miaugladiator1 The good thing about the new DeepSeek release is that the 4.0 model is perfectly usable now. It has almost no throttling or capacity issues.
It's interesting because a lot of infrastructure is pretty poorly secured, but until now there wasn't really much incentive to go after everything at scale. Now, though, it's very cheap to put together a model and have it continuously probe systems and look for vulnerabilities 24/7.
I think things could get pretty rough in a few months or years.
Yo noto eso básicamente, que puede hacerlo todo, pero usualmente no querés que tome decisiones arquitecturales o en especial cosas de UX. Y eso es algo bastante interesante, porque la gente todo el tiempo te dice mirá este sitio web que hizo Codex o Fable, pero la realidad es que si vos querés hacer un flujo de usuario intuitivo te va a hacer cualquier porquería y tenés que medio razonarlo con la ia y decirle el flujo y entonces ahí te lo hace bastante bien.
@composio It would actually be pretty useful to see how much time was spent on each harness for the failed task, so we could have some sort of metric. Otherwise, I don't think this is particularly useful on its own.
@adamlyttleapps Sometimes I kind of doubt myself and wonder whether maybe this time it really is as good as people say. But then the hype dies down, like, two days after the release, and nothing really changes in any fundamental way.
@Polymarket What could be interesting about Chinese AI chips is that they may start building their own hardware ecosystems instead of relying so heavily on CUDA, which could lead to some interesting developments.
I remember that, about a year ago, Anthropic said that one of their models had noticed it was being tested during one of those needle-in-a-haystack evaluations. And they framed it almost like, "Oh, it became self-aware."
And let's remember that this was with models that were still pretty dumb by today's standards.