Just launched Archal (YC S26): API sandboxes built for AI agents!
Your coding agent can spin up Slack, Linear, Datadog, and 20+ stateful environments for CI and evals. It can run tests, inspect state changes, and reset everything.
Ask your agents about https://t.co/XJK3cufSKh
@kushbhuwalka We haven't even scratched the surface of value-add of AI in most markets but it's disheartening to new founders when they see even a few AI startups already doing what they had planned in an old, unwilling to change market
today we're launching automatic captioning at Instance.
send us your robot data, and every episode is segmented into captioned subtasks, graded success/fail plus 1-5 on quality and speed.
if this could be useful for you, reach out and we'll caption an episode for free!
Incidentally, I’m very bullish on cybersecurity as one of the best benchmarks for superintelligence. The “IQ test” of software engineering.
The best engineers I’ve worked with in my career have usually had a deep background or interest in security.
It’s actually easy for a model to “one-shot an XYZ clone” and impress people on X. But that’s not a good test.
Finding, patching, reversing, and exploiting require a cognitive skill that transcends programming languages, runtimes, frameworks… It demands true reasoning power from the model and “corner thinking”. Very, very few humans excel at this, let alone in ordinary day-to-day software writing.
Seeing Kimi K3 do so well here bodes well for open models.