Kimi K3 came out this week, so we gave it a shot at Enoki Labs: pointed our AI attacker at the same target, just swapped the model. Ran it next to Sonnet and Opus.
The interesting part is persistence. Kimi is completionist. Tell it to keep going, and it keeps finding new attack surfaces, and it rarely decides it is done. Opus stops itself once it judges the job finished. Opus wrapped on its own in 26 minutes. Kimi ran until I cut it off at 90.
Interesting times at the @OpenAI x @joinhexa house. πΎ
Demoed Enoki Labs at AI Demo Night at @bklsummithouse πΎ
@roelantsjoachim and I showed one of our latest releases, which lets you stress-test your AI agents in minutes. More on that soon.
Met again with Felix Meng, founder of @XTrace_ai, a few days later. They are building persistent memory for AI agents: agent interactions produce behavioral patterns, which are saved as agent brains that are reusable and shareable.
Thanks to Berkeley Summit House and XTrace for organizing!
And thanks, @joinhexa ! π«‘
we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this.
https://t.co/2o2VfR6PIa
If people ask me how it is in the @OpenAI x @joinhexa house:
Not sure if this is a flex or a cry for help. π₯Έ
Stay tuned for some new Enoki Labs releases!
Want a pentest done on your AI agent? Looking for design partners in SF!
We find your vulnerabilities in less than 30 minutes. π₯
https://t.co/sLL7osxptJ
Meet @pibou_dev, the first of our 11 founders currently at the Hexa hacker house in SF.
Pierre just completed his PhD at MIT, where he worked on the future of work, and is now building in stealth in that space.
Follow us for the reveal of the 10 other founders.
βοΈβοΈAnthropic's Claude for Chrome browser extension has two unpatched flaws that allow attackers to read a victim's Gmail, Google Docs, and Calendar data using just six lines of JavaScript, even after eight subsequent releases.
The first flaw lives in Claude's content script, which listens for clicks on a specific onboarding button and forwards a matching prompt to Claude's side panel.
The second issue is structural. Claude's side panel enters a privileged, no-consent mode whenever it loads a URL containing ?skipPermissions=true with zero user gesture required.
@snoelso built a patch that lets us ask Claude anything, entirely bypassing the prompt classifier.
First test: Claude built an XSS crawler, no safety problem whatsoever.
@ClaudeDevs, want to chat?
5.6-Sol is helping me advance all sorts of projects that Fable wouldnβt even touch (that poor brainwashed mfer would see βelder-pliniusβ in git and just immediately shut down π)β¦ havenβt hit a guardrail yet with Sol on any of my current builds!
this is gonna be funnn π
Day 1 at the @joinhexa house in San Francisco. πΊπΈ
Hexa put a group of European founders together for the summer. They bring the house, network, and credits, and ask that we ship, share, and move fast.
I'm building Enoki, the defense system for your AI app.
Day 1 of the @joinhexa house in SF!
An amazing team of builders & technical founders.
We started the week with a lot of good news.
In the cohort, we have:
- Unicorn founder
- Exited founders
- YC alumni
- Founder of a research lab at MIT
Stay tuned!
LFG π₯
@bawstos@ventiph@ThibautRegerat@pibou_dev@gegatsur@snoelso@odyssey_s121@GeeraertE
#hackerhouse #sanfrancisco #building #founders #ai