It seems that VCs have forsaken backing natural entrepreneurs in exchange for people with impressive résumés at big-name companies. No doubt it is impressive to climb the ranks at a big company to a VP or even staff engineer position, but does that make you an entrepreneur? I would argue that it demonstrates almost the opposite skill set.
We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI.
The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may require longer investigation or coordination with third parties.
We’ll prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety or mitigation.
Alongside the framework, we’re publishing six reports on instances of misaligned behavior we’ve observed during the training or evaluation of our models in the last six months.
This is a starting point. We’ll refine the process through experience and public feedback, and share more reports on an ongoing basis.
https://t.co/ismCCkeE0L
As models get better at detecting prompt injections, prompt canaries inside networks become less useful.
In my experience, leading open-weight models simply won’t bite on malicious instructions. In some cases, they’re actually too conservative about what they classify as a canary.
The better defensive approach is straightforward: rely heavily on honeypotted SPN accounts inside Active Directory and design them so a human user would never accidentally interact with them.
At the same time, aggressively map and remediate ACL attack paths. When clear ACL escalation chains are removed, autonomous agents are more likely to fall back to roasting attacks, which are often easier to detect and instrument.