Sakana AI の Applied Research Engineer 合田晴紀が、Auth0社主催「Camp AI」に登壇しました。本人が開発責任を担ったUltra Deep Research Assistant「Sakana Marlin」を含む当社プロダクトをご紹介し、エージェント時代のプロダクト開発について議論しました。
Fugu-Cyber is available as a new endpoint at https://t.co/FxG5zvZ5pb
Achieving high scores on a cybersecurity evaluation is only the beginning of the story. What does it actually take to bring frontier AI into enterprise defense?
Recently, there has been a lot of fearmongering about the cyber capabilities of frontier models. Much of the industry narrative suggests that simply granting an organization access to a frontier model with cyber capabilities will instantly solve its security challenges.
We believe it is time to ground this conversation in reality.
As highlighted in a recent Nikkei Digital Governance report, having access to a frontier model like Mythos does not automatically solve enterprise security. Large organizations often struggle to operationalize these tools without specialized internal talent and deep integration into proprietary source code. Successful deployments require human expertise in cybersecurity alongside frontier capabilities.
A highly capable API with strong cyber reasoning is an incredibly important piece of the puzzle, but it is not the entire solution. When deployed in isolation, raw models will inevitably generate false positives. True enterprise defense requires orchestrating these models with sub-agents specialized for cybersecurity and human-in-the-loop processes to confirm whether a vulnerability would actually trigger in a real environment.
This is the exact challenge Sakana AI’s Applied Enterprise team is solving. We are working closely with major Japanese institutions to build the specialized harnesses required to deploy these models, including Fugu-Cyber, into production safely.
Reference: Nikkei Digital Governance report (Japanese)
https://t.co/spEbQWBca0 🛡️
Introducing Fugu-Cyber: an update to our Fugu orchestration model.
It achieves state-of-the-art performance on real-world security benchmarks, matching cyber-focused frontier models like GPT-5.5-Cyber and Mythos Preview.
https://t.co/5Nh1eBPhHg 🐡
"Feedback-to-Rubrics: Can We Extract Expert Criteria from Inline Comments?" will be presented at the Workshop on Human-AI Co-Creativity: Advances, Opportunities, and Challenges on July 11 at #ICML2026
Paper: https://t.co/MPfNX6l2z1 (Extended full-paper version)
LLMs are increasingly used for writing and review support, but their usefulness depends on context-dependent criteria, such as expert preferences or organization-specific conventions, that are often tacit, undocumented, and difficult to elicit directly. We propose a problem setting for learning reusable natural-language rubrics from accumulated inline comments on artifacts such as human-written or LLM-generated drafts.