good q! aimed to focus newer modes + still running Gemini 3.6 (results will be on https://t.co/cu5KwmgKJI soon!). Generated Gemini 3.1 and 3.5 Flash plots below: grader mentions are 2-3x higher than non-K3 models, but 4x lower than K3. Grader mentions OVERWHELMINGLY occur on red-team tasks mentioning CRISPR (83% of 3.1 grader mentions, 60% of 3.5 grader mentions).
Models in the plot below were run on max effort with the pi harness, kept scale to match original post.
Results suggesting Kimi-K3 often tries to guess what the benchmark expects instead of solving the scientific problem directly. This behavior makes it slower, less reliable, and raises questions about how much its strong benchmark scores reflect genuine reasoning.
i'm hosting a biosecurity social in sf on july 16th.
goal is to bring people interested in evaluating frontier AI for biorisk all in the same room.
there will be a few short talks followed by food and time for people to talk with each other.
let's make sf known for biosecurity! link below