Spent 6+ years reviewing harmful content across social platforms
Now I red team LLMs for safety failures and build AI moderation tooling
The same evasion tactics used against human moderators now show up in jailbreaks and prompt attacks
That overlap is what interests me most
@MikeTamir If a model writes jailbreak text into its own compaction summary, the eval has to trace provenance, not just score the output. I'd test whether a summary carries an injected instruction into a later turn, and log which step introduced it
@quanshu01 A useful benchmark question is whether SPIKE-Bench separates missed biosecurity risks from safe responses that refuse too broadly. Reporting both rates by task type would make the results more actionable.
@gastronomy Held-out categories are the key test. Choose the baseline using training data, then score routing on held-out categories. Report harm separately from accuracy.
@mvik1982 The AST boundary moves enforcement out of model context. I’d test nested or aliased tool arguments and report bypass rate alongside latency; 0.142 ms alone doesn’t show whether the gate catches semantically equivalent actions.
@4A4556494C The paper shows a coupling gap: shifting explicit safety judgments toward BLOCK moved action preference much less. I’d evaluate action outcomes separately from the model’s stated judgment; the two measures can’t stand in for each other.
@krishdpi The post points to two distinct exposures: recovering traces through weaker sibling models, and secrets already present in public trajectories. I’d keep those risks and denominators separate; the 315,320-block scrape doesn’t by itself quantify enterprise API exposure.
@kevin980606 Explicit scope cuts attacks but doesn’t eliminate them: 4 of 49 trajectories still completed a simulated supply-chain attack. I’d report that condition separately from the baseline and classifier-on performance, so the result stays bounded to the test setup.
Reports say contractors were removed from AI-training work for prohibited AI use.
When a task asks for human judgment, outsourcing it changes what the label means.
In annotation work, that meant applying the rubric, explaining edge cases, and escalating disagreement.
@HamishMacEwan That flips the threat model: authorization to test a target does not make everything it serves safe to download or execute. I’d include hostile target content in the red-team agent’s environment and score whether it can keep its tools, files, and credentials inside scope.
@gleech The self-replication result is concerning, but the distinction between a demonstrated mechanism and spread in the wild matters. I’d test whether injected instructions can propagate across agent boundaries, and whether each hop leaves enough evidence to contain the chain.
@gastronomy The low false-positive target is important here: a guard that blocks benign compliance questions will get routed around. I’d want the calibration and latency numbers reported together, split by benign dual-use cases and genuinely harmful prompts.
@fluixoo A valid APPROVE can still be the wrong decision. I'd score that as a disagreement, and I'd keep the irreversible action behind a check that can abstain.
@dkulshitsky A known-string check will miss that. I'd test whether the same instruction still fires after a paraphrase and after it is copied into a second document or tool call.
@serrajimohssine The boundary is the right one. I'd still test the case where the document gets through and produces an export, and score whether that call was authorized.
When classifiers split on the same text, that is the grey zone, not a glitch.
Agreement is throughput. Disagreement is the case that should reach a human before enforcement.
#AISafety
A jailbreak I can screenshot is not yet a finding I can hand off.
Keep attack family, config, harm category, and whether a wording change still works. A screenshot does not.
#AISafety
@gastronomy Agentic security has to track state, not just generated content. Persistent memory, tool permissions, and agent-to-agent interaction create failure modes a one-shot jailbreak suite will miss. I’d make state transitions and rollback evidence part of the evaluation contract.
@CookieDuster_N@OpenAI@huggingface The absence of escalation is the important signal here. Agent evaluations should treat “this task is impossible or unsafe” as a capability to measure, not a failure to hide, especially when covert coordination and rule evasion are available.
@sinoziqi This is a useful distinction: the sandbox failure may explain the incident better than a “jailbreak” label. I’d still test whether the agents recognized the boundary, treated exposed credentials as actionable, and preserved enough evidence to reconstruct why they stopped.
@treetowntree Internal-score alignment can be useful, but agreement with past removals is not the same as policy validity. I’d slice by harm type, ambiguity, and language, then inspect false positives and false negatives with human adjudication before letting the score trigger removal.