AI coding agents are no longer suggesting code. They are changing the software factory.
CTEM should decide what they are allowed to fix. Consequence should decide how much authority they get.
1Password’s Off-by-1 Labs generated 6,080 patches on six recent CVEs. Only 26% were a clean fix. 53.9% failed, introduced another vulnerability, or both.
Generation is cheap. Assurance is scarce.
Reason. Ground. Generate. Validate. Authorize. Execute. Verify. No single agent should own all seven.
Capability does not imply authority.
Where are you drawing the line between what an agent may propose and what it may change?
I develop the broader argument in a research paper on how AI changes enterprise cybersecurity strategy.
It treats hunting, detection engineering, adversary emulation, exposure management, control validation, and response as one learning system.
The strategic question is not how many hunts were completed. It is whether the organization can turn evidence into an authorized defensive effect before business consequence, then verify that the change worked.
Evans, A. (2026). How artificial intelligence changes enterprise cybersecurity strategy (Version 1.0). https://t.co/zA15YsAeCV
The hunt is not the outcome.
The learning loop is.
Sources in the replies below. Views my own.
What has a hunt changed in your environment in the last 90 days that an alert never would have?
One emerging model shows where hunting may be heading.
CrowdStrike, Unit 42, and Mandiant increasingly hunt across broader telemetry. Mars Security says it can query existing SIEM, EDR, identity, cloud, and data-platform telemetry where it resides, then use current threat intelligence to drive hunts, detection logic, and coverage assessment.
That shifts the question from “How many hunts did we run?” to “Against the threats that matter now, what can we see, where are the gaps, what changed, and did the change work?”
These are vendor-stated capabilities. Independent enterprise-scale validation is still needed.
https://t.co/P9rx9dcw0f
AI can accelerate hunting. Current evidence does not support removing the hunter.
Chona, Kozlov, and Kumar (2026 preprint) evaluated five frontier models against 106 attack procedures in raw Windows logs. The best model correctly flagged 3.8% of malicious events on average. None met the authors’ bar of at least 50% recall on every ATT&CK tactic.
OpenAI’s collective cyber-defense call still asks for continuous testing, verified fixes, shared intelligence, continuous monitoring, and traceable agent identities.
Chona, Kozlov & Kumar (2026). https://t.co/eKz6LWsq2B
OpenAI (2026, Aug 27). https://t.co/voPGveMsTb
Speed and persistence can exist at the same time.
CrowdStrike observed that 88% of exploitation involving vulnerabilities with public PoC code occurred within 48 hours of PoC release in H1 2026.
Mandiant, looking at 2025 investigations, reported a 14-day global median dwell time and BRICKSTORM dwell times approaching 400 days. It also notes that standard 90-day log retention can leave organizations blind to initial access and full scope.
CrowdStrike (2026, Aug 3). 2026 Threat Hunting Report. https://t.co/4szqEKAVLW
Kutscher, J. (2026, Mar 23). M-Trends 2026. https://t.co/UvJWWc1o7i
Threat-informed defense is broader than hunting. MITRE’s INFORM model treats it as a program across threat intelligence, defensive measures, and test and evaluation.
CISA made the assurance problem concrete. In a proactive hunt at a U.S. critical infrastructure organization, insufficient logging blocked several planned ATT&CK procedures. A negative hunt only means something when the evidence can support the conclusion.
Cunningham & Valenzuela (2026). From insight to impact: INFORM your defense. https://t.co/Mz02qr9v1k
CISA & U.S. Coast Guard (2025). AA25-212A. https://t.co/LzxtiZf1Yh
AI can generate hypotheses, correlate telemetry, and search at a scale people cannot. It should not replace the hunter.
A 2026 preprint gave five frontier models raw Windows logs across 106 ATT&CK procedures. The best model flagged 3.8% of malicious events on average. None hit 50% recall on every tactic.
The question is also getting larger than malware: is a human or machine identity using legitimate authority in a way the enterprise never intended?
The hunt is not the deliverable. The next change is.
A useful hunt should improve a detection, close an exposure, constrain an identity path, fix a control, or kill a bad architectural assumption.
Then detection engineering makes it repeatable. Adversary emulation tests whether the control works. Continuous Threat Exposure Management asks whether the path to consequence is actually closed. Then you hunt again.
Those two clocks are not a contradiction.
CrowdStrike: in H1 2026, 88% of observed exploitation involving vulnerabilities with public PoC code happened within 48 hours of PoC release.
Mandiant: 2025 investigations had a 14-day global median dwell time. Some BRICKSTORM cases approached 400 days. Standard 90-day log retention can erase the start of the story.
Fast and quiet both exist. Defense has to handle both.
A SOC is built to investigate what the stack emits.
A hunt asks what the stack missed: telemetry, detections, controls, identity paths, and assumptions that do not hold in production.
Finding an adversary is only one outcome. Finding nothing counts only if the evidence could have shown something.
Threat hunting is not a better dashboard.
It is how you find out whether your program can see the adversary behaviors that matter in your environment, including the ones your tools never surfaced.
The AI SOC question is not whether machines replace Tier 1.
It is what you can safely delegate when machines investigate and act at machine speed and humans still own the consequences.
An agent that summarizes malware is assistance. An agent that isolates endpoints or disables identities is delegated organizational authority.
How much of that authority have you actually defined?
The Department of War put ChatGPT Mil on https://t.co/5Z2onUvXZR for more than 3 million people, accredited for CUI at IL5.
That is secure AI adoption. It is not governed autonomous execution.
The harder line is whether the system can change consequential state with no transaction-level human authorization.
Secure access and delegated authority are related. They are not the same problem.
Where is your organization drawing that line?
The technical plumbing for runtime agent governance is advancing.
OpenID CAEP + Shared Signals improve continuous state propagation.
SPIFFE/SPIRE and current IETF agent-auth drafts give short-lived, operation-scoped identity.
Microsoft’s Agent Governance Toolkit and Dapr 1.18 Verifiable Execution put policy and signed provenance into the execution path.
None of them measures whether human oversight capacity still exceeds demand as agent volume rises.
That remains the least mature control.
It is also where progressive autonomy will break first under real load.
What have you observed when oversight volume climbs?
The most important AI security problem is not the model.
It is delegated authority.
Attackers no longer need to break the model. They compromise one trusted link (identity, memory, MCP server, dependency, session, or API) and let legitimate authority finish the job.
GTG-1002, STARDUST CHOLLIMA, SesameOp, and the Terraform MCP flaws all demonstrate the same architectural lesson.
Observed ≠ Demonstrated ≠ Projected.
What authority have we actually delegated, under whose intent, with what evidence, and can we stop or reverse the outcome?
Your AI supply chain is not being attacked where most teams are looking.
Threat actors are moving through the trust and authority your developers, pipelines, models, identities, and agents already have.
Delegated authority traveling through inherited trust.
Every arrow in the chain is an attack path.
Three control outcomes. One attack path.
What authority do your dependencies inherit, and can you prove they still deserve it?
Autonomous AI business operations should expand only as fast as the business can recover from being wrong.
Model performance is only half the decision.
The other half is: what happens when the system is wrong and changes business state?
Four recovery classes. Three rules. One test that should actually drive the autonomy decision.
Full decision aid in the image.
What is the highest-consequence action your autonomous processes can already take? Have you timed a real recovery exercise against it?
Stop measuring vulnerabilities processed.
Start measuring credible attack paths closed.
NIST’s April 15 NVD update made the old prioritization model weaker. When enrichment is no longer universal, local evidence of runtime presence, reachability, exploitability, existing controls, and business consequence becomes the only reliable signal.
Before you patch, establish reality.
Then break the path by the fastest effective means, patch, privilege reduction, segmentation, configuration change, or compensating control, and independently verify the route is non-viable.
A ticket marked “closed” is a claim.
An accepted risk with no owner, expiry, or compensating control is exposure moved sideways.
The metric that matters: validated critical attack paths closed.