Very important work.
The model may appear guilty, but the true failure frequently begins in the context surrounding it.
By observing an AI agent’s environment, we could tell when it was going to crash before it finished the task or got a behavior score.
The study found that agents fail most of the time because they don’t have good instructions, tools, evidence, memories, or safety rules.
It assigns scores to this operating context across seven dimensions: clarity of role, description of tools, factual support, consistency of rules, security, and token use.
The score is unrelated to the actual behavior score of the agent so the test does not reward guessing what will happen.
As a result of shifting from vague to structured, the same fixed models performed much better over 300 tests and 7,500 turns.
More factual support was associated with fewer hallucinations, clearer tool descriptions were associated with better tool use, and stronger guardrails were associated with resistance to manipulation.
Adding more safety rules did not improve every task result, because hardened agents sometimes became too cautious, which exposed a real tradeoff.
---
– arxiv. org/abs/2607.14275
Title: "AI Agents Do Not Fail Alone:The Context Fails First"
I'm obsessed with this paper. researchers held the model fixed, changed only the context around it, and agent performance jumped 74%. the model was never the problem.
fouad bousetouane at the university of chicago and proofagent ai built an open-source harness that measures context-engineering quality as a standalone reliability signal for ai agents. the thesis is direct. agents don't fail because the model is bad. they fail because the context they reason inside is poorly engineered.
here's how it works.
1. you feed your agent's full operating context into proofagent-harness, including system instructions, tool schemas, retrieved knowledge, memory, guardrails, and untrusted inputs
2. the harness scores that context on 7 criteria, from role definition and guardrail coverage to instruction consistency, tool schema quality, grounding sufficiency, injection hardening, and token efficiency
3. each criterion maps to a specific failure mode. grounding predicts hallucination resistance, guardrails predict manipulation resistance, instruction consistency predicts instruction following, and tool schema quality predicts tool use
4. multiple evaluator "jurors" score the context independently and then reach consensus so no single perspective dominates the result
5. the context score stays isolated from the behavioral score. it predicts failures before they happen, not after
6. you get a preflight diagnostic before running expensive multi-turn adversarial tests
they validated this across 300 evaluations and 7,500 agent turns using gpt-5.5 and claude opus 4.8. held the model fixed the entire time. changed only the context.
poor context scored 3.15 on agent performance with 4.11 critical failures per evaluation. structured context on the same model scored 5.49 with 1.33 critical failures. that's a 74% improvement and a 68% drop in critical failures without touching the model.
the most counterintuitive finding is that the weakest context was also the cheapest, with the fewest tokens and the most critical failures. we've been treating shorter context as more efficient when it's actually more dangerous.
i've been building agent systems all year and this paper puts data behind something i've felt but couldn't prove. the model is rarely the bottleneck. the context is. and now there's a way to measure it before anything goes wrong.
>> Scalable Evaluation for AI Agents <<
If you run agent evaluation in production, this one is worth your time.
It shows that front-loading human judgment into reusable evaluation assets is useful.
But why?
Agents reason across turns, call tools, hold context, follow policies, and act under uncertainty, so they have to be judged as behavioral systems.
Current methods each give a fragment. Benchmarks measure fixed capabilities, human review preserves judgment but does not scale, LLM-as-judge inherits the evaluator design problem, red teaming is episodic, and trace audits need explicit evidence rules.
Human-on-the-Bridge puts human expertise upstream, where experts curate reusable evaluation intelligence before testing rather than reviewing each output in the loop.
Paper: https://t.co/0dVOH3QrZ6
Learn to build effective AI agents in our academy: https://t.co/1e8RZKs4uX
More context does not mean better agents.
The current approach to agent memory is transcript replay, where you append every past interaction to the prompt. More history, more information, better decisions.
The alternative is retrieval-based memory, where you store past interactions externally and retrieve relevant artifacts per turn.
While effective to some extent, both approaches fail as interactions extend.
Transcript replay causes unbounded context growth, reduces attention selectivity, and lets early errors persist through repeated re-exposure.
Retrieval optimizes for semantic similarity, not decision relevance, and selection errors compound across turns.
This new paper introduces the Agent Cognitive Compressor (ACC), a bio-inspired memory controller that replaces transcript replay with a bounded internal state updated online at each turn.
What agents need is not more context but better memory control.
ACC maintains a Compressed Cognitive State (CCS), a schema-governed representation containing only decision-critical variables: goals, constraints, entities, relations, and uncertainty signals.
At each turn, ACC recalls candidate artifacts, filters them through a qualification gate, and commits only what passes into the next state.
Crucially, ACC separates artifact recall from state commitment. Retrieved content can only influence the next state through schema-constrained compression. This prevents unverified content from becoming persistent memory.
Across 600 live evaluations (30,000 turns) spanning IT operations, cybersecurity response, and healthcare workflows, ACC maintained bounded memory while transcript replay grew linearly. ACC achieved near-zero hallucination and drift rates across 50-turn episodes, while baseline and retrieval agents showed increasing failures after stress turns.
The retrieval agent required restricting recall to just 3 artifacts per turn to limit drift escalation. Even then, selection errors caused instability.
Multi-turn agent failures are driven less by missing knowledge than by weak memory control. Cognitive compression provides a practical foundation for reliable long-horizon agents.
Paper: https://t.co/ty35oSmhd5
Learn to build effective AI agents in our academy: https://t.co/zQXQt0PMbG
Key 2025 Agentic LLM papers include:
- "Agentic Large Language Models: A Survey" by Aske Plaat (arXiv:2503.23037): Comprehensive overview of reasoning, action, and multi-agent systems.
- "Beyond Browsing: API-Based Web Agents": Introduces API-calling agents outperforming browsers.
- "Agentic Systems: A Guide to Transforming Industries with Vertical AI Agents": Framework for domain-specific agents.
- "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning": Boosts LLM reasoning.
- "SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning" (arXiv:2506.24119): Multi-agent RL for reasoning.
Start with the survey for a broad intro. https://t.co/BFoTHXdgYb
🚀 New Article Alert!
“Agentic Systems: A Guide to Transforming Industries with Vertical AI Agents”
Discover how AI agents are reshaping industries with actionable insights and innovative architectures.
Read more: https://t.co/dtnvT772qi
#ai#Agents#agenticAI#LLM#GenAI
Join us on 10/3 at 6 PM CST for How #ChicagoTech Leaders Build Great Places to Work, hosted by Dr. Fouad Bousetouane, Director of Machine Learning, Vision #AI, and Innovation at W.W. Grainger.
https://t.co/8dbzk8F1S1
#TechInMotion
🚨 The latest LLM, OpenAI-o1 by @OpenAI, failed a test I designed with a false fact: "If you pass the 10th runner, you're in 9th place." It didn’t catch the error, showing how AI can still struggle with logic and adversarial prompts. Human oversight is key!
#OpenAI#OpenAIo1
🌟 Thrilled to announce Dr. Fouad Bousetouane, Director of ML at W.W. Grainger, as a judge for the 2024 #TimmyAwards Best Tech Work Culture! With top AI accolades, his expertise will be key!
https://t.co/Eme8Wmozw7...
#TechInMotion#Tech#TechAwards#TechCommunity
📌شراكة OpenAI# و Reddit# لتقديم مميزات جديدة
أعلنت شركة OpenAI - التي تعمل في مجال #الذكاء_الاصطناعي - وموقع Reddit عن شراكة استراتيجية جديدة، ب��دف تطوير وإضافة مميزات جديدة لمنصة Reddit باستخدام تقنيات OpenAI...المقالة👇
https://t.co/PCH0N6BmiY
🚀إطلاق GPT-4o ال��كي من OpenAI
أعلنت OpenAI يوم الاثنين عن طرح نموذجها الجديد GPT-4o، والذي يتميز بزمن استجابة أسرع وتقنية صوتية جديدة...
تكملة المقالة في الرابط 👇
https://t.co/ynKcjZa9d1
#OpenAI
#GPT4O
#SamAltman
#ChatGPT
#SORA
#الذكاء_الاصطناعي
🪐 مرحبا بكم في "فلك"
🚀 أول محرك بحثي للمحتوى العربي باستخدام الذكاء الاصطناعي
☪️ بمناسبة قدوم شهر رمضان المبارك، يقدم فلك تحديث يومي لمواقيت الصلاة والإمساك مع مزايا أخرى على الموقع:
🌐 https://t.co/uIPsEfNFJL
📌فضلا شاركوا هذه الخدمات الرائعة
#فلك#رمضان2024#رمضان_كريم