Editing a multi-agent plan is not the same as steering it.
We introduce ๐๐ ๐๐๐ฃ๐ข๐ : ๐ฎ ๐ฝ๐ฟ๐ผ๐๐ผ๐๐๐ฝ๐ฒ ๐ณ๐ผ๐ฟ ๐ต๐๐บ๐ฎ๐ป-๐๐๐ ๐ฐ๐ผ-๐ฝ๐น๐ฎ๐ป๐ป๐ถ๐ป๐ด ๐๐ต๐ฎ๐ ๐ฒ๐ป๐ฎ๐ฏ๐น๐ฒ๐ ๐๐๐ฒ๐ฟ๐ ๐๐ผ ๐ฒ๐ฑ๐ถ๐ ๐ฎ๐ด๐ฒ๐ป๐ ๐ฝ๐น๐ฎ๐ป๐ ๐๐ถ๐ฎ ๐ฐ๐ต๐ฎ๐, ๐๐ฎ๐ฟ๐ด๐ฒ๐๐ฒ๐ฑ ๐ณ๐ฒ๐ฒ๐ฑ๐ฏ๐ฎ๐ฐ๐ธ, ๐ด๐ฟ๐ฎ๐ฝ๐ต ๐บ๐ฎ๐ป๐ถ๐ฝ๐๐น๐ฎ๐๐ถ๐ผ๐ป, ๐ฎ๐ป๐ฑ ๐บ๐ฒ๐ฟ๐ด๐ฒ, ๐๐ฝ๐น๐ถ๐, ๐ฎ๐ป๐ฑ ๐ฟ๐ฒ๐ฝ๐น๐ฎ๐ป ๐ผ๐ฝ๐ฒ๐ฟ๐ฎ๐๐ถ๐ผ๐ป๐.
The key finding:
More controls reduce some effort, but they also create new verification work.
Users built hybrid workflows rather than choosing one interaction type.
Global feedback was most robust for complex revisions. Targeted feedback preserved unchanged parts of the plan, but could introduce structural errors. Edit sequences were easier to audit, but less reliable for complex changes.
For agentic AI, the design challenge is to โhelp users know whether the edited plan still works.โ
๐ Read the article: https://t.co/rXTzxVAHp2
@estevamhruschka
#AI #Agents #LLMs #HumanCenteredAI #CAIS2026
Weโre hiring a Fall/Winter PhD Research Intern at Megagon Labs.
Weโre looking for candidates with research or technical experience in any one or more of the following areas:
โข Retrieval-Augmented Generation (RAG)
โข Function Calling
โข Document Understanding / PDF Parsing
You do not need experience across all three areas. Strong experience in any one of these topics is relevant.
Interested in joining us this fall? Learn more and apply โ https://t.co/N9xgO2NjSf
Applicants must have or be able to obtain appropriate U.S. work authorization for the internship period.
#AI #NLP #LLMs #internships #phdposition
Correct answers are no longer enough. As LLMs become part of systems that use tools, memory, and evolving knowledge, we also need to understand how those components change model behavior.
That question shaped much of the conversation at ACL 2026.
Read the article!
https://t.co/UKGs5zL39u
#AI #LLMs #CompoundAI
Your AI agent can cite a source. That doesn't mean the source says what the row claims.
Here's the failure mode we kept seeing: an LLM curator builds a plausible structured table from parametric memory; it "just knows" what should go in the cell, then bolts on a page-level citation afterward to make it look sourced. The citation is real, and the row is not grounded in it. It's structured output with decorative provenance.
We built Stage-Audit to close that gap in the context of what we call Seed2Frontier discovery: starting from a single Wikipedia seed page and autonomously finding the additional pages needed to complete a structured table that answers a natural-language query.
https://t.co/KCaCPmJS3z
#AIResearch #AgenticAI #CompoundAI #LLM #DataManagement #AIAgents
Most LLM extraction systems forget what they learned the moment the document ends. At ACL 2026, we introduce DySECT: a self-evolving extraction system that builds an explicit KB as it extracts.
This is a glance at our loop:
extract triples -> update KB -> organize and score knowledge -> feed it back into extraction
On DocRED, KB-guided extraction improved recall across GPT-4.1, GPT-4.1-mini, LLaMA-3.3 70B, and Kimi K2.5.
The key result: 5-8% recall gains on the first guided iteration, without fine-tuning or synthetic data generation.
Adaptation can happen through inspectable memory, not only opaque model updates.
๐Read the article: https://t.co/987Q2Gnkaz
#ACL2026 #NLProc #LLM #InformationExtraction #EnterpriseAI #KnowledgeGraphs #AIResearch #MegagonLabs #ContinuousLearning
As AI agents become more capable, understanding how they make decisions is becoming just as important as the answers they produce.
In our latest blog, we explore why agent observability is emerging as a critical capability for enterprise AI. We also introduce L.A.K.E. (Logic Agent for Knowledge Extraction), our agentic data planning framework that makes planning visible through executable workflows and interactive DAGs, helping developers inspect, evaluate, and improve how AI agents reason across complex data sources.
Transparency isn't just about explainability; it enables debugging, optimization, and building more reliable AI systems.
Read our blog post: https://t.co/9yVv2Ebkzf
#AI #EnterpriseAI #AgenticAI #MachineLearning #Transparency #MegagonLabs
Most ๐๐๐ ๐ฒ๐ ๐๐ฟ๐ฎ๐ฐ๐๐ถ๐ผ๐ป ๐๐๐๐๐ฒ๐บ๐ treat each document in isolation, failing to retain and reuse knowledge acquired from previous extractions. That is a problem in domains where terminology changes, taxonomies evolve, and rare concepts matter.
At ACL 2026 in San Diego, we introduce ๐๐๐ฆ๐๐๐ง: ๐ฎ ๐๐๐ป๐ฎ๐บ๐ถ๐ฐ ๐ฆ๐ฒ๐น๐ณ-๐๐๐ผ๐น๐๐ถ๐ป๐ด ๐๐ ๐๐ฟ๐ฎ๐ฐ๐๐ถ๐ผ๐ป ๐ฎ๐ป๐ฑ ๐๐๐ฟ๐ฎ๐๐ถ๐ผ๐ป ๐ง๐ผ๐ผ๐น๐ธ๐ถ๐.
Extraction should create reusable knowledge. So, DySECT uses an LLM to extract structured triples, stores them in a growing knowledge base, and assigns confidence based on source reliability, repeated evidence, and conflicts.
This accumulated knowledge is now a resource that should improve subsequent extraction rounds. The KB enriches previous extractions through hierarchical concept abstractions, confidence calibration, mutual exclusion constraints, and automated property validation. This enriched knowledge is then fed back into the extractor via prompt augmentation, encouraging broader coverage and discovery of new information in later iterations.
On DocRED, KB-guided extraction improved recall across GPT-4.1, GPT-4.1-mini, LLaMA-3.3 70B, and Kimi K2.5. The first KB-guided iteration ๐ถ๐บ๐ฝ๐ฟ๐ผ๐๐ฒ๐ฑ ๐ฟ๐ฒ๐ฐ๐ฎ๐น๐น ๐ฏ๐ ๐ฑ-๐ด% ๐๐ถ๐๐ต๐ผ๐๐ ๐ณ๐ถ๐ป๐ฒ-๐๐๐ป๐ถ๐ป๐ด ๐ผ๐ฟ ๐๐๐ป๐๐ต๐ฒ๐๐ถ๐ฐ ๐ฑ๐ฎ๐๐ฎ ๐ด๐ฒ๐ป๐ฒ๐ฟ๐ฎ๐๐ถ๐ผ๐ป.
Adaptive extraction does not have to mean opaque model updates. DySECT keeps accumulated knowledge explicit, inspectable, and editable.
For enterprise AI systems, that matters. Knowledge can evolve, but oversight and transparency stay in the loop.
๐ Read the research: https://t.co/6T7ykVTDmw
Work by @moinnas, Hannah Kim, @estevamhruschka Hruschka
#ACL2026 #continuouslearning #AI #LLMs
As LLM systems become more agentic, improving individual outputs may not be enough.
In a recent Megagon Labs Seminar Series talk, Jay-Yoon Lee (Seoul National University / UIUC) explored the idea that correctness and reliability often emerge at the set level. Modern LLM systems increasingly operate over collections of retrieved documents, generated statements, and intermediate reasoning steps, yet most frameworks still treat these elements independently.
Interactions such as contradictions, redundancy, and joint relevance can shape outcomes just as much as the information itself. The talk introduced a unified perspective centered on set-level verification, retrieval, and editing as important directions for building more reliable and controllable LLM systems.
Thank you to Jay-Yoon for sharing his insights.
#LLM #AgenticAI #AIResearch
The way AI agents plan over enterprise data can dramatically change both latency and reliability.
This process of data planning, how an agent sequences operations like database queries, document retrievals, and API calls to resolve complex prompts, is foundational to enterprise AI performance.
In our latest research, we found that different planning regimes produce very different trade-offs:
- Single-shot tree planning achieved the strongest overall performance and lowest latency.
- Iterative planning provided stronger self-correction at a significantly higher execution cost.
We also observed that the planning structure itself, including generation direction and refinement strategy, can substantially impact agent behavior.
These findings highlight a broader challenge in agentic AI: you cannot optimize or trust what you cannot inspect.
As agents increasingly orchestrate tools across databases, documents, and external systems, understanding how plans are generated and executed becomes critical for debugging, evaluation, and deployment.
To address this, we introduce L.A.K.E. (Logic Agent for Knowledge Extraction), an agentic data planning framework designed to make planning observable. L.A.K.E. exposes execution paths as interactive DAGs with step-level provenance, enabling developers to inspect workflows, compare planning strategies, and understand why agents succeed or fail.
This work reflects a broader direction we are exploring at Megagon Labs: building agentic systems that are not only capable of reasoning over complex enterprise data, but also transparent, debuggable, and efficient to operate.
#Agentic #NLProc #LLMs
https://t.co/K8seDSN74z
Can LLMs automate prompt engineering?
RPT says: Apparently yesโby turning prompt tuning into a function-calling workflow: evaluate, diagnose recurring failures, use memory, and revise.
The result: stronger prompts on reasoning tasks and better-calibrated confidence.
๐ข New Preprint
Prompting is a flexible way to adapt LLMs, but prompt engineering is still manual and hard to scale. Can function-calling LLMs act like prompt engineers: inspect failures, diagnose patterns, and revise prompts?
Meet Reflective Prompt Tuning (RPT): a diagnosis-driven framework for iterative prompt tuning with function-calling LLMs.
Paper: https://t.co/wNczUx32KM
@MegagonLabs
1/n
๐ช๐ฒ ๐ณ๐ผ๐๐ป๐ฑ ๐๐ต๐ฎ๐ ๐๐๐ ๐ฎ๐ด๐ฒ๐ป๐๐ ๐ฐ๐ฎ๐ป ๐บ๐ฎ๐๐ฐ๐ต ๐๐ต๐ฒ ๐ฎ๐ฐ๐ฐ๐๐ฟ๐ฎ๐ฐ๐ ๐ผ๐ณ ๐ฅ๐ฒ๐๐ฐ๐-๐๐๐๐น๐ฒ ๐ฎ๐ด๐ฒ๐ป๐๐ ๐๐ต๐ถ๐น๐ฒ ๐๐๐ถ๐ป๐ด ๐ฎโ๐ฏ๐ ๐ณ๐ฒ๐๐ฒ๐ฟ ๐๐ผ๐ธ๐ฒ๐ป๐.
In our latest research, we revisit a core assumption in agentic AI: that agents must interleave reasoning and execution at every step to remain adaptable.
We compare single-step horizon planning against full-horizon planning for data-centric tasks such as Knowledge Base QA and Multi-hop QA.
Our experiments show that full-horizon planning with lazy replanning achieves accuracy parity with step-wise approaches across varying task depths, branching complexity, and tool robustness levels, while substantially reducing token usage.
In many cases, eager execution monitoring provided little benefit relative to its computational overhead.
These findings suggest that planning horizon is not merely an implementation choice, but a fundamental systems design decision for LLM agents.
For well-defined data-centric tasks, generating complete plans upfront and replanning only when necessary may offer a more efficient default for tool-using AI systems.
This work will be presented at CAIS 2026
๐ Read the research: https://t.co/CjPme0X3mm
#AI #AgenticAI #LLMs #CAIS
๐๐ด๐ฒ๐ป๐๐ ๐ฐ๐ฎ๐ป ๐น๐ฒ๐ฎ๐ฟ๐ป ๐ฐ๐ผ๐ป๐๐ถ๐ป๐๐ผ๐๐๐น๐ ๐ณ๐ฟ๐ผ๐บ ๐น๐ฎ๐ฏ๐ฒ๐น๐ฒ๐ฑ ๐ณ๐ฒ๐ฒ๐ฑ๐ฏ๐ฎ๐ฐ๐ธ, ๐๐ถ๐๐ต๐ผ๐๐ ๐ฝ๐ฎ๐ฟ๐ฎ๐บ๐ฒ๐๐ฒ๐ฟ ๐๐ฝ๐ฑ๐ฎ๐๐ฒ๐. ๐ง๐ต๐ถ๐ ๐ถ๐บ๐ฝ๐ฟ๐ผ๐๐ฒ๐ ๐ฎ๐ฐ๐ฐ๐๐ฟ๐ฎ๐ฐ๐ ๐๐ต๐ถ๐น๐ฒ ๐ฐ๐๐๐๐ถ๐ป๐ด ๐ฟ๐ฒ๐ฎ๐๐ผ๐ป๐ถ๐ป๐ด ๐ฐ๐ผ๐๐๐.
Our approach borrows from human cognition: agents build two kinds of memory from past examples:
๐๐ฝ๐ถ๐๐ผ๐ฑ๐ถ๐ฐ: instance-level critiques of what went right or wrong on specific cases
๐ฆ๐ฒ๐บ๐ฎ๐ป๐๐ถ๐ฐ: task-level patterns distilled from those critiques
At inference time, the agent retrieves relevant memories and uses them to guide its decision.
Across 7 datasets and 6 models (GPT-5, GPT-4o-mini, GPT-oss-20b, Llama-4-Scout, and two Qwen3-235B variants), the combined strategy delivered:
โ +8.1 pp over zero-shot
โ +4.6 pp over RAG few-shot baselines
โ 31.95% fewer thinking tokens for reasoning models compared to RAG few-shot.
Pre-computed critiques substitute for reasoning the model would otherwise do at inference, resulting in both accuracy gains and lower per-query costs, with significant implications for agents deployed in production.
We also introduce suggestibility, a metric for how receptive a model is to in-context reasoning. It helps explain why some models gain dramatically from memory augmentation while others barely move, and it raises questions about the tradeoff between learnability and adversarial robustness.
๐ง Code & datasets: https://t.co/IPl07dtht6
Read the paper: https://t.co/yz7C3yStjP
Research by Jackson Hassell, Dan Zhang, Hannah Kim, @tommmitchell, and @estevamhruschka โ presented at ACM CAIS '26.
#AI #LLM #AIAgents #MachineLearning
Where is AI actually headed? #EACL2026 highlighted a shift toward agentic systems, stronger evaluation, and more reliable LLMs.
From multi-agent workflows and tool use to deeper evaluations of RAG, reasoning, and personalization, the field is moving beyond isolated model improvements toward systems that operate effectively in complex, real-world environments.
Read the article to learn about the major trends across agentic AI, LLMs, and RAG.
๐ https://t.co/68eOYh7YS0
#AI #LLM #AI #AgenticAI @estevamhruschka
๐๐ก๐ฒ ๐๐จ ๐ญ๐ก๐ ๐ฌ๐๐ฆ๐ ๐ฉ๐ซ๐จ๐ฆ๐ฉ๐ญ ๐๐ง๐ ๐๐๐ญ๐ ๐ฉ๐ซ๐จ๐๐ฎ๐๐ ๐ฐ๐ข๐ฅ๐๐ฅ๐ฒ ๐๐ข๐๐๐๐ซ๐๐ง๐ญ ๐ซ๐๐ฌ๐ฎ๐ฅ๐ญ๐ฌ ๐๐๐ซ๐จ๐ฌ๐ฌ ๐๐ ๐ฌ๐ฒ๐ฌ๐ญ๐๐ฆ๐ฌ? ๐๐จ๐ฌ๐ญ ๐๐ฅ๐๐ฆ๐ ๐ฆ๐จ๐๐๐ฅ ๐ฌ๐ข๐ณ๐ ๐จ๐ซ ๐ญ๐ซ๐๐ข๐ง๐ข๐ง๐ ๐๐๐ญ๐ ๐๐ฎ๐ญ ๐จ๐ฎ๐ซ ๐ฅ๐๐ญ๐๐ฌ๐ญ ๐ซ๐๐ฌ๐๐๐ซ๐๐ก ๐ซ๐๐ฏ๐๐๐ฅ๐ฌ ๐ญ๐ก๐ ๐ซ๐๐๐ฅ ๐๐ฎ๐ฅ๐ฉ๐ซ๐ข๐ญ: ๐๐๐ญ๐ ๐ซ๐๐ฉ๐ซ๐๐ฌ๐๐ง๐ญ๐๐ญ๐ข๐จ๐ง.
We ran a controlled study, keeping content identical while changing only format. The results challenge how we think about building compound AI systems.
Key finding:
โข SQL systems dropped 30-45% in performance on the same data when representation changed.
โข LLMs showed the opposite pattern, with performance increasing.
โข Hybrid systems? We found something completely different.
For teams building RAG systems, multi-agent workflows, or agentic AI, this isn't just an academic insightโit's a critical design decision most are overlooking.
Representation is part of the computation, not just preprocessing.
๐๐๐๐ ๐ฐ๐ก๐ฒ ๐ญ๐ก๐ข๐ฌ ๐ฆ๐๐ญ๐ญ๐๐ซ๐ฌ ๐๐จ๐ซ ๐ฒ๐จ๐ฎ๐ซ ๐๐ ๐ฌ๐ฒ๐ฌ๐ญ๐๐ฆ๐ฌ: https://t.co/xBBrg0QMNV
By Megagon Labs and @SAYg_7
#AI #CompoundAI #LLM #MachineLearning #AIEngineering #Research
๐ขNew Preprint๐ข
LLMs can solve many tasks. But who verifies an answer when the judge is also an LLM?
We introduce AutoPyVerifier: learning compact Python verifier sets from labeled LLM outputs, improving objective prediction by up to +55 F1.
Paper: https://t.co/VwJKCHR81u 1/n
Welcoming @MegagonLabs as a Founding Gold Sponsor of CAIS 2026. Megagon is an AI research lab building at the intersection of NLP, data management, and LLM agents.
Find them in the Bayshore Foyer May 26-29. https://t.co/MxpkxyQuwm
Starting an AI research internship soon? ๐จโ๐ป๐ฉโ๐ป
We put together 7 tips for AI research interns on how to:
- hit the ground running
- adapt when research shifts
- build relationships that last beyond the internship
Read the full guide here โ https://t.co/JjQH1eBEhp
#PhDLife #AIResearch #ML #NLP #LLM
Attending #ICLR 2026 in Brazil? Learn how data representation affects LLM response, with a poster presentation on ๐ฆ๐ฎ๐บ๐ฒ ๐๐ผ๐ป๐๐ฒ๐ป๐, ๐๐ถ๐ณ๐ณ๐ฒ๐ฟ๐ฒ๐ป๐ ๐ฅ๐ฒ๐ฝ๐ฟ๐ฒ๐๐ฒ๐ป๐๐ฎ๐๐ถ๐ผ๐ป๐: ๐ ๐๐ผ๐ป๐๐ฟ๐ผ๐น๐น๐ฒ๐ฑ ๐ฆ๐๐๐ฑ๐ ๐ณ๐ผ๐ฟ ๐ง๐ฎ๐ฏ๏ฟฝ๏ฟฝ๏ฟฝ๏ฟฝ๐ฒ ๐ค๐ with @estevamhruschka
Not attending? Read the paper here: https://t.co/7x5beVWi1P
Paper by Yue Zhang, @SAYg_7, @nikita_bhutani
#AI #LLMs #Research #PhD