For more details, please see our paper and demo video:
๐ Paper: https://t.co/JIs3XQYDSd
๐ฅ Demo video: https://t.co/oVAoZ2ZTSA
Thanks to my co-authors Parke Godfrey, Lukasz Golab, Divesh Srivastava, and Jarek Szlichta!
#AI#LLM#DevTools#Debugging#Agents
Last week at #EDBT2025, I presented our latest demo, LADYBUG โ an interactive debugger for LLM agents ๐
LLM agents are powerful, but when they fail, debugging can be a nightmare. LADYBUG helps you trace agent steps, intervene, and even repair issues with LLM assistance.
๐งต๐
LADYBUG is like a software debugger, but built for LLM agents.
๐ Step-by-step execution traces
๐ ๏ธ Intervene at any step & re-run only whatโs affected
๐ค LLM-powered self-reflection to find & fix errors
๐ Inspect intermediate state & metrics in complex workflows
๐As we ring in the New Year, check out our 2024 roundup story!
๐From revolutionizing healthcare to unravelling the secrets of AI, #UWaterloo's Computer Science is pushing the boundaries of human curiosity.
๐Read more: https://t.co/TwmmN8JUqQ
Had a great time chatting with @WCDSBInnovates on their Deep Learning Dialogues podcast! We discussed my recent LLM-focused #ExplainableAI research and its implications for AI trust and safety. Check it out! ๐๏ธ
New episode! ๐ง "RAGE: Bringing Clarity to the Fuzzy World of AI Sources" featuring @JoelExplainsAI from @UWaterloo uncovers how the RAGE tool demystifies LLM sources, reshaping AI trust in education. Listen now on all podcast platforms or our website! #WCDSBInnovates
A team of #UWaterloo researchers have created a new tool โ nicknamed โRAGEโ โ that reveals where large language models (LLMs) like ChatGPT are getting their information and whether that information can be trusted.
More: https://t.co/awA2evmxoe | #UWaterlooNews
7/ ๐ฎ There are many aspects of LLMs, RAG, and beyond that remain unexplained. We're currently researching many of these, with several extensions in the works. Stay tuned for more updates and an open-source RAGE code release in the near future! ๐๐ก
6/ ๐งโ๐ฌ This work is a collaboration between @UWaterloo, @YorkUniversity, and @ATT. Huge thanks to my co-authors for their incredible contributions! ๐
Ever wonder how exactly your knowledge sources are being used during RAG?
We did too. I'm excited to share our early work in explaining LLMs, "RAGE Against the Machine: Retrieval-Augmented LLM Explanations"! ๐ธโจ
๐ Paper: https://t.co/xQ7QhQmv37
๐งต [Thread]
4/ ๐ Whether it's determining the greatest tennis player or answering your own unique questions, RAGE supports various use cases. Our code (coming soon) is designed with interfaces, making it highly adaptable to any LLM, retriever, or text data source! ๐พ๐
3/ ๐ Key features include:
- Source Combination Tests: Which sources lead to which answers? ๐
- Source Permutation Tests: How does the answer change if sources are reordered? ๐
- Interactive Demo: Explore behaviors of real LLMs using real data. ๐
2/ ๐ ๏ธ RAGE helps users understand how LLMs generate answers by identifying parts of the input context that, when moved or removed, change the LLM's response. This counterfactual approach makes AI decision-making more transparent, and allows us to derive a form of citation. ๐๐
1/ ๐ Our paper introduces RAGE, an interactive tool designed to explain the outputs of LLMs augmented with retrieval capabilities (RAG). RAG needs explaining since it is unclear how the presence and order of knowledge sources (among other things) affects the LLM's answer! ๐ค๐ก
Being able to interpret an #ML modelโs hidden representations is key to understanding its behavior. Today we introduce Patchscopes, an approach that trains #LLMs to provide natural language explanations of their own hidden representations. Learn more โ https://t.co/WfY1FYa1Wt