Our paper on influence functions titled Better TDA via Better Inverse Hessian-Vector Products will be presented at #neurips25 San Diego on Friday from 4:30 PM - 7:30 PM in Exhibit hall C,D,E (Poster 3907) Work with my fantastic colleagues @_elinguyen, @Runshi_Yang , @juhan_bae, @SheilaMcIlraith, @RogerGrosse!
Proud to introduce Llemma, the first models trained on OpenWebMath (part of the 55B mathematical tokens in ProofPile II).
Llemma serves as a platform for future research on quantitative reasoning and is a very powerful base model for mathematical tasks!
Should you let LMs control your email? terminal? bank account? or even your smart home?🤔
🔥Introducing ToolEmu for identifying risks associated with LM agents at scale!
🛠️Featuring LM-emulation of tools & automated realistic risk detection
🚨GPT4 is risky in 40% of our cases!
Identifying the Risks of LM Agents with an LM-Emulated Sandbox
Presents a framework that uses a LM to emulate tool execution and enables scalable testing of LM agents against a diverse range of tools and scenarios
proj: https://t.co/6iVW6ONkYH
abs: https://t.co/t7VQVT7VhX
This paper was a central contribution of @RToroIcarte 's PhD thesis at U of T, and done in collaboration with Toryn Klassen, @RickValenzano, and supervisor, @SheilaMcIlraith. Early results were first published at ICML18 and AAMAS18. We are proud to have received this recognition.
Meet STEVE-1, an instructable generative model for Minecraft. STEVE-1 follows both text and visual instructions and acts on raw pixel inputs with keyboard and mouse controls. Best of all - it only cost $60 to train!
w/ @Shalev_lif@SirrahChan@jimmybajimmyba@SheilaMcIlraith
Early LLM induced multibillion dollar casualty. What is the value of text-based content, if it can be paraphrased and accessed for free?
https://t.co/YQ5CQr1NNV
Excited to share Boosted Prompt Ensembles! Inspired by boosting algorithms, we propose an algorithm that automatically grows an ensembles of prompts to cover a target problem space.
Led by @silviupitis and with @andrew_wang10 and @jimmybajimmyba
What are boosted prompt ensembles, and how can they be used?
New paper released yesterday significantly improves output quality through automating n-shot learning
Paper breakdown with useful takeaways 📃🧵👇
https://t.co/c1igJYqPB4
we are starting our rollout of ChatGPT plugins.
you can install plugins to help with a wide variety of tasks. we are excited to see what developers create!
https://t.co/NQ684Yp2LK