Overhearing LLM Agents
Having AI agents proactively suggest and provide additional context can improve many workflows.
This is an exciting direction on AI agents that I've been exploring for the last couple of months, mostly for note-taking and research.
Good morning. Stay in bed today. Early reports suggest Crowdstrike Falcon - a computer threat checker used by lots (and lots and lots) of businesses pushed out an update that might have broken a lot of computers. Airlines, businesses etc affected
🚨New Paper🚨: Are AI text detectors *really* as good as they claim? (#ACL2024)
We release RAID—The largest & most challenging detection benchmark with 6M+ outputs from 11 LLMs, 8 domains, 4 decoding strategies, and 11 adv attacks
https://t.co/mPlNymdchJ
https://t.co/VR0mMvqXRm
Have you ever done a dense grid search over neural network hyperparameters? Like a *really dense* grid search? It looks like this (!!). Blueish colors correspond to hyperparameters for which training converges, redish colors to hyperparameters for which training diverges.
We are pleased to announce that the first Conference on Language Modeling will be held at the University of Pennsylvania in Philadelphia at the Zellerbach Theatre.
Thanks so much to UPenn CS as well as Mark Yatskar and Zachary Ives for facilitating the amazing venue.
The PEARLS Lab at @UCSD_CSE is now open for business! I'm recruiting Fall 24 PhD students in all things interactive and grounded AI, RL, and NLP!! Join us in the land of 🏖️ beach (🧋pearl tea included). Apply by Dec 20. Please help spread the word!
More: https://t.co/nPZT3KbuUC
Excited to announce that Kani has officially been accepted and will appear at the NLP Open Source Software Workshop at #EMNLP2023!
See you all in Singapore :)
✨New Paper✨: We release Kani 🦀 a highly-hackable open-source library for building LM apps with tool usage (e.g. plugins)
Kani lets you easily write LM-callable functions in pure Python w/ robust type checks + model retry
https://t.co/CfqvZTboFo
https://t.co/NmsjD5fI1Q
✨New Paper📷: FIREBALL! is the largest dataset containing structured game states from real Dungeons and Dragons games, capturing 8 million utterances and 1.3 million game states. https://t.co/sKGAJLjh0f