I recently started a podcast on AI and "taste".
Link: https://t.co/no1WyvCPDu
Also on Spotify and Apple Podcast.
With my co-host Philipp Zahn, we explore how the next generation of AI might develop the ability to judge what is good, choose what matters, and create work worth caring about.
Broader project: https://t.co/6IxM6SmmiQ
(π§΅) Happy to release AIRS-Bench, a benchmark to test the autonomous machine learning abilities of AI research agents π€
AIRS-Bench includes 20 tasks sourced from machine learning papers that assess the autonomous research abilities of LLM agents throughout the full research lifecycle, from hypothesis generation π‘ and implementation π οΈ to experimentation π§ͺ and analysis π
Each task is extracted from a paper with a state-of-the-art result and consists of a:
π problem description (e.g. text similarity)
ποΈ a dataset (e.g. SICK) and
π a metric (e.g. Spearman correlation) to optimise over
The agent is then given a GPU and 24 hours to develop and submit a Python solution that matches or exceeds the paper SOTA π
Read on for baseline results and examples of agents surpassing human SOTA οΏ½οΏ½οΏ½
π±We open-source the AIRS-Bench task definitions and evaluation code to accelerate in autonomous scientific research:
π» GitHub: https://t.co/UXzNXyGdU5
π ArXiv: https://t.co/badN0jq0IA
π€ HF paper: https://t.co/6FIWxF0Bsw
π Meta AI website: https://t.co/wcIWLrlYBU
Huge shoutout to the team from Meta FAIR who painstakingly crafted, debugged and inspected every single of these tasks and its runs across more than a dozen of agents @alisia_lupidi, @_tomwithanh, @BhavulGauri, @basselralomari, @albertomariape, Alexis Audran-Reiss, Muna Aghamelu, Nicolas Baldwin, @LuciaCKun, @GagnonAudet, Chee Hau Leow, Sandra Lefdal, Abhinav Moudgil, Saba Nazir, Emanuel Tewolde, Isabel Urrego, @mahnerak, @ishitamed, @EdanToledo and @rybolos, @alex_h_miller, @j_foerst, @yorambac for their leadership and support
I also believe strongly in vibe-writing.
Curious, @thulme, did you just use a mainstream AI tool like Claude or ChatGPT (voice + Canvas) for this use case? Or did you feel the need for a more powerful interface?
Asking because we're building something in that space at https://t.co/3NA8PLeVFT, taking inspirations from Cursor/Windsurf to give writers more control.
Excited to unveil my latest project!
Over the last few months, my co-founder @FlynnDevine and I have been researching ways to make AI work better for thinkers and writers who care deeply about quality and intellectual rigor.
There's a lot of free AI chat applications out there, and these can be really helpful as a research tool, but when it comes to writing, a lot of people are underwhelmed by the bland and generic feedback they get.
Then there's a lot of specialised apps helping marketing teams generate large volumes of SEO-optimised blog posts. But these are often just slop generators, and you'd never want to use them to produce any serious piece of writing.
We're exploring a new angle: giving people a single space to think and write, with an AI assistant helping them (or just observing) at every step, to try and really understand the user's taste and intentions.
This is still early days, but we'd love your feedback on the direction we're exploring at @TryThoughtly , so we're launching a first website today with a quick demo video. Please see link in comment to sign up for early access!
@dabit3 Why considering only 3 when they are at least 50 options out there? ;) See https://t.co/IWXZFc94cz for a fairly exhaustive list (please make a PR if something's missing or inaccurate!)