DEADLINE ANNOUNCED: Apply by April 26th to LASR Labs AI Safety research program. This is a 3-month paid opportunity to write a paper in a small team with expert supervision. Past work has been published at workshops/conferences. Previous alumni now at Open AI & UKAISI.
When will AI systems be able to carry out long projects independently?
In new research, we find a kind of “Moore’s Law for AI agents”: the length of tasks that AIs can do is doubling about every 7 months.
METR is running a pilot field experiment to measure how AI tools affect open source developer productivity.
If you're an open source developer who wants to make $150/hour to work on issues of your own choosing - consider expressing interest 👇
Can frontier models cost-effectively accelerate ML workloads via optimizing GPU kernels? Our take at METR: yes, and they’re improving pretty steeply – but it’s easy to miss these capabilities without good elicitation and “fair” compute spend.
How close are current AI agents to automating AI R&D? Our new ML research engineering benchmark (RE-Bench) addresses this question by directly comparing frontier models such as Claude 3.5 Sonnet and o1-preview with 50+ human experts on 7 challenging research engineering tasks.
How well can LLM agents complete diverse tasks compared to skilled humans? Our preliminary results indicate that our baseline agents based on several public models (Claude 3.5 Sonnet and GPT-4o) complete a proportion of tasks similar to what humans can do in ~30 minutes. 🧵
@elicitorg Maybe a bit too concise? I've noticed a tendency for the section summaries to mention the topics covered in those sections but not much of the views/arguments expressed by the authors on those topics.
@james_elicit Especially appreciated the nuance about optimising within the *known* landscape. Promoting the meme of safety as a "windfall" might be a surprisingly effective cognitive vaccine against risk-taking behaviour among practitioners.
@smfleming great talk at TCPW, and an exciting response to my question about the link with awareness/intrusive thoughts. Looking forward to exploring that connection further
Anyone looking for a broad and deep perspective on current computational neuroscience & its intersection with AI could do a lot worse than going through the entire back catalouge of @pgmid 's BrainInspired podcast https://t.co/lphViYmmS9
@dileeplearning@pgmid@ProfData Loved this talk (and the rerun on @neuromatch). Have been captivated since the CHMM paper at @CogCompNeuro. Thinking now about how "HC as inference on cloned cognitive graph" fits so well with "default mode network as MDP" from https://t.co/0fJkOSdiBN
It's rough, meandering, and overlong, but at least I finally finished getting my thoughts down in a somewhat intelligible form. Something I've been thinking about for a while: what does the mind process tells us about responsibility and justice?
https://t.co/shptPJLEL6
@MaryamShanechi@neuromatch@seanescola@Pascallisch@bgalbraith Loved your outro talk, and just finished reading your recent motor/mood BMI review paper. Feeling inspired! I'm about to start my PhD in simulation neuro and would love to get involved in this kind of research. Any advice?
The latest in my series of articles on the "mind process" is a practical guide to mental training and mindfulness. It's not woo, it's just physics!
https://t.co/zJBPzUf7CV