We're entering a new era of AI agents, but who is asking workers what they actually want?
We studied how current AI development aligns with what workers need, identifying gaps and areas of opportunity. 👇
🚨 70 million US workers are about to face their biggest workplace transmission due to AI agents. But nobody asks them what they want.
While AI races to automate everything, we took a different approach: auditing what workers want vs. what AI can do across the US workforce.🧵
Left NeurIPS feeling so inspired — especially by @YejinChoinka’s keynote. The future is socially intelligent AI: a world with adaptive collaboration and models that genuinely understand our intent
1/8 Pareto Frontier 🤠for Human-centered AI 📈: We all want to build AI that is good for humans, but the path is often paralyzed by complexity. Either “oh my god, it’s too complicated😱” or delusional “I have a warm and fuzzy feeling of understanding 🥴”? "It’s hard because it depends.🤷" is the enemy of progress. We need a Pareto Frontier for Human-centered AI. 🧵👇
Ever ask a coding agent for a fix, only to get something totally misaligned?
Zhou et. al introduces TOM-SWE, which models user mental state from interaction history and includes user satisfaction in evaluation!
Love to see agents being optimized for human-AI collaboration.
Hoping your coding agents could understand you and adapt to your preferences?
Meet TOM-SWE, our new framework for coding agents that don’t just write code, but model the user's mind persistently (ranging from general preferences to small details)
arxiv: https://t.co/5SmiplW7AZ
❓Motivation: Most coding agents today can plan, edit, run, and test code. But they still fail at a key part of real-world development, understanding the user! Underspecified, shifting, or context-dependent instructions can easily break them.
You must have those moments when coding agents were running for 10 minutes and ended up producing things largely misaligned. (1/)
Today's AI agents are optimized to complete tasks in one shot. But real-world tasks are iterative, with evolving goals that need collaboration with users.
We introduce collaborative effort scaling to evaluate how well agents work with people—not just complete tasks 🧵
Super exciting to see this MSR paper building on ideas from Future of Work with AI Agents! Love seeing more data-driven analysis on how people actually use AI at work.
New paper from my group at @MSFTResearch!
📄https://t.co/bRwk7auUAn
Promises about how AI will change work are cheap. What does the actual data say?
We measured which work activities people use AI for, how successful they are, and which jobs do those tasks. 🧵1/8
i've been thinking lately about how future ai systems will interact with us and how we can make systems that care about people and wanted to put words to it -- hopefully it resonates a bit!
Just wrapped up a research internship @MSFTResearch in the Machine Learning + Optimization Group!
Grateful to the MLO team for an incredible summer 🙌Our project was on improving LLMs for optimization problem-formulation...more coming soon:)
From @Diyi_Yang’s amazing talk at @MIT today!
The misalignment between investment in both startups/academia with Public needs is intriguing.
High desire but low capability for AI Agent research papers per task is the highest of all, potentially indicating the deployment gap.
41% of YC AI startups are solving tasks workers don't need automated
New Stanford study shows workers actually DO want AI, but for repetitive work that frees them up for higher value tasks
Startups are chasing full automation where partnership would work better
Some tasks are painful to do.
But some are fulfilling and fun.
How do they line up with the tasks that AI agents are set to automate?
Not that well, based on our new paper "Future of Work with AI Agents: Auditing Automation and Augmentation Potential across the U.S. Workforce"
1/2
From Future of Work with AI Agents: Auditing Automation and Augmentation Potential across the U.S. Workforce.
A link to the paper.
https://t.co/AuFOpKKkBy
Co-authored with @EchoShao8899 Humishka Zope, @YuchengJiang0, @jiaxin_pei, David Nguyen and @Diyi_Yang
Some interesting @ycombinator data too.