Agents trained on population-scale data can do a lot of things.
But ask a professional to stake their reputation on an AI-generated artifact? “Pretty good” isn’t good enough.
We introduce TAHI: a Test-time Adaptive agent framework through Human-agent Interaction. Featuring:
📈 Test-time adaptation via context (memory, skills) and weight training
⚡ Efficient adaptation to individual expertise within tens of tasks
✔️ Creating comprehensive rubrics for “non-verifiable” tasks
🔍 Analysis of shared community guidelines vs. personalized tacit expertise
We’re releasing a broad range of new mathematical results produced by an internal frontier model.
We’ve been consulting with the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, and we have drawn on their advice and public recommendations to inform how we release these results.
https://t.co/7N6TPlft1P
Claude now works inside Google Docs, Sheets, and Slides, and those files also open inside Claude.
In Google Workspace, Claude sits in a sidebar next to your file, reads what you have open, and edits it in place. You can approve each edit before it lands.
I got quite a few CHI ’27 review invitations, and after reading the submissions, I started wondering if HCI is getting less interesting. Too many papers now feel like fancy LLM wrappers. I attended three top HCI conferences recently, and purely in terms of how interesting the papers felt to me: CHI ’25 > DIS ’26 > CHI ’26. Where are we going?
I study AI systems for human goals. For complex tasks involving sensemaking, evidence gathering, and decision-making, a label alone tells us very little. Humans and AI can reach the same label through very different processes. With test-time scaling, maybe labels are no longer enough. Should we start annotating the process too?
I study AI systems for human goals. For complex tasks involving sensemaking, evidence gathering, and decision-making, a label alone tells us very little. Humans and AI can reach the same label through very different processes. With test-time scaling, maybe labels are no longer enough. Should we start annotating the process too?