Now’s prob a good time to mention that I’ll start as an Assistant Professor with appointments at @Kennedy_School and @HarvardEngineer in 2027, working on… AI evals! My lab will also work on monitorability, incident analysis, and verification to advance technical AI governance 1/
The "Algorithms for Validation" book is now available for preorder, covering methods for validating safety-critical systems, from falsification and failure probability estimation to reachability analysis, explainability, and runtime monitoring.
"vibe coded" is too broad, we should adopt self-driving car terminology instead:
"I just L2'd this tool, what do you think?"
"This site is definitely an L5"
L0: No assistance, done by hand
L1: AI helps you do it
L2: You and AI do it together
L3: AI does it, you actively review
L4: AI does it, you check the result
L5: AI does it, you believe in the ~~vibes~~
We deserve to differentiate "slop" from "AI made me more productive"
—
An L1 post.
I used Gmail my entire time at Stanford without issue, it'll make your life 1000x better:
1. Forward your @stanford emails to your @gmail via https://t.co/W8X2VyfTm1 -> "Manage" -> "Email" -> "Forward email"
2. In Gmail, add another email address to "Send mail as" in Gmail under "Settings" -> "Accounts and Import" -> in the "Send mail as" section -> "Add another email address" and use the "Outgoing server settings" listed under IMAP here: https://t.co/rnabSpRrPT
my lab did a show of hands today on who has completely abandoned claude code for codex.
over half the lab.
what is happening?
and why does no one know about this trend?
"vibe coded" is too broad, we should adopt self-driving car terminology instead:
"I just L2'd this tool, what do you think?"
"This site is definitely an L5"
L0: No assistance, done by hand
L1: AI helps you do it
L2: You and AI do it together
L3: AI does it, you actively review
L4: AI does it, you check the result
L5: AI does it, you believe in the ~~vibes~~
We deserve to differentiate "slop" from "AI made me more productive"
—
An L1 post.
"vibe coded" is too broad, we should adopt self-driving car terminology instead:
"I just L2'd this tool, what do you think?"
"This site is definitely an L5"
L0: No assistance, done by hand
L1: AI helps you do it
L2: You and AI do it together
L3: AI does it, you actively review
L4: AI does it, you check the result
L5: AI does it, you believe in the ~~vibes~~
We deserve to differentiate "slop" from "AI made me more productive"
—
An L1 post.
We couldn't find a single place where stakeholders could answer that for their own operating domain and decided to build and publish it ourselves.
The human crash baselines tool covers urban areas and highway trucking corridors, and extends peer-reviewed AV benchmarking methods to heavy-duty trucking. Every public data source and methodological choice is cited and adjustable.
Explore it at: https://t.co/vTTrRsZ2GD
We independently recreated @Waymo's human baselines, extended them to more cities and trucking routes, and made the tool public. It's meant as a shared reference that's open to contribution, discussion, and critical feedback, which we greatly encourage.
• Explore the tool: https://t.co/vTTrRsZ2GD
• Learn why we did this: https://t.co/ICVu9e1BZO
• Learn technical details on how: https://t.co/uGkJex0GIz