The economics of AI should not become a race to produce the most dramatic estimate.
We need careful measurement before making sweeping claims about the future of work.
The goal is not to dismiss AI risk.
AI will absolutely reshape parts of the labor market.
The goal is to improve how we measure and communicate those risks.
The goal is not to dismiss AI risk.
AI will absolutely reshape parts of the labor market.
The goal is to improve how we measure and communicate those risks.
In our recent research, different AI models produced dramatically different estimates of which occupations are exposed to AI.
That means estimates of “job risk” may depend heavily on which model generated the rating.
In our recent research, different AI models produced dramatically different estimates of which occupations are exposed to AI.
That means estimates of “job risk” may depend heavily on which model generated the rating.
Headlines about “AI replacing jobs” often rely on occupational exposure estimates generated by large language models.
But most people never see how uncertain those estimates actually are.
Headlines about “AI replacing jobs” often rely on occupational exposure estimates generated by large language models.
But most people never see how uncertain those estimates actually are.
New @IZAWorldofLabor opinion piece:
“Workers Deserve More Honest Estimates of AI Job Risk”
https://t.co/QBfUeWRpeG
A short thread on why public conversations about AI and jobs need more humility and transparency.
New @IZAWorldofLabor opinion piece:
“Workers Deserve More Honest Estimates of AI Job Risk”
https://t.co/QBfUeWRpeG
A short thread on why public conversations about AI and jobs need more humility and transparency.
@voxeu@NorthwesternU Different models can produce very different estimates of occupational AI exposure — even under the same framework.
That means policy conclusions may sometimes depend heavily on the model generating the ratings.
@voxeu@NorthwesternU But we probably need to pay much closer attention to:
how these measures are constructed
how stable they are across models
what assumptions are embedded in the ratings
how uncertainty propagates into downstream conclusions
The point is NOT that we should stop using these methods.
LLM-based occupational exposure measures are incredibly valuable and have pushed the literature forward in important ways.
But we probably need to pay much closer attention to:
how these measures are constructed
how stable they are across models
what assumptions are embedded in the ratings
how uncertainty propagates into downstream conclusions
But we probably need to pay much closer attention to:
how these measures are constructed
how stable they are across models
what assumptions are embedded in the ratings
how uncertainty propagates into downstream conclusions
@voxeu@NorthwesternU Thank you for the opportunity to contribute to the conversation. As AI increasingly shapes research and policy discussions, I hope this article encourages more attention to the transparency, reproducibility, and uncertainty behind AI-generated occupational exposure estimates.
Our @VoxEU piece discusses why different large language models can generate very different occupational AI exposure estimates — and why that matters for downstream economic conclusions and policy discussions.
When the ruler is made of the thing it measures: Multi-model evidence on AI occupational exposure scores
Michelle Yin @NorthwesternU
https://t.co/Ye0ppVVlso
When the ruler is made of the thing it measures: Multi-model evidence on AI occupational exposure scores
Michelle Yin @NorthwesternU
https://t.co/Ye0ppVVlso
New Wall Street Journal piece on paper number one of a series. Same task list, four AI models. The share of US occupations flagged "high AI exposure" runs from 14% under one model to 51% under another, on identical content. A 19-fold spread. Thread. https://t.co/XwKbbOhYei
Two more papers and several policy briefs on this topic are coming. But if there is one takeaway right now, it is what I told the Wall Street Journal: “I personally would not rely on just one measure to say, ‘Oh, I should change my job,’ or ‘I should change my kid’s major.’”
I care about this because I have sat across the table from workers whose career decisions depend on what researchers like me put into the world. If those numbers are not credible, we are failing the people we are supposed to serve.
I care about this because I have sat across the table from workers whose career decisions depend on what researchers like me put into the world. If those numbers are not credible, we are failing the people we are supposed to serve.
AI companies are profit-driven and their models reflect their training data, and their design choices. Researchers who use these proprietary outputs as scientific instruments have an obligation to verify what they produce before families, communities, and governments act on it.
AI companies are profit-driven and their models reflect their training data, and their design choices. Researchers who use these proprietary outputs as scientific instruments have an obligation to verify what they produce before families, communities, and governments act on it.
When I discovered the instability in AI exposure scores, I was working with workforce programs in Maine and Virginia trying to help real people navigate a changing labor market, and I realized the numbers we were relying on gave fundamentally different answers!
When I discovered the instability in AI exposure scores, I was working with workforce programs in Maine and Virginia trying to help real people navigate a changing labor market, and I realized the numbers we were relying on gave fundamentally different answers!
I want to share why this research is personal to me, not just professional. I came to this country as an immigrant and built my career as a labor economist because I believe how we measure work shapes how we value workers.