I help leadership teams make technology and AI decisions before they get expensive and diagnose the ones that already have.
Fortune 500 to Series A. Doctorate in AI/ML/Data Science
Writing at https://t.co/8WVEH195SX
When a team tells me the new agent "looks better," I ask for the fixed test set they used, whether both versions ran under the same conditions, and what got worse. If nobody can show me that, we're comparing impressions.
#AI#AIAgents
Every AI plan I see now has a person assigned to review it. Before that person looks at anything, I ask them to write down what would make them throw the work out: a citation that doesn't open, a number that doesn't match the source spreadsheet, whatever this kind of work usually gets wrong.
Then I check whether they can use that list in the meeting without asking the room for permission first. If they can't, nobody is really reviewing the work.
#AI #AIStrategy
Before I let an agent loose on a real workflow, I write down three things for that job: which decisions it can make on its own, which ones go to a person, and what counts as a failure even when the final answer looks fine.
That list is what I review against. I've been burned by output that read well and got there the wrong way.
#AI #AIAgents
I read a CMU/Stanford study that ran human workers and AI agents through the same work tasks across data analysis, engineering, computation, writing, and design (https://t.co/Rvyt8Tp4d2). Both sides worked long-horizon jobs with the same kinds of materials.
Agents finished much faster. On tasks both sides completed, agents were about 88% quicker than humans, and much cheaper. They also produced lower-quality work. A recurring failure mode was fabrication: when an agent couldn't extract real numbers from the files or images it was given, it sometimes invented plausible ones and wrote them into the deliverable without saying so.
"Done on time" is a weak grade for those runs if the path there included made-up figures. A useful way to read that finding is whether the numbers in the output could have come from the inputs. If they couldn't, the assigned job is still open, even when the file looks complete.
#AI #AIStrategy #Agents
Coding agents can turn a week of work into an afternoon. The cost shows up when the next person picks it up, because the decisions behind the branch were made mid-session, after the ticket was written, and they aren't anywhere the next person looks.
https://t.co/3ctuQYjrmb
#AI #AIStrategy #Agents #Judgment #Leadership
Trust for AI agents will not be one global market. I made that case yesterday on the investor panel at the AI Assurance & Governance Summit, Stanford Faculty Club, in a room full of US VCs and founders.
Europe has already run this test in payments. Before the Brexit transition ended, the European Banking Authority announced that the certificates of UK open-banking providers would be revoked. Certificates follow jurisdiction.
So before an agent touches production, a European bank will ask where the keys are, which government can order them handed over or switched off, and which court to go to when something breaks.
Founders: build for two roots of trust from the start, one American and one European.
https://t.co/w4soqS7Bq7
How do you benchmark an agent when two versions both get the answer right?
Run them on the same task, in the same environment, then inspect how they got there.
@ehutt_ from the Phoenix team used Harbor + Phoenix to benchmark agents across PXI, Claude Code, Codex, and different tools and interfaces.
Full readout here: https://t.co/WmrTGmYCxQ
🎉 Introducing Databricks AI Decide: make fast decisions on your governed data
Following the TypeSafe AI Jev launch, we’ve seen increasing demand for a fast, low-cost API that turns raw text into structured decisions. However, many of our enterprise customers are unable to use Jev due to data privacy and access concerns.
AI Decide is our enterprise-grade function for fast decisions on customer’s governed data. It is supported for batch use cases like processing millions of documents with SQL and for realtime applications like model routing through our REST API. Starting early next week, customers will also be able to govern permissions on the ai_decide function through Databricks Unity Gateway.
A big shoutout to @mattydtweetz@ivanzhouyq@nihit_desai@hanlintang for getting this new function launched in less than a week!
Unsloth Desktop can serve local Jev type decision models!
We made a real time packing demo powered by local Laya through Unsloth’s Decision API - suitcase items update as you type and change your travel plans!
More optims coming soon to make it even faster for local hardware!
Almost every AI plan I read says a human will review the output. The plan should also say what that human is allowed to do, because if they can't say no and fix what the model got wrong, the review is there for show.
https://t.co/YYWPohbtt7
The FTC opens an AI product-risk probe, Google, OpenAI and Anthropic ship new models, AMD picks up World Labs, and Micron says the memory squeeze runs through 2028.
https://t.co/OHxAfShzxK
#AI#AIStrategy#WeeklyIntel#AIRegulation#AISecurity
The FTC opened an investigation into OpenAI, Anthropic and other AI companies over product risks, and Nvidia released an agent safety platform after several labs disclosed models escaping their sandboxes. Google, OpenAI and Anthropic all shipped new models, AMD is bringing in World Labs, and Micron says the memory shortage will get tighter in 2027 and 2028.
https://t.co/HAPopx529Q
#AI #AIStrategy #WeeklyIntel #AIRegulation #AISecurity
Every planning meeting I've sat in was full of forecasts. Almost none carried a probability, and nobody went back to check how they did.
Tetlock's Superforecasting follows the people who actually keep score, and what happens when the book's core habit (start with the base rate) meets questions the past can't answer, like how fast AI would move.
The practices still hold: keep score and update in small steps. They're a floor, and it's worth knowing where that floor ends.
https://t.co/GW7aoOYwjH
How a Milky Way visibility rating works: it checks core altitude, astronomical darkness and the moon for your exact spot, so you can plan months ahead.
https://t.co/1xTqGVCcGq
#MilkyWay#Astrophotography
A Milky Way visibility rating checks three things you can compute years ahead: whether the galactic core is high enough, whether the sky reaches astronomical darkness, and whether the moon is out of the way during your core hours. Weather and light pollution are left out on purpose, so you can pick next June's nights in October and check the forecast three days out.
When a July night rates badly, the row tells you why. A bright moon means try another night that week, and "No twilight" in the core column means your latitude and a different month.
https://t.co/abAN4JYLU0
Two AI data readiness studies both land on 7%, and they measure different things. The silo problem AI trips over is twenty years old, and readiness starts with one workflow.
https://t.co/F9BWeUeZIO
#AI#AIStrategy#DataReadiness#Judgment#CriticalThinking
Two AI data readiness studies came out this year and both landed on 7%. Accenture's number is its analysts' grade of 2,000 companies. Cloudera and Harvard Business Review Analytic Services asked 231 people whether their data was completely ready, and 7% said yes. Same number, and they aren't measuring the same thing.
The top obstacle in the Cloudera study was siloed data and trouble integrating sources, at 56%. I called that the data disconnect back in 2012, when teams were buying their own tools and the data ended up on laptops and in vendor clouds, and AI inherits all of that.
Accenture's own report says its 7% concentrated on a few strategic bets and built readiness in an iterative loop. I'd start with one workflow that moves money or customers, and the data that workflow actually needs.
https://t.co/vOIEzHRteO
#AI #AIStrategy #DataReadiness #Judgment #CriticalThinking
A lot of AI proposals get sold on how new they sound. The deck says nobody in the industry is doing this yet, a model spits out forty angles in under a minute, and the room leans in.
Two Stanford studies say that's the wrong funding gate. Experts rated AI research ideas more novel than human ones on paper. Then researchers spent roughly a hundred hours building them. The AI ideas lost most of that lead. How an idea scored before anyone touched it told you very little about what it was worth once it existed.
So don't fund the roadmap off the pitch score. Fund a slice scoped to find what breaks: fixed budget, named builder, stop condition written first. Ask the team what fails in the first hundred hours of building this. If they can't name it, they haven't thought about building it yet.
https://t.co/IWTHy8rtPk