@reidhoffman said back in the day that building a startup is like assembling a plane while you're falling off a cliff.
In the AI version, a lab can swap your engine out mid-fall, and the physical laws you're falling through can change at the same time. Gravity near Earth's surface holds at roughly 9.8 meters per second squared, the same on Tuesday as it was on Monday, which is why you can design a plane against it.
Imagine that number changing overnight because the law itself moved. A plane built against physics that no longer applies isn't the right thing to keep building, no matter how well it was designed the first time.
This is going to keep being the reality for founders and enterprise buyers alike, though it hits enterprises hardest, since they're the ones built to wait for stability before committing anything real.
Knowing which parts of your stack quietly assume today's version of the reality, whether that's a specific model's quirks, this quarter's cost curve, or a capability that didn't exist six months ago, is going to matter a lot.
Ironically, the fix isn't waiting for the rules to settle, since they won't. It's building the part of the operation that can play under any set of rules it gets handed, the team, the process, the judgment that decides what to do next, regardless of which model or capability happens to be available that month. That's the part you want stable.
There's no last mile in enterprise AI.
The industry treats deploying agents as a one-off, built, handed off, and considered finished. It doesn't matter whether that handoff comes from FDEs or your own team, since it's the same misreading of reality either way.
The workflow you built for keeps changing, so the agent has to keep adapting long after the build is finished, or it goes quietly wrong, unnoticed.
We run on two loops for this. One catches the exception the moment it happens. The other lets the agent revise itself as the workflow moves. The knowledge stays inside the enterprise, not with whoever built it.
Where I'd add a layer: the systems thinking in point 1 isn't just an engineering or design skill right now, it's becoming an operations skill. The people managing agents in a real back office need the exact same zoom-out reflex Elizabeth describes. Not "did the agent get this case right" but "does this whole workflow still make sense, and where does it break as the business shifts." That's a genuinely new muscle for most operators, and we are not yet training for it.
What I take from Mira's article is really about what work means for humans. To me, it's the thing a team figures out together. The satisfaction of a chef creating a new recipe after the tenth try. A group of people finally cracking the one workflow that always broke on the weird exception. For a lot of us, that is work.
So when we talk about building AI that does all of it for us, I think we skip the actual question. Not can the model do the task. But if it does, what happens to the part of us that got better by doing it ourselves. Practice and repetition make us sharp, not watching.
I believe humans are problem-solvers. We're at our best when we're challenged, not when we're handed the answer.
A year ago we set out to empower humanity with a focus on multimodal AI, custom models, and open science.
We previewed interaction models that collaborate the way people do. Tinker lets anyone train their own open weights models. We published our research on Connectionism.
Who supports this in three months?
The build very impressive, and 'the workflow is the unit of automation' makes sense. But these 16 workflows will keep changing, and the pods that built them have moved on. You've also put ~30 of your best engineers on back office automation that isn't Uber's core, and that math only gets heavier going from 16 functions to dozens.
If you don't give the non-technical users a way to take the training wheels off and run these themselves, you're going to run out of wheels. And engineers.
Very impressive result from Weco. An agent that rewrote its own harness and beat two years of hand-tuning in eight days.
The reason this is harder in an enterprise is what you're improving against. Weco has a fixed benchmark that sits still while you climb it. A regulated workflow never sits still. Policy shifts, new edge cases show up, the ground truth moves, and there's no clean score to chase. Just the work, and an auditor who checks every case, not the average.
The first experimental evidence of recursive self-improvement (RSI).
Autoresearching the autoresearch agent for eight days.
The result beats the harness we hand-tuned for two years, on held-out benchmarks: 🧵(1/7)
The first experimental evidence of recursive self-improvement (RSI).
Autoresearching the autoresearch agent for eight days.
The result beats the harness we hand-tuned for two years, on held-out benchmarks: 🧵(1/7)
Customer support has a clean unit to price against. An incoming email, a minute of a call. The back office behind it tends to be light, get an answer and update a record, so you can reason about cost per interaction fairly directly.
Core business operations doesn't behave that way, for a few reasons.
First, it isn't one thing. It's hundreds of distinct workflows, from AP to contract validation to AR and cash application, loan approvals, credit memo generation, expense approval. Each has its own shape.
Second, it varies by company, and honestly often by department within the same company. There isn't a standard unit to anchor to.
Third, the complexity isn't fixed. It varies across workflows and grows over time as the business does.
The practical consequence is that solving one workflow tells you very little about the cost of the next. There's no clean read-through. What tends to happen in that situation is people default to running on tokens, and costs drift upward with no ceiling.
Clear sense forming that a real solution has to have a real hold on cost. That means owning the harness rather than renting it, being able to reason about the tradeoff between performance, speed, and cost for a given task, and running across model types, frontier, open source, LLMs, SMLs, so the business gets clear alternatives.
Still early, but this feels like one of the harder unsolved problems in the space, and the one that quietly decides whether a lot of these deployments are economically viable at scale.
Big agree on the flywheel. The loop where the system gets better from its own interactions is the part that actually matters here.
I see the framing differently though. For most enterprises the real bottleneck isn't open vs closed, at least not yet. It's that everyone treats this as a last-mile problem and nobody owns what happens after deployment. Frontier models still perform better right now, so use them, build a system that keeps agents accurate as the business shifts underneath them, and optimize model and price later once you can see your distribution.
Enterprises don't have a model problem. They have a 'who maintains this in month six' problem. The agents that learn from their own exceptions are the only ones still running a year later.
Context is the hard battle, and the applied layer is where most of it happens. Governed domain knowledge, model routing, purpose-built post-training. I agree with most of it.
But when we say "whichever companies can solve that completely in an end-to-end fashion will have the greatest moats," this can't be FDEs. The change management of the workflow is not a single moment.
Keeping it accurate as the workflow changes takes forever, and that's where most of the value and most of the difficulty actually sit. Whoever owns that continuous part, not just the install, is the one who's hard to displace. Trying to solve that with FDEs, the math won't work.
FDEs solve the last mile: getting to production. But a regulated workflow isn't a last-mile problem, it's a road with no end. AWS knows it, which is why half the post promises you won't be stranded after the team leaves: runbooks, knowledge graphs, trained champions. A runbook is a photograph of a process that's already moving.
And it's not just AWS. OpenAI and Anthropic stood up their own deployment ventures this year. The whole industry is betting deployment is the constraint. It isn't. The constraint is what happens after, when the work changes and the build doesn't.
We built Reindeer for the thousand miles, not the last one. No team that hands you a system and rolls off, but an agent that carries the road itself and changes when the work does. The last mile is a human problem. The rest has to be the system's, or it was never really solved.
The government let Mythos back online, for small-picked US companies tied to critical infrastructure. Fable's still dark for everyone else. That partial reopening is the part worth sitting with, more than the shutdown two weeks ago.
It shows you the new normal. Frontier access is now something you're granted, based on who you are and where you sit. OpenAI shipped new models the same day and agreed to keep them to a small group of trusted partners.
We built Reindeer model agnostic, on our own harness, from the start, because we always figured whatever model we're on today is temporary. None of that was a prediction about export controls. The government just keeps proving the point for us.
למה כתבנו* Harness משלנו? 🧵
בתחילת** הדרך בריינדיר הרצנו את כל ה workflows של הלקוחות על גבי ה sdk של קלוד קוד. זה עבד יופי, אבל עם הסקייל התחלנו להיתקל בבעיות.
*קיסטמנו את https://t.co/iKaDZKhDnT
**ממש בהתחלה קראנו ישירות ל API של המודלים אבל די מהר הבנו שהארנס זה קריטי.
מסכים עם כמעט הכל פה, אבל מרגיש לי שהדיון הזה מקדים את זמנו בחצי שנה לפחות.
זה נכון שלפני שלושה שבועות ביום אחד כולם עברו מtokenmaxing לדבר על כלכלת טוקנים, אבל אתן קצת מידע פנים ממה שאנחנו רואים ב@Reindeer_AI :
1. לפני שנה חששנו שרוב האגנטים לא יהיו רווחיים, לפני חצי שנה בפועל זה ממש לא היה issue, והיום אכן כן רואים עלויות creeping up
2. אנחנו עדיין בגישה של מקס ערך ללקוח, נדאג לזה אח״כ. האתגר האמיתי באנטרפרייז הוא כרגע לא העלות (למרות שגם אתגר, הROI על כוח עבודה לא קל להם) אלא בעיקר האימוץ והתרבות הארגונית, מודל חלש יותר יקשה על זה ולפשל שם זה יכול להיות הזדמנות שלא חוזרת עם צוות בארגון.
3. זה פחות המודל עצמו, הכוונה שברמת הskill אנחנו עדיין יכולים לקרוא למודלים מאד חזקים, אבל העניין האמיתי הוא הharness שהפך להיות בזבזני ונוטה ללכת לטיולים ארוכים. לא יודע אם זה כי אנתרופיק מפתחים אותו בקלוד קוד ויש שם המון סלופ, שזה dealer mentality שמאפשר להם לשרוף טוקנים בצורה לא שקופה או סיבה אחרת, אבל לפחות אצלנו שם הסיפור
4. אנחנו עברנו לharness שלנו בחלק גדול מהמקרים. מעבר לשליטה במחיר, זה גם שיפור ביצועים בזמן בצורה מדהימה. אם יש עניין אפשר לכתוב על זה עוד. לדעתי בתור אצל @yairwein :)
5. בסוף גם השוק ידאג למחיר. שילוב של חוק מור, אופן סורס שיתפתח, יעילות שלנו וערך ללקוח
1/18 אני חושב לאחרונה הרבה על מה קורה עם Open-Source AI models, ואיך זה משפיע על סטארטאפים. חשבתי לשתף כמה מהמחשבות כ-VC בשלב הסיד. וגם, איזה dark patterns חדשים די בטוח מחכים לנו מעבר לפינה? למה בנצ'מרקס לביצועי AI זה חסר סיכוי? ואיך הזווית הגיאופוליטית משפיעה על Open Source AI?
למה כתבנו* Harness משלנו? 🧵
בתחילת** הדרך בריינדיר הרצנו את כל ה workflows של הלקוחות על גבי ה sdk של קלוד קוד. זה עבד יופי, אבל עם הסקייל התחלנו להיתקל בבעיות.
*קיסטמנו את https://t.co/iKaDZKhDnT
**ממש בהתחלה קראנו ישירות ל API של המודלים אבל די מהר הבנו שהארנס זה קריטי.
@NetanelBollag@Reindeer_AI@yairwein מחכים לך גם פה!
הבסיס הוא pi עם קוסטומיזציה, אבל ליאיר יש סיפור יפה על איך הוא בילה טיסה לוושינגטון כדי לבנות את הגרסא הראשונה של זה לקראת דמו ללקוח