Wow, an out of control moving truck crashed into several vehicles today in Maryland today.
One of those vehicles was a @cybertruck, that ended up being split in half by the impact! The cabin is entirely intact, protecting the occupants, who only received minor injuries.
AI Mania Is Eviscerating Global Decision-Making | Nikhil Suresh, Hermit Tech
I strongly believe there are entire companies right now under heavy AI psychosis and it’s impossible to have rational conversations with it about them. I can’t name any specific people because they include personal friends I deeply respect, but I worry about how this plays out.
- Mitchell Hashimoto, of HashiCorp and Ghostty fame
Over the past year, I’ve run point on all of our company’s sales, led the technical components of all but two of our engagements, and over the lifetime of this blog have had something like 300 catchups with professionals from around the world. This has ranged from people on the ground in niche service industries to executives at Fortune 500 companies. Because of this, I've had a front-row view to our collective institutions across both the private and public sector undergoing breath-taking mass psychosis. This essay is an attempt to describe the bizarre dynamics that are currently at play, as I am in the rare position where my wellbeing is not contingent on paying lip service to madness, and to reassure the people trying to survive amidst all of this that they are not crazy.
The reality is thus: the people in charge either have no plan, or see no path forwards other than keeping their heads down. Not at banks, not at hospitals, not in our government institutions. The world’s organisations have been captured by people in the throes of frothing excitement, and saner people who now live in a state of constant commingled fear and frustration.
I. AI Investments Are Generally Total Failures
Are companies actually seeing massive productivity gains from their AI adoption? Does any of this sordid affair make sense?
This should be an easy question, but it is surprisingly hard to get a straight answer to it. Executives that tell the press that their company has gone insane will quickly find themselves removed from their positions. Employees who are honest will find themselves fired in short-order, or “randomly” selected for a round of layoffs. In fact, it is in the interests of almost every actor in the space – boards, executives, employees, vendors, consultants – to obfuscate and misrepresent the success rate of AI projects. Many publicly traded companies are putting out announcements about their AI productivity gains when I know for a fact that the businesses have done nothing other than purchase Copilot licenses and declare victory.
Yet we need to know if these projects are panning out – if the total focus on AI as a core tenet of business strategy is succeeding at a reasonable rate, then a discussion about the relative risk and reward is warranted.
Unfortunately, we live in a dark timeline. All of the AI projects we have observed as a team are failing. Every single one – we have seen 0% success in a year and a half, not only amongst projects we have been asked to participate in22 We have rejected all AI implementation work. It is absolutely a gigantic bubble and we have minimized our exposure to it – every single one of our current contracts would be totally unaffected by OpenAI collapsing, save for perhaps some second-order effects such a recession causing a client to become unable to pay us. And there's nothing we can do to insulate ourselves from that anyway., but even within projects that we have observed in passing while doing totally unrelated work. Even if you grant that AI tooling accelerates specific workloads, the method and scale of the current investments is senseless. Frequently the failure is not related to AI itself, but rather that companies are terminally bad at running software projects effectively, and as I have remarked previously, AI projects are subject to all the failure modes of normal projects plus you can get everything right and then still fail because of the method's novelty. Very few companies are so good at shipping software that they can afford the extra risk profile.
Often enough, though, it’s an actual failure in what LLMs can accomplish. The most common version of this, being rolled out across businesses around the world, is the internally-facing chatbot, or for the more daring company, the customer-facing chatbot. The story is always the same. For the former, I’ve never seen substantial internal uptake from inside a business. Employees don’t use internal chatbots because companies tend to have low-quality documentation and an LLM is not psychic – it can only know things that have been written down and made accessible. For the latter customer-facing applications, I have rarely had a pleasant experience as a consumer, with perhaps the exception of live transcription during medical appointments – hardly something worth pivoting an entire organisation around. In both cases, project leaders are very careful to avoid tracking basic metrics, such as whether the tools are being used at all, or they track metrics that are easily gamed.
For example, my last consumer interaction was attempting to get help from Mitsubishi following an automotive failure, where a very polite robot asked me to describe the problem and that I’d receive a call back as soon as someone was available. This was the single most competent implementation of such a project I’ve seen in the wild, in that the voice was natural sounding, responded quickly, was clearly “live” in production, and promised a swift resolution.
That was six months ago, and I did not, in fact, get a call back.
When Mitsubishi did not call me back, what happened? Did that request just go into the void, showing one less incident for the year? Does it appear that the phone bot resolved my query without the need for human intervention? All we know is that it didn’t show up as an error, or I’d have received a call. I’m sure it looks great in all sorts of ways except the one that matters, which is that I was planning to buy a car and decided not to buy another one of theirs.
For this reason, our team has quickly learned while on an engagement not to ask anything about ongoing AI projects in any context – by the time that project has started, it is too late for the management team, and intervention is not possible until a crisis point is inevitably reached. There is no conceivable positive outcome. The failure rate is so high that even basic inquiry leaves us in an untenable position. Any coherent question about how it’s going, what the goal is, who is using it, constitutes an inadvertent attack on the chain of command responsible for the work because there are no good answers to anything. Even in rare cases where my interlocutor has stated that things are going well (usually while the project is still mid-flight and failure has not had a chance to manifest), it is generally obvious that they are doomed, but at least in these cases I can simply agree and then go home to scream into a pillow for six hours straight.3
All of this is to say that I am very confident that almost every report at a company about “massive AI productivity gains” is untrue as a matter of brute fact. Even if some companies are seeing clear gains, this is the exception, not the norm. With that assumption in place, we can talk about the dynamics at play, and how it has become impossible for many organisations to stay focused on things that actually matter to their long-term (or even short-term) health.
II. Heretics Will Be Shot
It has become outright dangerous to even raise the possibility that AI might not be the solution to a problem, let alone be the sole focus of a company’s entire strategy.
In every sufficiently large business we have observed (say, with 500+ employees), we have noted that continued advancement, and increasingly continued employment, has started to require repeated professions of belief in the transformative power of AI for said business. I am not talking about providing ideas about how to use AI in the business – I mean religious profession, declarations of faith. Overwhelmingly these statements are made by non-technicians, though it is not uncommon for technicians to emit deranged statements to curry favour.
There have been several occasions where I have seen someone, apropos of nothing, blurt out almost word-for-word “AI is changing everything”, only to concede moments later that their organisation does not currently use LLMs for anything, and indeed, that they cannot name a single thing that has changed other than they get some use out of ChatGPT (frequently the free-tier). In one extreme case, I have seen an executive confess that they had never even used ChatGPT or any AI tool in their life, immediately after producing a technical strategy for an organisation with $2B+ in revenue which was entirely centered around AI.
Initially these statements were so absurd on their face that I thought it was some cynical ploy to achieve thought leader status, and there are certainly some people doing this – I have had it admitted to me. But the broader reality is so much worse: people who have no background in the technology at all actually believe what they are saying. As a general rule you should avoid getting into business with a liar, but if you must, you can at least reason with them even if only in private. A true believer is much more threatening because they are impervious to even inducement by self-interest.
The turning point in my belief was watching someone with a spectacular amount of money on the line fire their highest performers because they were achieving that performance without LLMs. When an employer publicly talks about AI innovation, we have to ask ourselves if they’re simply trying to manipulate the market or customers. When they privately commit to strategies like this with their own money at stake, with no attempt to communicate that strategy to external clients, I can only assume they really mean what they’re saying.
A while ago, I wrote “Contra Ptacek’s Terrible Article On AI”, which was focused on the fact that many of Ptacek’s points in his own essay “My AI Skeptic Friends Are All Nuts” were internally inconsistent.44 We have since kissed and made up in private, though I don't think we've budged at all on the core points of our viewpoints. I maintain that Thomas is a very talented writer with a lot of good advice who just happened to blow it massively that one time because he takes Hackernews commenters too seriously. We all have our weaknesses. Mine is people telling me that "Scrum is good if you do it right". But on the crux of the matter, we are actually in total agreement, because he opens his essay with this:
Tech execs are mandating LLM adoption. That’s bad strategy.
Which is to say that we can sidestep arguments about the precise utility of LLMs entirely and we’re left in a very simple place – it is entirely obvious to both myself and Ptacek, two people that are coming at this from fairly opposed views, that people are being really, really stupid about this, and that organisations are demanding bizarre workflow constraints from their specialist staff.5
These mandates have led to extremely strange places. Several of my peers now “AI-wash” their work, meaning that even when they can perfectly competently execute on their jobs to the satisfaction of their management teams, said managers are unhappy if the engineers haven’t used AI in the work… so now they’re lying about using LLMs even in contexts where their professional judgement is that they aren’t the appropriate tool. They just do the work, the same way they have for decades, and say Claude did it. Others are being measured on their AI bills with “token leaderboards”, where higher is better because I have evidently fallen into the pocket of Hell where the demons torment me by doing elaborate impressions of absolute fucking morons, so the people hired for their freakish ability to perform system optimisation do the obvious thing. They set the LLMs prompting themselves in a semi-plausible loop in case someone inspects the token consumption and then they watch Netflix. Not a single one has been caught, even when their own assessment of the output is that it isn’t suitable for deployment.
Checking out a parallel copy of our Go repository and telling the AI to rewrite the whole thing in Zig while I work on something else just so I can keep my job. I hate this shit so much. My job has usage tracking and quotas. I don’t use it for actual work, I just spin it up and disregard the output.
- An actual software engineer
In fact, the only people I know of to be fired over this whole thing are people that have expressed visible doubt about this organisational strategy, which again, even Ptacek thinks is transparently dumb. The net result is that everyone has learned very quickly to praise executives on their visionary AI prowess, or they will be gunned down in the proverbial streets.
III. AI Demos Are The Mind-Killer
Bless me, Father, for I have sinned. It has been ∞ days since my last confession. I accuse myself of the following sins:
One of the main pieces of infrastructure we deploy at our clients is an analytics-focused database called Snowflake – for a typical business, the bill is tiny because it’s a pay-as-you-go situation and we can process all their data in one minute a day, you get a very hands-off deployment, and in short it has many characteristics that are very pleasant for our work. One of the features in Snowflake that we don’t use is called Cortex.
Cortex is their AI chatbot layer, with the ability to plug into metadata (for non-nerds, descriptions of your data, like what a column in a spreadsheet means) and query a company’s database autonomously. In theory, you can ask a question like “What was our revenue for last week?” and it will spit out an answer.
It is not really suitable for production usage. From memory, the last time I was given a presentation on it, by actual Snowflake staff, they reported that ideal configuration results in something like ~92% accuracy due to the complexity of data at a large business (see: probably best-in-class for these tools, but imagine your CFO having one in every ten of their numbers be outright wrong) and there were serious issues with managing deployments. Nonetheless, it can be used to produce some very flashy demonstrations.
On several occasions, we’ve been exposed to folks that have been sort of lukewarm on our main offerings, but they really, really wanted to use AI to perform a natural language query on their data. And we thought “Okay, if you really want to see it, maybe we can caveat this appropriately and show you what it might look like.”
This was a terrible mistake. It backfired in the most predictable way imaginable – every lukewarm client that saw the chatbot in action, even with us telling them that it was not going to accomplish what they wanted, wanted to buy it immediately. Every other consideration, including millions of dollars that we could plausibly help them achieve by non-AI means, was swept aside. It was like a dark and terrible force seized control of their limbs, plunged their hands into their own chests, and presented their still-beating credit cards to us in grim supplication. We were so mortified by the inexplicable shift in energy that we (wisely) declined to take the money and ended the sales process, and soon thereafter removed Cortex from our list of demonstrations. It would have been too irresponsible to exploit this gap in their reasoning, and frankly, it was already irresponsible to have even run the demonstration – doctors don’t walk around showing off cool pills that they’d never prescribe.
Watching the total 180°, that shift from ice-cold to red-hot buying frenzy, was a deeply unsettling experience. It was personally uncomfortable to see people that clearly didn’t gel with us interpersonally suddenly dying to enter an ongoing relationship, but more broadly uncomfortable because for a brief moment I began to understand what is happening in sales meetings around the world. There was no warning I could have given that would have made them refuse to buy the damn thing – their appetite was as large as their budget could stretch, and some part of me wonders if this is because they knew that their ravenous hunger would be present in their own customers. They’d just buy it from us, then pivot right to a larger company and mind control their leadership team until the buck finally stops with the loser that needs to justify the expense. The main protection against this seems to be that the median vendor is so bad at their jobs that we had presented the first even somewhat-working products these people had seen, and this included an ASX-listed company that was already bragging about their AI usage. It took our team two hours to produce something that was frankly not that good – basically just typing text descriptions of data into a web browser – and it was still better than anything the leads had seen because they had nothing to show for all the investment.
In fact, we have been forced to opt out of every sale where the lead has expressed anything beyond the most fleeting curiosity in the use of AI in their business. I don’t mean that we’ve heard that they’re interested in AI and elected to drop the contract on moral grounds. I mean that, over the course of the engagement, these people have exhibited a pattern of behavior that has made it near-impossible to sell to them without incurring reputational and legal risk, and are furthermore crafting management environments that I can only describe as cultish, ineffective, and “please dear God, do not let it be on earth as it is on LinkedIn”.
IV. Executives, Game Theory, and The Emperor’s Clothes
The good news is, CISOs are used to having to protect the business from their hare-brained initiatives, and this one isn’t really that different, except that there’s a cult-like atmosphere to it that you didn’t see with, say, the cloud. It almost doesn’t matter whether you embrace the initiative or not; there’s work to be done to manage the risk, so that’s what you do. From talking to CISOs everywhere, I would say most of them are quietly skeptical but afraid to speak up.
- Career CISO and well-known speaker that asked to remain anonymous
Despite the substantial prevalence of true believers, many of the people running large AI initiatives, or making public statements about them, do not believe what they are saying. There are “heads of AI” who read this blog, at companies with $1B+ in annually recurring revenue, who have written in to say they believe their job is totally fraudulent but it was the only promotion pathway remaining at the organisation.
On a trip overseas, I had the privilege of a meeting with one of the Fortune 500 executives mentioned at the beginning of the post, who will remain anonymous so that they are not executed by firing squad by their board. As we were chatting, it became clear that they were very switched-on and technically competent, and they also happened to be at a company that had committed to the usual battery of exorbitant claims about their recent innovations – we’ve 100x’d our productivity, AI is the future of everything, I am but a vessel for OpenAI to make love to my wife. You know, normal things. But since I had them there without any microphones around, I asked why this was being repeated without opposition. Was it just sales fluff?
The answer was a lot more interesting. It was partially ridiculous sales material being delivered to an easily excitable audience, but this was not the dominant factor constraining honesty. Executives at their customers were saying absurd things about achieving 100x productivity, and this meant that if any executive at the vendor said that these gains were not plausible, it would undermine the credibility of the customer’s executive, be perceived as an attack (or heresy), and possibly result in an enterprise contract cancellation. And getting enterprise contracts cancelled because you wanted to opine on something that doesn’t really matter to your organisation’s mission is a great way to get fired.
But this company was also a major player, of the kind that signs enormous enterprise contracts with other companies. So presumably there is another vendor that has sold to them, and their CEO is worried that saying something sane will contradict this executive, and very quickly we can see how we can have executives around the world nervously pointing guns at each other, not wanting to be shot first but also watching everything gradually spiral out of control. This is to say that we’re facing a coordination problem around executives being honest around the AI gains they’ve witnessed – if they co-operate, they keep their jobs. If they defect, they will possibly be fired by their embarrassed peers (who have now been implicitly called liars, cowards, or incompetents) and then replaced with someone that will toe the line anyway. If they could all admit the truth at once there might be some hope, but there is no way to coordinate that event.
This sounds deeply concerning, but it is worth noting that it means that some executives who are emitting nonsensical statements are not as dull as they might seem at first – they’re in a fraught political environment, where they are surrounded by many people that are gunning for their roles, and subject to the whims of a board that is undergoing similar pressure. Against all the dictates of reason, I have presented on navigating AI hype to people on S&P 500 boards and they are in exactly the same situation – the main comments I remember from the session were board members admitting they were skeptical, but expressing anxiety that their positions were contingent on demanding AI investment. One of them commented “investing this early seems like risk without much upside”. About two years later, I can see now that their decade-old multi-billion dollar organisation is now branded as “AI-native”, whatever the hell that means.
V. You Must Be This AI-Native To Ride
All of the above converges on the state that we find ourselves in now, where effective decisionmaking has ground to a halt. Collectively, what started as a few people undergoing either destabilising psychological events or being caught up in hype has now resulted in an environment where leaders cannot speak honestly about their beliefs on how best to guide organisations, for fear of being removed, creating a sort of distributed government by assassination. This means that the least sensible recommendations are going totally unchallenged, resulting in employees being evaluated on totally gameable metrics such as “money spent on AI”, and those employees must play along to avoid being terminated. This has also created an insatiable appetite for purchasing “AI” solutions, which target both true believers that will believe implausible claims, and also non-believers that cannot decline the purchases without having their commitment to the cause coming into question.
This means that all offers that are subject to internal politics at an ideologically captured organisation must include AI alignment, even if the value proposition is patently ambiguous. My assessment of the market so far is that a substantial component of the outburst of AI projects are actually non-AI projects with an AI element slapped on after the fact to pass the purity test.
For example, I recently witnessed an organisation handling a database migration from an Oracle database to Snowflake – instead of handling the migration directly, the vendor bolted on a preliminary phase which involved trying to get an LLM to automate the translation of the Oracle-flavored SQL to Snowflake-flavored SQL. When the project failed (due to issues getting enough permissions to automate the work, not because an LLM can’t do something that easy), the vendor simply started handling the translation by hand but the company billed it as an AI-driven success because some inconsequential portion of the SQL had been translated by AI before being pasted over.
What was actually purchased? A totally standard database migration to help an executive meet the strategic deliverable of decommissioning a system prior to license renewal. What was sold to their superiors? “I allocated a substantial percentage of my budget to AI and it helped me accomplish my mandate.” True AI projects, of the kind that is driven by an LLM as the sole mechanism underlying it, where the project can clearly fail to deliver specific numbers, are actually very rare. We mostly see them in the context of startups, and frankly we have stopped engaging with them because we kept getting to the end of the sales conversation and finding out they wanted us to build the product that they were marketing as completed.
Read more:
https://t.co/YrrD1lL1uY
Australian discovers Texas Roadhouse…
First, he calls it a “fancy restaurant,” and he couldn’t be more wrong. Texas Roadhouse isn’t just a fancy ole restaurant. It’s a giant slice of heaven brought down to earth.
This man loves the bread, the free refills, the service.
In response to those outside the US complaining about tipping there: “You don’t have to tip… you WANT to tip. These are the most wonderful people on earth.”
“I don’t even know why some of you Americans are so angry all the time; you guys have Texas Roadhouse in your country.”
You got that right. 🇺🇸🇺🇸🇺🇸
#worldcup #usa
Everyone's primary focus is the driver can't speak English. Issue? Sure, but licensing and carrier hiring standards are the primary root issues. This company had the same exact crash in NC on 8/23/2024. Company remained on the road then got a Satisfactory rating last month. https://t.co/hZvTCxYlyl
It's Not The Terminator. It's worse. | xtalBiskit, silencing machine
I carry a concealed handgun, legally. I’ve thought pretty carefully about what I’m willing to do with it and under what circumstances. Last week I realized it would be completely useless against a new threat I’m really worried about.
Let me back up.
I used to think the idea of a robot apocalypse was stupid. Not because AI isn’t potentially dangerous, but because the specific fantasy (Skynet wakes up, builds an army of chrome humanoids, declares war on humanity) skips about forty steps and gets basically everything wrong. Humanoid robots are hard to build, expensive to operate, and worse than humans at most of the things you’d actually want a weapon to do. The form factor is a liability.
I used to stop there. Dismiss the robot apocalypse, move on. I don’t anymore.
I’ve been playing a lot of ARC Raiders lately, which is a game where you are part of the last dregs of humanity slugging it out against (among other terrible things) autonomous flying drones. The developers had to nerf the drones to make the game playable. Give the AI realistic tactical behavior and semi-realistic physical capability and the honest answer is “this would be unwinnable.” So they tuned it down. Made the drones a little dumber, a little slower, a little less coordinated than they plausibly could be. And even nerfed, the drones are scary enough that they spontaneously caused human players to cooperate with each other. If you’ve spent any time in online multiplayer shooters, you know how remarkable that is. These games are not known for bringing out the best in people.
Real-world drone developers are working in the opposite direction.
I recognize how strange it is to be using a video game as a jumping-off point for threat analysis. But the thing is, I’m not actually speculating. Cheap autonomous drones are actively killing people in Ukraine right now. That conflict has become the world’s involuntary laboratory for exactly this technology, and the lessons are being absorbed by every military and non-state actor paying attention. The video game is fiction. Ukraine is not.
Here’s what a weaponized drone actually is: a flying landmine that goes looking for targets. It doesn’t need to navigate stairs or open doors or carry a rifle. It moves in three dimensions without caring about terrain. It can approach from any angle, including directly above whatever cover you’re hiding behind. It sees in infrared. It follows you into buildings. It doesn’t need to shoot you. It just needs to get close enough and detonate. The physical form factor isn’t a liability. It’s optimized for exactly this.
And here’s what one costs: not much. A decent FPV drone with a small explosive payload runs maybe $500-1000 fully built. A $200,000 home equity loan (the kind a moderately successful engineer could get on a Tuesday) buys you 200 to 400 of them. That’s not a thought experiment. That’s a shopping list.
But the cost isn’t even the scary part. The scary part is what happens when you have a lot of them at once. You’ve probably seen drone light shows at stadium events or big outdoor concerts. Hundreds of drones moving in precise coordinated patterns, forming shapes in the sky, no human pilot for each one. Just software telling the swarm where to go. That’s the technology. Now replace the lights with explosives and replace “form a shape” with “find a heat signature.” A single drone operated by a single person is scary. A coordinated autonomous swarm doesn’t need a pilot at all. It needs a target description and a launch command. After that it handles the rest. Traditional air defense systems are designed to stop a small number of fast expensive threats. Not 400 slow cheap ones arriving simultaneously from every direction.
I could build one in my garage. All the components are legal. I won’t, because I’m not a monster, but the reason I won’t is entirely ethics and not at all logistics. That is a very thin layer of protection to be relying on.
And you don’t even need rogue AI for this to go badly. You don’t need Skynet. You need one disgruntled engineer with a home equity loan and a Timothy McVeigh moment. The AI angle is almost beside the point. The tools are already here. The only thing standing between right now and a catastrophic autonomous drone attack on a civilian target is the moral character of everyone who knows how to build one. That’s it. That’s the whole security apparatus.
What doesn’t exist yet is any meaningful civilian defense against any of this.
A handgun is optimized for the threat model that drones bypass entirely. Wrong geometry, wrong speed, wrong approach vector. A shotgun is marginally better. Birdshot, wide choke, if you see it coming, in daylight, and there’s only one of them. Against a coordinated swarm at night you’re not bringing a knife to a gunfight. You’re bringing a gun to a drone fight. Same problem, different century.
There are some rudimentary civilian devices that do exist, technically. The net gun, for instance. Range of about 115 feet, can intercept a single medium-sized drone. Here’s the thing nobody mentions about that: you haven’t neutralized the threat. You’ve caught it. A kamikaze drone on a tether is now at chest height, closer to you than it was before, with you holding the other end. It’s like building a rat trap out of a bucket of water. Quite effective. Now instead of a rat problem you have an angry, wet rat problem.
The tools that actually work (RF jammers, directed energy weapons, laser systems) are federally restricted. A corporation can’t deploy them. You definitely can’t. The legal framework for civilian drone defense was written for the threat model of “annoying neighbor filming your backyard,” and it hasn’t been updated for anything else.
So there’s no civilian defense that can be realistically or legally deployed. What about civilian governments? Police forces? I’d love to believe they could effectively address this, I really do. These are also the same institutions that can’t reliably prevent a school shooting, which is one person with one gun. The idea that they’re going to develop and deploy effective infrastructure against coordinated autonomous drone swarms is, to put it kindly, a stretch. So that means relying on defense contractors, which has the potential to be so much worse.
The honest answer to “what actually stops a hostile drone swarm” is, probably, a defensive drone swarm. What’s a defensive drone swarm? Great question. It’s a thing my friend and I made up last week that does not currently exist, is barely even defined, and would require solving enormous engineering problems before it could. I also just watched Benn Jordan’s series on the security vulnerabilities in municipally deployed Flock cameras and robot dogs. Surveillance technology with genuinely terrible security holes, deployed by city governments that didn’t understand what they were buying. Now scale that institutional failure up to autonomous weapons systems. The same procurement process, the same IT department, the same gap between the people specifying the technology and the people accountable for what it does. Except instead of leaking license plate data, a security hole means someone else controls your city’s defensive drone swarm.
The doom loop is pretty tight. Individual defense is legally impossible to prepare in advance. Private deployment at scale requires political will that doesn’t exist. Government deployment at scale inherits the security failures we already see everywhere else. And the threat keeps getting cheaper and more accessible regardless.
When I try to explain any of this to people, I can see the moment they file it next to alien invasions. It sounds like that. Drone swarms attacking civilian targets in America sounds as plausible as a weather control weapon or a mind control satellite. I get it. I sound a little unhinged saying it out loud.
But the logic isn’t actually complicated. The technology is mature. The components are legal and cheap. The defensive infrastructure doesn’t exist. The regulatory framework is twenty years behind. And we know it works because we can watch it working, right now, in Ukraine.
I’m going to hate being right about this one.
https://t.co/1vCSysseOt
☁️ CLOUD BREAD 🍞 Only 2 Ingredients! (High Protein, Low Carb)
Just tried this viral fluff and it’s insane. Light as a cloud, tastes like actual bread, zero guilt. Perfect for keto, high-protein diets, or just when you’re craving something fresh out the oven.
Recipe (makes 1 loaf):
• 3 eggs (yolks + whites separated)
• 150g Greek yogurt (or lactose-free)
• Pinch of salt (optional)
How to make it:
1. Whip egg whites to stiff peaks (electric mixer works best).
2. Mix yolks + Greek yogurt until smooth.
3. Gently fold whites into yolk mix (don’t overmix – keep it airy!).
4. Pour into parchment-lined loaf tin.
5. Bake at 170°C / 340°F for 25-30 min until golden.
6. Cool slightly, slice & enjoy.
Macros (whole loaf approx): ~300 kcal | 33g protein | ~5-7g carbs | ~15g fat.
This thing pulls apart like fresh sourdough but it’s basically pure protein. Game changer for breakfast, sandwiches, or snacking.
If you ever have made this feel free to reply.
How AI Productivity Fails | Shrivu Shankar, Shrivu’s Substack
Why we are not actually 10x more productive (yet).
Most AI users today get ~10–20% more productive no matter how “game changing” they claim it is or how many lines of code they output. Yet, I still think 2x or even 10x+ is both real and reasonably expected. Real transformation requires two changes at once: personal practice and organizational refactoring. Whether the output lands as 10x leverage or 10x slop depends on the practice and the org around it.
So far in 2026, I’ve seen exponential increases in output but linear increases in realized impact. This post covers some of the issues I’ve found “debugging” the issue.
I’ve grouped observations and recommendations into:
Personal Pitfalls — guidance for individuals
Organization Pitfalls — guidance for organizations and leadership
Personal Pitfalls
You don’t shift left, so you don’t understand what you ship.
AI removes the friction that used to force planning. Without something pushing back on bad abstractions, the upfront thinking gets skipped silently. You ship systems you can’t debug or extend, full of blanks AI quietly filled in. Outline first: headers, structure, audience, plan, principles, what-done-looks-like, then let AI fill it in. Review shifts up the stack: outcome-vs-plan instead of line-by-line. The interrogation belongs inside planning: spawn subagent critics to red-team the plan before you let anything generate against it. The easy review at the end is only easy because the hard review at the start was good. If you find it taxing to review your own AI generated output, you didn’t shift left.
Small tasks are slower with AI, not faster.
Context overhead doesn’t shrink with task size, and small tasks are mostly edges. AI handles the middle 80% of work well but can be brittle on the first and last 10%: setup, edge cases, final-mile review. With AI, often a two-line fix pays the same setup cost as a full feature. So you spend more time briefing the agent than just writing it. Even then it lacks context, ships something subtly wrong, and you redo it. Lift task ambition until the context cost is justified. The heuristic: if it’s smaller than a meaningful unit of work (a PR, a section, a chart, a campaign), it’s probably too small. Starting “small” with AI might be the reason you don’t graduate to larger leaps of AI-driven outcomes.
Your parallelism is too low to matter or too high to manage.
AI scales generation, not your human-bound working memory. The ceiling on parallel agents is how many threads a human can hold without dropping context or understanding, and that number is small. You either single-thread with a coffee break while you watch it code or overshoot and abandon five threads. Stay at ~three or fewer active in your cognitive window, or close the loop (next section) and pull fully out. If you are only running a single session at a time, you are probably not delegating enough. If you are struggling to manage dozens of sessions, consider whether you even need to be “managing” them, whether sessions should take longer leaps (fewer agents that do more work), or to intentionally stay linear until you invest in the right skills and context.
You stay in the loop where you could close it.
Closing the last 1% is verification infrastructure and an incredibly well-defined definition of done: a different discipline than generation, so people skip it and layer AI on top of existing handoffs instead. You become the screenshot human, the agent waiting on your click or upload to validate. Close the loop end-to-end: tests, queries, sandboxes, browser tools, type checks, real API responses. Whatever lets the agent see its own output and iterate without you. Before automating the surrounding process, delete it. Closed loops are how you exceed the ~3-agent ceiling: work runs underneath your cognitive window because it doesn’t need to be in it. Believe it or not, AI can handle the “fuzzy” verifications well too with the right pre-defined context, including the “does this actually make sense for long-term system architecture” and “is this the right feature to be adding to the product” questions that I often hear naively claimed as only something a human can do in the loop.
You don’t build leverage from your AI use.
Our chat interfaces often treat every session as ephemeral. Taste is a reusable artifact, but the interface gives you nowhere to deposit it. So encoding never starts, or you overcorrect and write one skill per task, ending up with a directory of one-shots nobody else can use. Build skills for the class of task, not the instance. Some are specs (how to do a thing); others are principles (how to think about a class of things). The durable artifact is the skill (often literally some markdown file), not the prompt. The recognition signal is concrete: any time you find yourself editing or “bullying” the output for a better answer, that’s a rule you can codify once and stop bullying forever. Consider even meta skills that regularly take feedback (from you or across the users of a process) and make the right skill or context modifications from accumulated learnings. Refuse to edit an AI output manually (i.e. typing over it literally) or in a way that doesn’t feed into learnings for next time. Tell it why it’s being dumb, how you think, and make sure that sticks for next time.
The only skill you’re leveling up is asking Claude.
Skill comes from cognitive struggle, and AI removes the struggle by completing the thought before you’ve had it. The learning loop never closes—you can’t tell when the model is wrong, you can’t operate without it, the domain skill quietly atrophies. You grow through resistance: edit, interrogate, override. Bootstrap by using AI on tasks where you’re the domain owner, so the friction is real and the corrections you push back are correct. Juniors get hit hardest: they offload cognition before building the capacity to evaluate output. The people who can evaluate (taste-holders, domain owners) need to be the ones encoding skills and quality-bars even if traditionally the Super Senior ICs and managers used to not touch the codebase. No task should be permanently too hard for an AI system and as it learns, human effort moves up the stack to architecting the system AI builds: a place where AI is genuinely worse and the friction lives. AI will get good at that as well, so you just keep moving up to harder and harder meta derivative skill building.
Organizational Pitfalls
The personal pitfalls above and the organizational ones below are one problem at two scales. AI optimizes individual roles but leaves the process that constrains them intact. Personal practice is where you collapse the steps; organizational design is where you collapse the handoffs. The gains only show up when there’s shared ambition large enough to do both.
Promoting usage instead of outcomes.
Usage is easy to measure; impact is hard. Managers and leadership often carry the cultural pressure to praise visible use over invisible value. Tokens land in perf reviews and in the next cycle Claude loops get left running to inflate counts. Teams ship new AI-shaped systems instead of fixing the existing ones that matter. Reward what shipped, not what got used. Track usage as a leading indicator for enablement and resistance (where to focus training and tooling, especially early in a rollout), but never as the long-term goal. Critically, and this is where I’ve seen the most confusion, pure short-term impact with no recurring AI-powered leverage (often) does not maximize long-term business goals. Weight outcomes by the reusable leverage they leave behind: closed loops, AI-friendly architected systems, codified skills, shared context that make the next ship exponentially cheaper than this one.
Tool sprawl across build, buy, and provider.
Build cost used to be the de facto curator. Only worthwhile tools got built. AI removed that filter without replacing it, people keep shipping, and not everyone has the right taste filter. Discovery becomes harder than building another tool, and context gets inconsistently duplicated across teams and roles. An architect-owner at the top should be accountable for the taxonomy across what you build, what you buy, and which providers you standardize on. Pilots at the bottom—small, scoped, not broadcast prematurely—are how new tools earn graduation, with telemetry and explicit retirement criteria so dead tools get pulled. Consolidate tools where you can and enforce a maintained shared context layer they pull from.
Low-taste skills and context proliferate.
Authoring used to require the expertise being authored. AI broke that coupling, so production no longer signals authority. You get ten invokable ways to do anything, knowledge fragmented across wikis and CLAUDE.mds, personal skills bleeding into shared with no quality bar. A top-down architect or domain-expert should decide the core skill set and assign ownership to the taste-holders who’ll build it. If the people who know what good looks like aren’t writing the skills, the skills are mediocre by default. Personal vs. shared context is an explicit split: personal skills can be loose, shared skills are operating practice and have to be built like one, with review and a real quality bar.
AI output ships faster than anyone can absorb or own.
Generation outpaces review, AI doesn’t progressively disclose, and authorship gets fuzzy when “the prompt” wrote it. Long docs nobody reads, PRs reviewers can’t keep up with, slop ships because no one wants to compromise “productivity”. Hold people accountable for the artifact even when AI generated it; build pushback culture with harsh, specific feedback when output crosses into slop. Use AI and assume that at the end of the day all content will be AI-generated. And despite this, you have to understand what you ship well enough to defend it under questioning. Reviewers should refuse to be the debugger of last resort. People shipping AI output need domain ownership themselves with consistent reliance on skills built and maintained by domain owners.
Handoffs absorb the gains.
Most orgs are organized by function, so getting anything done means handoffs. Coding was always ~20%1 of the cycle; the other 80% (approvals, reviews, syncs) was the rest. AI (if you are doing it right) compressed the 20% to near-zero, leaving the 80% as the entire bottleneck. A 5-minute fix sits 3 days in review; sync meetings spring up to unblock work AI already finished. Loop ownership should replace function ownership: one person closes the chain from problem to deployment, with the right guardrails so they can move without sacrificing function-level taste. Specialists shift to platform—encoding their taste into the systems, prompts, and context that loop owners’ agents use. Bottom-up speed alone hits a wall here; the org has to reorient around loops to absorb the acceleration. I call this transposing your organization.
Top-down without clarity.
Mandates transmit behavior, not taste or judgment. Without ground truth at the top, each layer strips intent on the way down. Engineers get pushed into reviewing AI slop instead of being elevated into architects of it—the job becomes downstream cleanup, not upstream design, and replacement fear takes root underneath. Mandate is still a powerful lever. What’s often missing is clarity: explicit expectations, updated role-definitions, and the why behind the importance and urgency of AI adoption. Be intentional about when top-down is the right lever and when bottom-up is. And don’t kill the fun of building. People like building things and you want people to like what they do — it’s the top-down’s job to make sure they are building useful things.
Bottom-up without expectation setting.
Bottom-up energy needs a target to compound against. Without a shared definition of return, every team optimizes locally and the gains never aggregate. Usage looks great while token spend decouples from business outcomes, and sprawl goes unchecked. Make ROI legible: every team articulates return on its AI investment, even crudely. Wire working behaviors—closed loops, codified skills, outcome-aligned spend—into career pathways so the right behaviors get rewarded structurally, not just culturally.
So…
The 2x version of you isn’t a token-count away. AI multiplies what you already do well and whatever your org already enables. If your practice is loose, AI compounds the looseness. If your org runs on handoffs, AI accelerates the frequency of handoffs. Both have to change at once, or neither change matters.
The 10–20% is free at this point. Anything past that is rebuilding: personal practice on one side, organizational design on the other. I think most folks still have an uncomfortable amount of self-refactoring left to do.
https://t.co/OsbYlstwN5
More than 1.6K stars on terraform-skill! ❤️
Writing #Terraform with AI, but not always getting great results?
I have just shipped a major update to Terraform Skill v1.7.0, with significantly fewer hallucinations.
It was also fact-checked by seven independent reviewers: five Claude expert personas (Terraform, Security, DevOps, SRE, Cloud), plus GPT Codex and Gemini.
Terraform Skill v1.7.0 is ready for Claude, Cursor, Copilot, Gemini, Codex, and more:
npx skills add https://t.co/S6rOXICIBd
These .claude/settings.json options fixed most of the issues I was having with Claude:
{
"model": "claude-opus-4-6",
"effortLevel": "high",
"alwaysThinkingEnabled": true,
"env": {
"CLAUDE_CODE_DISABLE_ADAPTIVE_THINKING": "1",
"MAX_THINKING_TOKENS": "31999"
},
}
unpopular dockerfile takes (that actually work)
1 - stop using alpine — yes, it's tiny. but musl libc ≠ glibc. your python/node app will rebuild native deps from scratch or just... silently be slower. use -slim (debian-slim) instead. same size win, zero grief.
2 - layer order is your cache strategy. COPY your lockfile first, run install, then copy source. invalidating the install layer on every code change is a skill issue ngl
3 - multi-stage builds aren't just "best practice" — they're the actual reason your prod image doesn't ship gcc and 400mb of build tools. builder stage = bloat zone. final stage = lean mean container.
4 - COPY . . is fine actually — if your .dockerignore is correct. most pain here is from forgetting to ignore node_modules/, .git, *.log. fix the ignore file, not the COPY.
5 - one process per container is a vibe, not a law. if your app needs nginx + app server and you're not at k8s scale — just use supervisord. the "one process" dogma costs more complexity than it saves sometimes.
6 - pin your base image by digest, not tag. node:20 today ≠ node:20 in 6 months. prod broke because of a tag? that's a you problem tbh.
7 - BuildKit cache mounts (--mount=type=cache) will change your life. pip/apt/cargo cache between builds without it ending up in the final layer. nobody talks about this enough fr
there's no "best practice" in a vacuum. alpine is great for Go binaries. slim is great for Python. scratch is great for static bins. know your workload, then choose.
btw if you want something to catch all this stuff automatically -
check out dockerfile-roast — a linter written in Rust that literally roasts your Dockerfile. 63 rules, brutally honest output (but it can also provide just dry facts, no roast), runs on any OS or as a docker container
https://t.co/NVYpe8iD65
#docker #devops #kubernetes #backend #linux #rust #sre #containers
A single 𝗖𝗟𝗔𝗨𝗗𝗘.𝗺𝗱 file just hit 15K GitHub stars.
(derived from Karpathy's coding rules)
Andrej Karpathy observed that LLMs make the same predictable mistakes when writing code: over-engineering, ignoring existing patterns, and adding dependencies you never asked for.
If you've used AI coding assistants, you've hit all of these.
But here's the thing:
If the mistakes are predictable, you can prevent them with the right instructions.
That's exactly what this 𝗖𝗟𝗔𝗨𝗗𝗘.𝗺𝗱 does. You drop one markdown file into your repo, and it gives Claude Code a structured set of behavioral guidelines for your entire project.
This is a big deal.
- Built entirely around prompt engineering for AI coding assistants
- No framework, no complex tooling, just one .md file that shapes behavior
Developers are moving past "use AI to write code" and into "engineer the AI's behavior so the code is actually good."
The Claude Code ecosystem is growing fast, and the best tools in it aren't always software. Sometimes they're just well-crafted instructions.
100% open-source.
I've shared a link to the GitHub repo in the next tweet!
This 2 hour Stanford lecture shows exactly how Stanford trains it's engineers to build AI systems. It's more practical than every Claude tutorial & prompting threads you've seen.
Bookmark & give it 2 hours, no matter what. It'll be the most productive thing you do this weekend.