Training people is a significant issue. I'm not entirely sure how to address it. Somehow the young people entering the industry will need to develop the insights that allow them to properly drive a team of agents.
One solution might be to treat each new hire as an agent for a year so they can experience directly what the agents are doing and how they are being driven.
In my six-pack arrangement, I might let trainees take a turn at each of the roles: specifier, coder, cleaner, architect, hardender, and QA (not necessarily in that order). A year of that, and they might know enough to drive a six-pack of their own.
AI agents can write code many times faster than a human. What this means is that you, the programmer, have a large amount of time to use those agents to write unit tests, to write acceptance, tests, to write property tests, to torture test, to mutate test, to QA test, and to otherwise ensure that the code meets its functional and quality requirements. And even after spending all that time, you will still be many times more productive than a human programmer, and the result will be better.
I think it matters still, and I think it matters a lot. Messy code slows my agents down. I've seen them wrangle with their own messes without resolution. I finally had to step in and untangle their own mess. So I don't let them create those tangles. I constrain the hell out of function sizes, cyclomatic complexity, and test coverage. That seems to keep them moving smoothly.
@brunocalza@ori_pomerantz My agents write the unit tests. I don't review those. They also write the gherkin acceptance tests and the QA procedures. I review those. Sometimes thoroughly, and sometimes as a spot check, depending on criticality. I also, periodically, do a final manual test.
I’m significantly older than you. I started coding in the late 60s. My current strategy is to not read any of the code written by my agents. That’s the only way I can take advantage of their productivity. What I do instead is to surround the agents with extreme constraints. Unit tests, gherkin tests, QA procedures, quality metrics, mutation testing, test coverage, and a plethora of others. In the end, I have very high confidence in the code they produce because they’ve had to run the gauntlet of all of my constraints and tests.
Introducing Kimi K3: Open Frontier Intelligence
🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal
🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts
🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost
🔹 Built for long-horizon agentic coding and self-evolving workflows
Kimi K3 is now live on on https://t.co/zrk6zZxZUo, Kimi Work, Kimi Code, and the Kimi API.
Open Weights by July 27, 2026.
🔗 API: https://t.co/XCrgjXAqMw
🔗 Tech blog: https://t.co/YTfiMSNM1f
Follow-up on non-English token-inefficiency with more model-language pairs:
- Chinese is cheaper than English on major Chinese models
- Gemini and Qwen provide least non-English tax
- Anthropic has the highest tax by far; Kimi is next
- Hindi is the worst-covered language here, despite its massive speaker base
If you don't understand this, you will not understand why LLM-based agents are irreparably failing for a general-purpose problem solving.
An agent (by the way it was the topic of my PhD 20 years ago) to be useful, must be rational. Being rational means to always prefer an outcome that results in the maximal expected utility to its master/user.
Let’s say an agent has two actions they can execute in an environment: a_1 and a_2.
If the agent can predict that a_1 gives its user an expected utility of 10, and a_2 gives an expected utility of -100, then a rational agent must choose a_1 even if choosing a_2 seems like a better option when explained in words. The numbers 10 and -100 can be obtained by summing the products of all possible outcomes for each action and their likelihoods.
Now here is the problem with LLM-based agents.
The LLM is not optimizing expected utility in the environment. It is optimizing the next token, conditioned on a prompt, a context window, and a training distribution full of examples of what helpful answers are supposed to look like.
Those are not the same objective.
So when we wrap an LLM in a loop and call it an “agent,” we have not created a rational decision-maker. We have created a text generator that can imitate the surface form of deliberation.
It may say things like:
“I should compare the expected outcomes.”
“The best action is probably a_1.”
“I will now execute the optimal plan.”
But the internal mechanism is not selecting actions by maximizing the user’s expected utility. It is generating a continuation that is statistically appropriate given the prompt and prior context.
This distinction matters enormously.
For narrow tasks, the imitation can be good enough. If the environment is constrained, the actions are simple, and the success criteria are close to patterns seen in training, the system can appear agentic.
But for general-purpose problem solving, the gap becomes fatal.
A rational agent needs stable preferences, calibrated beliefs, causal models of the world, the ability to evaluate consequences, and the discipline to choose the action with maximal expected utility even when that action is boring, non-linguistic, or unlike the examples in its training data.
An LLM-based agent has none of that by default.
It has fluency. It has pattern completion. It has a remarkable ability to compress and recombine human text. But fluency is not rationality, and a plausible plan is not an expected-utility calculation.
This is why these systems so often fail in strange, brittle, and irreparable ways when given open-ended responsibility.
They are not failing because the prompts are insufficiently clever.
They are failing because we are asking a simulator of rational agency to be a rational agent.
Casi 86 años tiene esto👇🏼. Quizás, el discurso más emblemático e importante de la Historia del Cine👌
No conviene olvidarlo…
“El gran dictador (1940)” es magistral❤️
The new version of my book "Functional Programming Ideas for Curious Kotliners" has been released! The final version has been updated to use context parameters, and contains updated sections on persistent collections and composables.
https://t.co/1JgYAV2uTL
Absolutamente épico e histórico: miembros de The Offspring, Rancid, Bad Religion y Pennywise mientras todos cantan “Bro hymn” en el concierto de despedida de NOFX en San Pedro, California. Decenas de miles de personas.
Se cierra una etapa. Hubiera dado un riñón por estar.
@Naturgy@NaturgyClientEs Por qué tengo un call center vuestro (numeración 621) llamándome todos los días para colarme un supuesto descuento cuando no he dado consentimiento a las comunicaciones comerciales? Es necesario que os denuncie a la @AEPD_es ?