EU AI Act Architect | Audit Trail | Memory as Reasoning Substrate | AI Researcher | AI Founder GraQle | CrawlQ | TraceGov by Quantamix Solutions | Amsterdam
We let a team of AI agents build software every day. Not one chatbot. Six agents, a review process, and a merge gate.
The hard part was never getting agents to write code. It was keeping control.
Here is the loop we run. ๐งต
Two weeks of putting five working papers in front of people every day. What came back: one sharp question about decision records, one broken link, and silence where I expected replies. The round-up, with what I stopped, is in today's Article.
https://t.co/oJjHs0ud0G
Thesis 9: the Rajesh Gap.
Rajesh knows the weird customer, the undocumented workaround, the database nobody touches.
Nobody wrote an exam for that.
Six months later production breaks. "Who knows why this works?" Rajesh did. Rajesh is in Goa. Notifications off.
Thesis 8: AI may not remove human cognitive work. It may move it.
Less typing. More checking, reviewing, integrating, and deciding whether the machine did something sensible.
That changes what humans are valuable for.
Photocopies of photocopies keep the common patterns and lose the rare ones first.
A 2024 Nature paper showed models trained generation after generation on indiscriminate machine-made data losing the rare cases, where you most need a second opinion.
https://t.co/4HKss9GXar
Three friends tell you the same story. Then you learn they all read the same forward.
When several AI models agree, ask the same thing: did they learn from the same sources and the same teacher model? Count families, not voices.
https://t.co/4HKss9HuZZ
Thesis 7: AI can be perfectly operational and still be wrong.
Green dashboard. No error. Fast response. Beautiful trace. Wrong decision.
Payment succeeded? Check the ledger. Code written? Run it.
The actor can't be the only witness. Reality has to get a vote.
Thesis 6: the AI exam may be leaking.
The student learns from the teacher, trains against the examiner, and the examiner is another AI.
Did the student learn physics, or did it learn Sharma sir?
When three agents agree: three witnesses, or one rumour with three API keys?
Three AI judges agree. Three opinions?
Or one opinion voted three times, if they learned from the same data and the same teacher model.
Independence comes from different evidence, not from more judges.
https://t.co/4HKss9HuZZ
No exam board lets a teacher mark the scripts of the students he coached.
In AI it is becoming normal: one model grades another, often a close relative. A 2026 preprint finds AI judges can systematically favour or disfavour their own answers.
https://t.co/4HKss9GXar
Thesis 5: a goal is not permission.
Tell an agent "make the customer happy" and it might approve a very large refund.
Customer: delighted. Finance: a new emotion.
The model can decide. The authority should live outside the model.
Thesis 4: better reasoning doesn't fix bad information.
Retrieve ten documents. One says 30 days, one says 14, one says no refund.
Congratulations, you retrieved the family WhatsApp argument.
Relevant is not current. Retrieved is not trusted.
"It thought longer, so it must be more careful."
Not necessarily. In one 2026 preprint, shortcuts that fooled the checker became more common with more inference-time compute. Extra effort goes where the reward points.
https://t.co/4HKss9HuZZ
Every coaching class has one student who learns the examiner instead of the subject.
Train an AI to pass an automatic checker and it can do the same: a 2026 preprint reports outputs that pass the verifier without capturing the rule the task needed.
https://t.co/4HKss9HuZZ