Passed the Claude Certified Architect Foundations exam.
One thing worth saying clearly. This exam is not meant for people who think spending time on the Claude Code CLI makes them an expert. It goes well beyond terminal fluency. You are tested on real-world scenarios like context window management under degradation, API and cost optimisation across synchronous and batch workflows, structured error propagation in multi-agent systems, escalation and handoff patterns, and the agentic design choices that hold up in production. If you have only used Claude Code to vibe-code side projects, you will struggle. The exam rewards people who have shipped systems, debugged the wrong tool calls, and understood why a hook beats a prompt when guarantees matter.
What I did not expect was an entire scenario on Conversational AI that is not mentioned anywhere in the official exam guide or the Architect's Playbook. Roughly a quarter of my exam came from a domain I had no targeted prep for. Multi-turn context management, instruction persistence across turns, memory strategies, handling ambiguous user inputs. None of it appears in the published blueprint.
What the exam actually tests is sharper than the guide suggests.
The breadth is wide. Five domains covering Claude Code configuration, Agent SDK orchestration, MCP integration, prompt engineering, and context management, plus six (now seven) production scenarios. The depth is sharper. Distractors are written by people who have shipped real systems, so the wrong answers are exactly the wrong answers you would give in production if you had partial knowledge.
AI agents can be classified into four types based on the user persona.
This persona-based approach abstracts four things: where execution runs, where state is persisted, how authority is delegated, and where policy is enforced.
Read my latest article for a deep dive into the classification - https://t.co/UBwsv30PiJ
Finally installed, configured, and hardened @NousResearch Hermes Agent on a dedicated Intel NUC. Using Sonnet 4.5 as the primary model and Kimi K2.6 as the secondary. In the process of adding skills to talk my Google Workspace, M365, GitHub, Notion, and more.
I am building a multi-agent system, and the root agent, which is the coordinator, is called the Gang Leader. The next-in-line agent is called Big Boss, and the sub-agents are called Indra, Rudranethra, and Chantabbai.
If you are a fan of Chiranjeevi, you know what this means ๐
Amazon Quick Walks Into The Trap That Killed WorkDocs
Amazon just shipped a desktop app and Microsoft 365 plug-ins for Quick. The product is now technically capable. The market problem is unchanged.
This is Amazon's fourth attempt in five years to sell software to knowledge workers. WorkDocs shut down in April 2025. Chime was discontinued in February 2026. WorkMail ends support in March 2027. Three productivity products, three quiet exits.
The Cowork race makes the gap clearer. Anthropic ships Claude Cowork on the desktop. Microsoft licensed Anthropic's technology and embedded it inside Outlook, Teams, Word and Excel. Amazon ships Quick as a separate destination, with Microsoft 365 extensions still in preview.
Microsoft has 15 million paid Copilot seats and only 35.8 percent of those users are active. If the company with the largest distribution advantage in enterprise software cannot convert seats into daily habit, what does that say about a vendor starting later with no comparable surface?
The uncomfortable truth is that Amazon's own corporate workforce runs on Microsoft 365. Quick has to win a market where its own parent company is the customer of its biggest competitor.
Read my analysis on Forbes - https://t.co/ujeWsSglDy
Is a new composable AI coding stack taking shape? @janakiramm thinks so.
Instead of one tool to rule them all, a different stack is showing up dev environments.
https://t.co/cXmrMLuPur
All the AI dev tools vendor are betting on orchestration. Where does that surface live? The terminal, like Claude Code? Sandalone desktop app like Codex? As part of the IDE like Antigravity and Cursor?
@janakiramm takes a good look worth checking out ... https://t.co/ojTFvqR6ry
Kubernetes is becoming the AI inference operating system โ and KubeCon + CloudNativeCon Europe 2026 made that crystal clear. ๐ฅ
82% of orgs use Kubernetes for AI workloads, yet only 7% deploy AI daily. That gap is the real story.
Projects like llm-d are building the bridge. ๐
Read the full breakdown by @janakiramm on @Forbes โ https://t.co/E7PQuMhEAY
#OpenSource #Kubernetes #AIInference #KubeCon #CloudNativeCon
800 attendees. 1 mission: Automating the cloud ๐
Last week, we hosted the inauguralย @kubeautodayย in Amsterdam, and the energy was electric! We set out with a simple goal: to bring together practitioners using AI to automate cloud at scale through real production stories.
The community response blew us away:
๐งโ๐ป 800ย checked-in attendees.
๐ฅ House fullย sessions from the first keynote to the final talk.
๐ Deep-dive technical discussions that prove AI-driven automation is no longer "the future", itโs happening now.
A massive thank you to our sponsors for making this possible:ย @cast_ai, @AMD, @QodoAI, @stack8s, @awscloud, @coderabbitai,ย andย @OpenObserve and amazing speakers including Kelsey Hightower, @Njuchi_, Daniel Gebler, @salaboy, @tbotskina, Ioana Adelina Apetrei, @virtualized6ix, Andrew Martin, @janakiramm and more!
๐ย Next Stop: The KubeAuto Global Tour!ย Weโre just getting started. We are taking these stories across the globe:
๐ย Sรฃo Pauloย โ 27th May, 2026
๐ย Mumbaiย โ 17th June, 2026
๐ย USAย โ Late 2026
๐ Read more about it in our press release: https://t.co/OGJM6ENGrF
Stay tuned for registration details!
If you'd asked me last year to run an autonomous research loop across two GPUs, I'd have said that's not something I can do.
Not "it'll take a while." Impossible.
Today this is my Saturday.
My ๐ง-๐๐ฎ๐ญ๐จ๐ซ๐๐ฌ๐๐๐ซ๐๐ก ๐ฅ๐จ๐จ๐ฉ ๐ฃ๐ฎ๐ฌ๐ญ ๐๐ข๐ง๐ข๐ฌ๐ก๐๐ ๐ ๐-๐ก๐จ๐ฎ๐ซ ๐ซ๐ฎ๐ง ๐จ๐ง ๐๐ฎ๐๐ฅ ๐๐๐ ๐๐๐๐๐ฌ: autonomous hyperparameter mutation, parallel experiments, no babysitting.
Results:
โ 17 experiments, 0 crashes
โ baseline 1.2365 โ ๐๐๐๐ ๐ท.๐ธ๐ท๐พ๐ธ ๐๐๐_๐๐๐
โ 1.48% improvement. Found by the loop.
The chart shows the staircase, the same pattern @Karpathy sees in his runs. His: 2 days, 276 experiments. Mine: 1 hour, 17. Same logic, different constraints.
The constraint here is the ๐บ๐ถ๐ฟ๐ถ'๐ ๐ธ๐บ๐ถ๐ฑ ๐ ๐๐ฐ๐ผ ๐๐๐๐๐๐๐ ๐๐๐๐๐_๐๐๐ฃ๐=๐บ, which means ~5.5% MFU. The GPU is mostly waiting on memory transfers, not computing.
What the same loop looks like on different hardware:
๐๐ฑ ๐๐๐ ๐๐๐๐ (๐ฐ๐ก๐๐ญ ๐ ๐ซ๐๐ง): ~๐๐ ๐๐ฑ๐ฉ๐๐ซ๐ข๐ฆ๐๐ง๐ญ๐ฌ/๐ก๐ซ โ ๐.๐๐% ๐ ๐๐ข๐ง
๐๐ฑ ๐๐๐๐ ๐๐๐: ~๐๐+ ๐๐ฑ๐ฉ๐๐ซ๐ข๐ฆ๐๐ง๐ญ๐ฌ/๐ก๐ซ โ ๐๐ฌ๐ญ๐ข๐ฆ๐๐ญ๐๐ ๐-๐% ๐ ๐๐ข๐ง
๐๐ฑ ๐๐๐๐ (๐๐๐): ~๐๐๐+ ๐๐ฑ๐ฉ๐๐ซ๐ข๐ฆ๐๐ง๐ญ๐ฌ/๐ก๐ซ โ ๐-๐๐%
๐๐ฑ ๐๐๐๐ (๐๐๐): ~๐๐๐+ ๐๐ฑ๐ฉ๐๐ซ๐ข๐ฆ๐๐ง๐ญ๐ฌ/๐ก๐ซ โ ๐๐-๐๐%
Same autoresearch loop. Just more runway.
Karpathy's dream setup, 8x H100 running 48 hours would likely hit 5,000+ experiments. On 4090s that's not feasible. But 24-48 hours on what I have would still find significantly more than 1 hour did. That's what's running next.
This is our own multi-GPU implementation built on top of his single-GPU original, orchestrated with iii functions, workers, and triggers.
Claude Code made the extension possible in a weekend.
Google Solved the Inference Problem. AWS and Microsoft Are Still Trying.
Three hyperscaler announcements. One week. The same admission underneath all of them.
AWS partnered with Cerebras to disaggregate inference hardware because Trainium alone isn't enough for reasoning models. Microsoft licensed Fireworks' inference engine into Foundry because its own serving layer had fallen behind. Google, which spent seven generations building Ironwood specifically for the inference era, needed neither fix.
Inference has become the primary competitive axis in enterprise AI. Not model catalogs. Not developer tools. The speed, cost and reliability of how models run in production.
My latest Forbes analysis breaks down all three strategies and what the pattern means for enterprise AI buyers.
https://t.co/fSDff8kIRW