@ChaitanyaPro Agreed, orchestration drives a lot of it. But that's the app layer. The article is one layer down: even a well-orchestrated agent still has to decide where the generative core runs, and today that's gated by the runtime, not the orchestration. Keen to read your paper.
An agent that runs all day isn't a speed problem. It's a power problem. And there's exactly one block on your laptop's chip built to run all day without draining the battery: the NPU (Neural Processing Unit).
It's in every laptop sold this year. Your agent mostly isn't running on it yet. I looked across the major platforms to see whether that's a silicon problem or a software one, and the answer is more encouraging than it sounds: it's the runtime, not the chip, and the gap is closing fast.
Let me know your thoughts.
https://t.co/ctDoJiQ6mV
On-device agentic AI in 2026 nails the single tool-call demo. The instant you ask the same model to plan across three tools and thread results between them, the wheels come off. Same 3B local model: 8 of 8 right on one tool, 1 in 15 right on three. I ran the benchmark on my Mac mini and a Windows Copilot+ PC to figure out whether the cliff is a hardware problem or a model problem (full write-up with source code here).
Let me know your thoughts.
https://t.co/Bfh9g1pDdB
FLOPS are not the only binding constraint on agentic AI. The CPU layer is becoming the bottleneck.
That's why CPU-focused silicon companies are surging in the public market right now. More detail here (source code included).
Let me know your thoughts.
https://t.co/R91ovT9p3K
For GBrain I built a proper eval harness. 145 queries, Opus-generated corpus. The retrieval stack uses graph based, vector based and Grep based strategies in combination.
The graph layer is worth +31 points on precision. Vector-only misses 170/261 correct answers that the full system finds. Keyword + vector + graph are three separable wins, each load-bearing.
Standard information retrieval metrics: the same ones Google uses to measure search quality.
Precision at 5: You ask a question, the system returns 5 results. How many of those 5 are actually useful? If 3 out of 5 are relevant, P@5 = 60%. It measures: am I wasting your time with junk results?
Recall at 5: For a given question, there might be 3 pages in the entire brain that are genuinely relevant. If the system finds all 3 in its top 5, R@5 = 100%. If it only finds 1, R@5 = 33%. It measures: am I missing things you need?
High precision = low noise.
High recall = nothing slips through.
GBrain's 97.9% R@5 means it almost never misses the right answer. The 49.1% P@5 means about half the results are relevant — which is good when you realize that for most queries there are only 1-2 right answers out of 17,888 pages, so 2.5 hits out of 5 is strong signal.
Entity resolution is zero-LLM-call: regex extracts typed links (works_at, invested_in, founded) on every write. Re-embed on write not on a timer, so decay = stale pages, and stale pages get rewritten when new info lands.
Scorecards: https://t.co/PV2qDE11UC
Inference Chips for Agent Workflows
@sdianahu
Most AI chips are designed for "prompt in, response out." Agents don't work that way. They loop, branch, and hold context across dozens of steps, and current GPUs hit 30–40% utilization as a result.
That gap is where purpose-built silicon wins.
“We delivered robust Q1 results, reflecting the growing and essential role of the CPU in the AI era and unprecedented demand for silicon,” says Intel CFO @dzinsner. Read more about our business highlights from Q1 2026: https://t.co/tE8kwBdZvi
“With a solid foundation in place, we are addressing this opportunity by listening to our customers and driving their success with our technical expertise and differentiated IP. This deliberate reset to how we operate drove a sixth consecutive quarter of revenue above our expectations, as well as new and deepened relationships with strategic partners,” says Intel CEO @LipBuTan1. Read more in our Q1 2026 earnings report: https://t.co/NCw80CPLcU
Excited to share v0.11.0 of Exo!
highlights:
- meta prompting to learn your style
- better gmail syncing (including ability to turn it off)
- better interaction with external agents like openclaw
- many bug fixes, sourced from the community.
It's now pretty stable. Try it out!
GStack is an open-source toolkit built by YC President & CEO @garrytan that turns Claude Code into an AI engineering team — with skills for office hours, design, code review, QA, and browser testing.
In this video, Garry walks through how GStack works, starting with Office Hours, a skill modeled after real YC partner sessions that pressure-tests your idea before you write a line of code. He demos it live, going from idea through adversarial review, design mockups, and automated QA in a single session.
@aj_chan16 Spot on, AJ. We’re moving toward a triage model where efficient local agents handle the routine tasks, leaving the specialists in the cloud for the heavy lifting. It’s the only way the math works without breaking the grid.
I’ve spent much of the last few months deep diving into my roots as a programmer with Claude Code again. It’s addicting. I can now satisfy my curiosities by building things instead of just reading them.
In this piece about the Hardware Reality of AI Inference, I’m sharing an analysis of the upcoming compute bottlenecks. I argue why simply adding more data centers is not enough. We’re heavily memory, network, and thermal constrained. I even used Claude Code to reproduce one of these on my Mac, with code available.
https://t.co/cXwz4OkOwM
What do you think?