I'm sharing with you one of the best papers I've read on Harness Engineering
it takes apart Claude Code (Opus 5.5), Codex, Gemini CLI and more, and boils all of them down to one line:
agent = model + harness
what you walk away with:
> what a harness actually is
> how today's coding agents are built under the hood
> the patterns every one of them repeats
> where the whole category is heading
> their checklist for building your own
the model gets the headlines. the harness decides whether the agent works
loop and graph engineering on top of it - in my article below
用 Opus5.5 的朋友们必看 Claude 官方的这篇文章《Getting the most out of Opus 5.5 in Claude and Claude Code》充分发挥 Opus 5.5 的能力
https://t.co/GqBwPRR6JR
作者是 Addy Osmani 算个大名人了,之前在 google 呆了 14 年,写了很多前端的书,我也读过很多,2026年9月加入 A 社了。
这篇文章我觉得优势是特别可以通俗,没有什么复杂的原理 但是非常偏实战
几个方面:1 怎么下任务 2 控制 Long Run 的集中方法 3 验收结果 4 Claude App 的使用技巧 5 Opus 5.5 的安全开关 ,这个比较有趣 一旦被判敏感 Claude 不会停,而是自动换成更旧的模型接着答
you can prompt this entire facility
one model controls everything: equipment, researchers, and inventory
I spent two weeks living inside it, working on C5R's launch with Astra – here's what it felt like:
this is f*cking gold.
How to build your first AI agent with Jev.
This is everything you need to know to be ahead of the most people.
Once Jev is in your loop, your agent picks its own next action, rates its own outputs, and decides when the goal is met.
it runs without you watching it.
Jev is a decision model built specifically for this. it reads your system state and returns a typed answer with confidence:
Choice - which agent or action should happen next?
Score - is this result good enough to keep?
Noul - is the goal met?
the architecture:
LLM thinks. Jev decides. Agent executes.
10 minutes to get it running. Hosted on Vercel and Cloudflare via Typesafe. Waitlist access, people report getting in within a day.
The guide in the image maps the full setup, from architecture to your first decision call.
Wrote the full breakdown on agents, loops, and graphs below.
Such a great article! After talking with several brilliant researchers last night, I realized that beneath all the excitement around Physical AI, the fundamentals haven’t really changed. The real challenges still lie in hardware, deployment, human–robot interfaces, and data.
The hardest problems may only truly emerge after GPT-6. That’s bad news—but also a huge opportunity.
We have plenty of human demonstration data for movement. We don’t have enough for touch.
Without tactile feedback, foundation models can’t feel if an object is slipping, soft, or fragile during complex manipulation.
Matrix Inno is bringing three data-collection gloves to IROS 2026: the RigidCore Tactile+EMF, the Fabric Tactile + EMF, and the Fabric Tactile + UltraLite Motion Glove. The hardware tracks pressure and hand pose simultaneously during human demos, mapping exactly where contact happens and how it shifts in real time.
Wearables like this are how we finally get real, continuous contact data at scale.
Source: Matrix Inno @moxianmxkj
Google just dropped a 12-page PDF on Agentic Engineering - building agents and agents teams that will change your life
here's the 5-stage Google agentic engineering pipeline:
stage 1 → specification - define success with tests and automated checks
stage 2 → harness - give the agent tools, prompts, guardrails, and everything it needs to act
stage 3 → trajectory - let it execute the full workflow instead of judging one response at a time
stage 4 → verification - run linters, tests, and semantic checks before accepting the result
stage 5 → meta-debug - when the agent fails, fix the workflow, not the model
the shift is simple:
stop asking "which model should I use?"
start asking "what system do we build around the model?"
Read this 12-page PDF before building another agent - it breaks down five real Google systems most agent tutorials never show
save now, then read how to build your first agent in the article below