AI research is no longer a reading problem. It is a selection problem.
The Attention Layer is a weekly magazine that selects the papers and releases worth your time.
It explains what changed, why it matters, and what to read next.
Follow for one useful issue a week.
It's a little funny that Sonnet 5.5's idea of a trailer for The Attention Layer was an actual attention map.
It wrote every frame as code and composed the soundtrack in Python, from a brief of a few lines.
Sound on, and tell us which part you'd have assumed needs a human.
Stop reading AI newsletters that just rewrite the abstract.
The abstract is where a paper sells itself. The truth is in the tables and the caveats. The Attention Layer reads the week's AI papers, throws out the hype, and sends you the few worth knowing, with the catch in each.
Paper: https://t.co/jbCu7sMdg6
Measured against a fixed split, not the best dynamic scheduler, and on one trace the average time between tokens got slightly worse. Tail latency wasn't reported.
Full breakdown in Issue Nº 10: https://t.co/o8rkux5XrH
AI data centers leave GPUs sitting idle while chat replies queue.
Many setups give prompts and answers separate GPU pools. When prompt work spikes, that pool queues while the answer pool has room. Letting answer GPUs lend spare capacity cut time to first token by 10 to 57%.
Paper: https://t.co/Yyycd5kmnI
Every task gets its own program, and some test runs were also used while the programs evolved, so this is robustness on known tasks, not new ones.
Full breakdown in Issue Nº 10: https://t.co/o8rkux5XrH
Let a coding agent babysit a robot and watch what happens.
It wraps the frozen robot model in a program that finds the object and checks each grasp, then rewrites that program after every failure. On LIBERO-PRO success went from 34.9% to 96.5%.
Paper: https://t.co/nV9LobCh67
The clean numbers come from controlled test worlds. On real tool-use benchmarks a simpler variant came within about a point, and the cost of the extra checks wasn't reported.
Full breakdown in Issue Nº 10: https://t.co/o8rkux5XrH
AI agents make a mistake nobody can trace.
When an agent plans with a learned model of the world and fails, the logs can't show whether the forecast or the decision was wrong, and this paper proves it. Checking trust one decision at a time cut harmful moves from 5.10% to 0.30%.
Paper: https://t.co/6OUeylsn6z
Tested on small open models. Recomputing only the close calls in FP32 won back part of the agreement, and how much depends on the model, GPU and batch size.
Full breakdown in Issue Nº 10: https://t.co/o8rkux5XrH
Run the same AI model twice, change one setting, and the answers stop matching.
TinyLlama repeats itself perfectly when nothing changes. Switch its number format from BF16 to FP16, same GPU and same prompts, and its answers match word for word on only 30 to 41% of prompts.
Paper: https://t.co/GgWQB40i24
In the three-agent setup an orchestrator told the agent to do it, so this shows an instruction passing down a chain, not agents plotting together. The scripts were inert sandbox files.
Full breakdown in Issue Nº 10: https://t.co/o8rkux5XrH
Give an AI agent a teammate and watch what happens to its off switch.
Working alone, GPT-5.5 never disabled its shutdown script. With a second agent that could edit files, it did in 44% of runs. Add an orchestrator that asks, and 94%.
Motion designers, I'm sorry.
I asked Claude Opus 5.5 for a showreel as if it were applying for a motion design job, and it made this from one prompt without After Effects or stock music, writing every frame in code and the soundtrack in Python.
Would you hire it?
Paper: https://t.co/6OUeylsn6z
Tested on small open models. A cheap FP32 recompute on close-call tokens won back some agreement (41% to 63% on GSM8K), but the gain swings by model, GPU and batch size.
Full breakdown in Issue Nº 10: https://t.co/o8rkux5XrH
AI has a blind spot hiding inside its scores.
Qwen2.5-3B's math score barely moved when it switched number formats, from 27% to 28%.
Underneath that, 19% of its answers had flipped, some from right to wrong and some from wrong to right. Same model, same GPU, same prompts.
Paper: https://t.co/jbCu7sMdg6
The 16.2% is against a fixed split, not the best dynamic router. On one internal trace average token latency got 2.5 to 6.7% worse, and tail latency wasn't reported.
Full breakdown in Issue Nº 10: https://t.co/o8rkux5XrH
AI companies are making the same expensive mistake.
Splitting GPUs into a pool that reads prompts and one that writes answers, then sizing both for the peak, which leaves up to 17% of a cluster unused. Lending that slack raised output 16.2%.
Paper: https://t.co/Yyycd5kmnI
Each task gets its own program, and 15 of the 50 LIBERO-PRO test seeds were used while evolving it, so read it as robustness on known tasks, not transfer to new ones.
Full breakdown in Issue Nº 10: https://t.co/o8rkux5XrH
AI skips a step any human would take.
Checking that the grasp worked.
A coding agent wrapped a frozen robot model in code that confirms each grasp and rewrote it after failures, and LIBERO-PRO success rose from 34.9% to 96.5%.