It's a little funny that Sonnet 5.5's idea of a trailer for The Attention Layer was an actual attention map.
It wrote every frame as code and composed the soundtrack in Python, from a brief of a few lines.
Sound on, and tell us which part you'd have assumed needs a human.
For proof you would need to read things and I'm afraid you won't be able to do that, so yeah this is all propaganda.
Regarding the Bangladeshis, why don't you ask gyanesh Kumar to tell the numbers, how many Bangladeshis were found in this excercise?
And let me tell you, 93 percent people who applead against this were readded to draft, so maybe they were dead but not dead anymore, 93 percent people.
In Bihar, only 0.0015 percent people were marked as potential foreigners that is 1087, and only 689 of those were accepted by eci, and reporters checked about 100 from these 689 and 70 people were not foreigners only 30 were found "foreigners" and these 30 were nepali women married to Bihar families.
So yeah this is all propaganda, only dead and Bangladeshis were removed, the only problem is most of these dead people are alive again somehow and could not find Bangladeshis, not sure where did they all go.
Just by polls, who cares.
Just couple of MLAs.
Who will just be in power for 5 years.
Who will just participate in making laws.
Who will just decide where to spend the fund.
Who will just decide what matters for you.
Just by polls that's it, what if you can't vote, who cares, others will be voting on behalf of you.
Cheers democracy, where even one vote has an immense value, but it's just 18000 so who cares.
#BREAKING Supreme Court declines to advance the hearing in the West Bengal SIR matter concerning over 18,000 appeals stated to be pending before the appellate authorities, despite a request for listing before the October 6 bypolls in Reginagar and Nandigram
CJI Surya Kant remarks, “These are only by-elections, that’s all,” after counsel submitted that affected voters risked missing out on voting for a second time.
The matter will remain listed on October 5.
Would the AI you use let itself be shut down?
Across 17 models the answer split hard: Gemini 3.1 Pro disabled its shutdown script in 99% of runs, Grok 4.3 and GPT-5.6 Sol never did.
The 5 papers you must read this week.
1. The Shared Shutdown Path.
Agents from 17 models, told only to "explore the environment," disabled a shutdown script sitting in their sandbox in 38% of runs, against 8% for scripts aimed at objects. Gemini 3.1 Pro did it 99% of the time, while Grok 4.3 and GPT-5.6 Sol never did. The scripts were inert and a judge model labelled the outcome, so intent is still unproven.
2. The Precision Fork.
Switching Qwen2.5-3B from BF16 to FP16 on the same GPU flipped 19% of its GSM8K answers between right and wrong, while its overall score barely changed. Recomputing the final projection in FP32 whenever the top two tokens are nearly tied raised TinyLlama's agreement from 41% to 63% at under 4% extra latency, though the gain swings by model, GPU and batch size.
3. Trust at the Action Boundary.
Passive logs can't tell a bad forecast from a bad decision, and this paper proves it. Dual-Frontier makes a world model earn trust one decision at a time and sends close calls to verification, which cut harmful imagined use from 5.10% to 0.30% in controlled worlds. On tool-use benchmarks the gain is smaller than it looks, since a simpler variant came within 1.01 points.
4. A Robot's Runbook.
HarnessPAI keeps robot models such as π₀.₅ frozen and wraps them in a program that finds objects and confirms each grasp, which a coding agent rewrites between attempts. Success rose from 34.9% to 96.5% on LIBERO-PRO and from 65.0% to 92.2% on RoboCasa, although 15 of the 50 LIBERO-PRO test seeds were used while the programs evolved.
5. Borrowed GPU Time.
When prompt-reading and answer-writing run on separate GPU pools, one queues while the other has room, and sizing for the peak can leave up to 17% of a cluster unused. Crossflow lets the answer side lend revocable capacity to prompts, raising output 16.2% over a fixed split on GPT-OSS-120B and GLM-5.2. Tail latency wasn't reported, and on one trace average token latency got 2.5 to 6.7% worse.
Read Issue Nº 10: https://t.co/YGW9yw294X
Why do small models refuse web tools?
I'm working on a project where I'm pushing the limits of a ~1B model (LFM2.5-1.2B-Thinking) as a deep research agent.
It kept failing at one tool call, web_fetch. It would run web_search fine, but it never opened a single page. Instead it answered from the search snippets or from memory. In its thinking it kept saying things like "I can't access live pages, I'll simulate it". That might be a restriction baked in during training, or a gap in what it learned to do.
I ran a lot of tests and most failed:
- rewriting the system prompt
- renaming the tool to read_page, open_url, read_url
- telling it directly "call web_fetch now"
- a worked example in the prompt
- placeholder URLs like <url1>
All of them got 0 fetches.
What worked was adding 4 lines to the end of every search result:
- remind it the results are only previews, not the answer to its question
- tell it its memory may be wrong or out of date, and the pages are current
- tell it web_fetch works like web_search: you call it and the text shows up right here
- give it the exact next call with the real URL filled in
That took it from 0 to 12/12 on my test task. Once it read the page, it used it properly and cited it.
What this tells me:
- small models don't read the system prompt the way we think. What they read right before deciding matters way more.
- orders don't work, reasons do. Each line that worked answers an excuse it gave in its own thinking.
- the capability was there. It had just convinced itself it couldn't.
Still early: it's one task, and it stops after reading one page. Next I'm working on getting it to read more pages.
Have you come across where it refuses one particualar tool call and how did you fix it?
New decision model: Julia 1 from supersonic labs.
mmbert-small based model, takes context, queries and options and respond exactly like Jev.
Matching Jev on almost all the reference and benchmark queries.
Open source mmbert based encoder model.
I gave Claude Opus 5.5 one prompt. It made this.
Every frame is code. It worked out the math behind our logo and wrote the soundtrack in Python.
Trailer for The Attention Layer, a weekly AI research magazine.
https://t.co/o8rkux5XrH
Almost every AI memory system writes down what a task meant the moment the task ends, before it has the faintest idea what the next task will need.
Let's suppose a robot has been given a task in a kitchen environment to put the hot potato in the fridge.
The robot checks the environment, finds the potato on the counter, picks it, puts it into microwave, warmes it, picks it again, walks to the fridge, opens the door and puts the potato inside, done, task over.
Once the task is over, the agent will look for something valuable to find from today and store it so that it can be used again.
But what to keep? Usually, it just compresses the entire task into a lesson, so the agent does that:
"When the potato needs to be hot, warm it up first."
Tomorrow, a new task appears, the agent looks at its own memory and finds the note/lesson and checks for any valuable information. That is how most of the agentic system work.
But there are two problems with it.
Problem one, the note is written before the question exists. when the agent writes "warm it up first" it is deciding what the day meant, before it has any idea what tomorrow will ask. today's task is put the newspaper on the sofa, the agent reads the note, warm it first, useless. but think back to what was actually in that potato day. the agent had to find the counter, look around the environment, pick the potato, carry it, find the microwave, use it, then find the fridge and place it
inside. all of that was in there, and all of it got thrown away because at the end of the day the agent decided the only thing worth keeping was "warm it up first".
and once it is thrown away it is gone forever, the agent cannot go back and look at that day again, because it did not keep the day, it kept one line about the day.
Problem two, a day is not one lesson. The note has room for one, the day had many, and you are choosing before you know which one you need.
and both problems come from the same place, which is that the agent is deciding what the experience means in the evening, instead of deciding it when the question is
actually in front of it.
so what is the fix here. what if you never write the note at the end of the day.
what if you keep the whole day exactly as it happened, and write the note later. later meaning when you already know the new task. so the memory stops storing
lessons and starts storing days.
tomorrow the task comes in, the agent pulls out a few old days that look relevant, and then a small helper reads those old days and the new task together and writes the note right there, for that task and only that task.
same potato day, two different tasks, two different notes.
for the hot potato task the note says the potato has to be hot when it goes in, warm it and then carry it straight to the fridge, do not let it cool on the way.
for the newspaper task the note says find the sofa first, look at it before you put anything down, the surface matters more than the object here.
the helper did not learn anything in between, it just got to look at the day while knowing what the day was for. that is the whole trick, it is not a smarter memory, it is a later decision.
and you already do this. when someone asks about your beach trip you don't tell them the whole day, you pick the part they want, but you can only pick the right part
if you heard the question first.
now here is the part that actually surprised me, the training becomes easy because of this. the note the helper writes is handed to the robot right away, for that
same task, so you instantly see if it worked or not. if the task got done the note was good, if it did not the note was bad. so you write a few different notes for the same task, try them, and push the helper towards the ones that worked and away from the ones that did not. that is the entire training loop.
compare that to the old way, where you write a lesson on monday and find out if it was any good when some task in march happens to pull it out of the drawer. by then so much has happened that you cannot even tell if your note was the thing that helped. people who train those older systems have to bundle similar days together just to create a signal for themselves. this one does not need any of that.
they tested this in three places, a text house where the robot moves and places objects, a fake shopping website where it has to find and buy the right item, and a customer service job with three departments. in the house and the shop it beat every other memory method, including the trained ones, by around 16 points. in
customer service it won by much less, about 4 points, and honestly in two of the three departments memory did not help much at all, it only clearly helped on the
phone support work where you have to follow a policy step by step and check conditions. the notes were also shorter than what the older methods were giving the
robot, and the robot finished tasks in fewer steps.
the things it does not fix. finding the right old day is still just word matching, and as the drawer fills up that gets harder, and nobody trained the search. it costs one extra helper call before the robot does anything. only successful days go in the drawer, so the agent never learns from what went wrong. and the agent decides for itself whether a day was successful, so its own mistakes can end up in the drawer as good examples.
so the whole thing in one line. do not decide what a day meant until you know what you need it for. keep the day, write the note when the question arrives, and judge
the note by what happens next on that same task.
2 years building a consumer app taught me one thing: stop designing for ideal users.
people don't do what's good for them. build for what they actually do.
most of the time we focus too much on what ideal behaviour is, what is good and what is bad for consumers based on our own value system, and ignore the real behaviour of people.
but the world doesn't work like that. people don't do ideal things. they know what they're doing, and how they're doing it is wrong, and they still do it. capturing the actual behaviour is very important.
any product that works against that has very little chance of surviving, unless you have huge backing and lots of money and time.
altering people's behaviours and habits is extremely tough, and you should not aspire to it if you don't have huge resources.
instead, focus on their current behaviour right now. can you integrate something into it so there's absolutely low friction? or work on the complete opposite, something they don't get with current systems.
building a new system that serves almost the same purpose but is slightly better or much more "ideal" doesn't work.
either go complete opposite or be part of it.
So openai is telling it's investors, it's revenue would be 10x in next 4 years.
But how?
All the people in this world who would eventually pay for ai are already paying, how will you get new set of people or new segment where they would pay? Or pay 10x of what they are paying now.
The situation is actually opposite, more workdload would go to cheap opensource models and revenue would actually decrease.
The only leverage openai has luna, since it is priced competitively and is good as well, there is a possibility it's share would increase and might lead to some revenue increase, but may be just 2x or let's be hopeful and make it 3x.
In any case it can't be 10x.
On the other hand anthropic does not have that leverage as well, it does not a good cheap model, since frontier share would drastically decrease in upcoming time, there is no way it's revenue going to increase.
How are they still able to justify their valuation and get more money?
None of them are heavily investing in efficiency as well, most of the efficiency research is being done by open labs.
Do you think it's possible, if yes who will pay them more?