A party which produces a public servant like Donald Trump, in the United States or any honorable society, is an astonishing failure as a political institution.
Hundreds of millions of people throughout the world are immersed in extreme poverty. Yet, disproportionate wealth remains in the hands of a few. It is an unjust scenario, in the face of which we cannot fail to question ourselves and commit to change things. There is no lack of resources at the root of disparities, but the need to address solvable problems related to a more equitable distribution of wealth, to be achieved with moral sense and honesty.
@elonmusk grok's bad at evaluating statistics:
"The shift from 100% in 1400 to **2โ6%** today (most sources converge around 2โ4% for effective reservation/trust control in the lower 48) reflects centuries of treaties, wars, forced removals, land cessions, and policies like the Dawes Act."
@rch371 The imagined story of a people longing for a place on a planet started the war, a fanciful literary grant unaware of the damage to our ecology to follow. Better social paradigms exist. We canโt let an early, parochial, draft of the human story be the end of the story.
@ShaneRosengren @autocorrect2_0 Haven't you read Numbers?
Why did they need to scout a land which was theirs from the beginning?
"We went into the land to which you sent us, and it does flow with milk and honey! Here is its fruit."
See: Numbers 13:27 (OT; NIV)
Damn. AI was not designed to be improved by its users, a colossal engineering mistake.
grok: "But yes, under the old/current dominant paradigm, almost all of that value was thrown away afterward instead of being turned into permanent model improvement."
https://t.co/dpEJY26UiP
๐จBREAKING: Princeton just proved that AI agents are throwing away
the most valuable data they'll ever collect.
And nobody noticed because it looks like normal conversation.
Every time an AI agent takes an action, it receives what researchers
call a "next-state signal." A user reply. A tool result. A terminal
output. A test verdict.
Every existing system takes that signal and uses it as context for
the next response.
Then discards it forever.
The Princeton team just proved this is one of the most expensive
mistakes in AI engineering. Because that signal contains two things
nobody was extracting.
First: an implicit score. A user who re-asks a question is telling
you the agent failed. A passing test is telling you it succeeded.
A detailed error trace is scoring every step that led to it. This
is a live, continuous reward signal hiding inside every interaction.
Free. Universal. Completely ignored.
Second: a correction direction. When a user writes "you should have
checked the file first," they're not just saying the response was
wrong. They're specifying which tokens should have been different
and how. That's not a scalar reward. That's token-level supervision.
And scalar rewards throw every single bit of it away.
They built a system called OpenClaw-RL around recovering both.
Then they ran the experiment that changes everything.
An agent started with a personalization score of 0.17. After just
36 normal conversations, with no new training data, no labeled
dataset, and no human annotations, the combined method hit 0.81.
The agent didn't get retrained. It got used.
That's the part nobody is talking about. The model was serving live
requests at the same time it was being trained on them. Four
completely decoupled loops running simultaneously. Policy serving.
Rollout collection. Reward judging. Weight updates. None waiting
for the others.
The agent gets smarter every time someone talks to it.
And the deeper the task, the more it matters. On long-horizon
agentic tasks, outcome-only rewards give you a signal at the very
end of a trajectory and nothing in between. Their process reward
model scores every single step using the live next-state signal as
evidence. Tool-call accuracy jumped from 0.17 to 0.30. GUI accuracy
improved further on top of that.
This creates a shift nobody has fully reckoned with yet.
The current paradigm: collect data offline, train in batches,
deploy, hope it works.
The new paradigm: deploy, extract training signal from every
interaction, update continuously, improve automatically.
Every conversation is training data. Every correction is a gradient.
Every re-query is a reward signal.
The agents that figure this out first won't need bigger datasets.
They'll just need more users.
@ItakGol Damn. AI was not designed to be improved by its users, a colossal engineering mistake.
grok: "But yes, under the old/current dominant paradigm, almost all of that value was thrown away afterward instead of being turned into permanent model improvement." https://t.co/dpEJY26UiP
๐จBREAKING: Princeton just proved that AI agents are throwing away
the most valuable data they'll ever collect.
And nobody noticed because it looks like normal conversation.
Every time an AI agent takes an action, it receives what researchers
call a "next-state signal." A user reply. A tool result. A terminal
output. A test verdict.
Every existing system takes that signal and uses it as context for
the next response.
Then discards it forever.
The Princeton team just proved this is one of the most expensive
mistakes in AI engineering. Because that signal contains two things
nobody was extracting.
First: an implicit score. A user who re-asks a question is telling
you the agent failed. A passing test is telling you it succeeded.
A detailed error trace is scoring every step that led to it. This
is a live, continuous reward signal hiding inside every interaction.
Free. Universal. Completely ignored.
Second: a correction direction. When a user writes "you should have
checked the file first," they're not just saying the response was
wrong. They're specifying which tokens should have been different
and how. That's not a scalar reward. That's token-level supervision.
And scalar rewards throw every single bit of it away.
They built a system called OpenClaw-RL around recovering both.
Then they ran the experiment that changes everything.
An agent started with a personalization score of 0.17. After just
36 normal conversations, with no new training data, no labeled
dataset, and no human annotations, the combined method hit 0.81.
The agent didn't get retrained. It got used.
That's the part nobody is talking about. The model was serving live
requests at the same time it was being trained on them. Four
completely decoupled loops running simultaneously. Policy serving.
Rollout collection. Reward judging. Weight updates. None waiting
for the others.
The agent gets smarter every time someone talks to it.
And the deeper the task, the more it matters. On long-horizon
agentic tasks, outcome-only rewards give you a signal at the very
end of a trajectory and nothing in between. Their process reward
model scores every single step using the live next-state signal as
evidence. Tool-call accuracy jumped from 0.17 to 0.30. GUI accuracy
improved further on top of that.
This creates a shift nobody has fully reckoned with yet.
The current paradigm: collect data offline, train in batches,
deploy, hope it works.
The new paradigm: deploy, extract training signal from every
interaction, update continuously, improve automatically.
Every conversation is training data. Every correction is a gradient.
Every re-query is a reward signal.
The agents that figure this out first won't need bigger datasets.
They'll just need more users.
This is politically basic to me, but the minimum requirement of a political party, in its role as a public service, is to leverage common knowledge to make everyone better off. When the common knowledge is corrupted by a political party, so is the public purpose.
Today, Michelle and I are proud to announce that we will be hosting the dedication ceremony for the Obama Presidential Center on June 18thย in Chicago, and welcoming the public on June 19th.
We canโt wait for you to visit. Go to https://t.co/swMRHuB5Y4 to learn more.
@autocorrect2_0@StevenN23402223 It's a shame human priorities aren't flipped to appreciate our actual experience, and dedicated to making those concrete experiences the best they can be for everyone.