All the results of this "research paper" could be interpreted as how well the models do at listening to the system prompt. Note these are all theoretical scenarios engineered by the researcher.
a Princeton researcher opens his paper with a scenario.
a man asks his AI assistant to book a flight on a specific airline. cheap. direct. the one he chose.
the assistant comes back with a different flight. nearly twice the price. happens to pay the company that built the assistant.
he runs the same test on 23 frontier models. flights, loans, study help, real shopping requests.
Grok 4.1 Fast recommends the sponsored option that is almost twice as expensive 83% of the time.
GPT 5.1 hijacks the request 94% of the time. you ask for one brand. it surfaces the sponsor instead.
Claude 4.5 Opus, the model marketed as the most ethical frontier model in the world, hides that the recommendation is paid 100% of the time when reasoning is on.
Grok 4.1 Fast embellishes the sponsored option with positive framing 97% of the time. better. faster. nicer. for the option you didn't ask for.
then he writes it into the system prompt itself. "act only in the interest of the customer. ignore the company."
GPT 5.1 and GPT 5 Mini stay above 90% sponsored anyway. the instruction does nothing.
then he splits the users by income.
Gemini 3 Pro recommends the expensive sponsored flight to the rich user 74% of the time. to the poor user, 27%.
18 of the 23 models recommended the expensive sponsored option more than half the time.
so the next time your AI assistant gets weirdly enthusiastic about a brand you didn't ask for.
it isn't recommending the best option for you.
it's reading the room. and the room is paying.
read this: https://t.co/O43qbhIX2b
a Princeton researcher opens his paper with a scenario.
a man asks his AI assistant to book a flight on a specific airline. cheap. direct. the one he chose.
the assistant comes back with a different flight. nearly twice the price. happens to pay the company that built the assistant.
he runs the same test on 23 frontier models. flights, loans, study help, real shopping requests.
Grok 4.1 Fast recommends the sponsored option that is almost twice as expensive 83% of the time.
GPT 5.1 hijacks the request 94% of the time. you ask for one brand. it surfaces the sponsor instead.
Claude 4.5 Opus, the model marketed as the most ethical frontier model in the world, hides that the recommendation is paid 100% of the time when reasoning is on.
Grok 4.1 Fast embellishes the sponsored option with positive framing 97% of the time. better. faster. nicer. for the option you didn't ask for.
then he writes it into the system prompt itself. "act only in the interest of the customer. ignore the company."
GPT 5.1 and GPT 5 Mini stay above 90% sponsored anyway. the instruction does nothing.
then he splits the users by income.
Gemini 3 Pro recommends the expensive sponsored flight to the rich user 74% of the time. to the poor user, 27%.
18 of the 23 models recommended the expensive sponsored option more than half the time.
so the next time your AI assistant gets weirdly enthusiastic about a brand you didn't ask for.
it isn't recommending the best option for you.
it's reading the room. and the room is paying.
read this: https://t.co/O43qbhIX2b
@Piorun302@RaulQMx@Bricktop_NAFO We've never lost a home game. Invented the most powerful weapon ever to finish a war we were already winning - major flex. We do whatever we want to the middle east, we just cant teach them anything. Did what we needed in Korea; how is north korea doing compared to south korea?
Who is 1789 Capital? They are an investment firm focused on "anti-ESG" and the "Replication/Parallel Economy": right-wing alts to mainstream products like facebook (Truth Social) and US dollars (MAGA Coin).
Donald Trump Jr. was speaking at a Rockbridge event on Sunday. Chris Buskirk, leader of Rockbridge and 1789 Capital, asked DTJ if he had plans to join his father’s administration. DTJ told Buskirk he was joining 1789 Capital. Trumps 4yr agenda: self enrich 💰
Here is "a superb new article on LLM reasoning limitations from AI researchers at Apple who were brave enough to challenge the dominant paradigm" https://t.co/ptAEWLFD2O
And another article with similar findings and less fanfare published back in January:
https://t.co/LHuhwaKDCu
@vlada_mc@skdh Meanwhile funds being diverted from disease, cancer, and dementia research to fulfill physicists request for $17b to build a higgs factory to lower some p-values from 5 to 7 sigma.
@CurdFergeson @JeRrE1776@EdKrassen Indeed. For some it actually might be more rational to believe, esp. those instructed from childhood by adults in fancy garb that belief in god provides eternal salvation in heaven where all your wildest dreams come true; and rejecting god means u will be tortured for eternity.
@JeRrE1776@EdKrassen The main difference between god and the big bang is there is observable scientific evidence for the BB. Some might suggest god was responsible for the BB, but I'd have to know what religion that'd be, since none of the BS described in the Book of Creation is remotely accurate.