When Grok Called Trump a "Russian Asset" https://t.co/lkU3bZdAkZ (1/5)
A fascinating USA Today article shows how Grok confidently labeled Trump as a Russian asset - perfectly illustrating the fundamental problems with how AI systems search for and interpret information. 🧵
🚨SHOCKING: Apple just proved that AI models cannot do math. Not advanced math. Grade school math. The kind a 10-year-old solves.
And the way they proved it is devastating.
Apple researchers took the most popular math benchmark in AI — GSM8K, a set of grade-school math problems — and made one change. They swapped the numbers. Same problem. Same logic. Same steps. Different numbers.
Every model's performance dropped. Every single one. 25 state-of-the-art models tested.
But that wasn't the real experiment.
The real experiment broke everything.
They added one sentence to a math problem. One sentence that is completely irrelevant to the answer. It has nothing to do with the math. A human would read it and ignore it instantly.
Here's the actual example from the paper:
"Oliver picks 44 kiwis on Friday. Then he picks 58 kiwis on Saturday. On Sunday, he picks double the number of kiwis he did on Friday, but five of them were a bit smaller than average. How many kiwis does Oliver have?"
The correct answer is 190. The size of the kiwis has nothing to do with the count.
A 10-year-old would ignore "five of them were a bit smaller" because it's obviously irrelevant. It doesn't change how many kiwis there are.
But o1-mini, OpenAI's reasoning model, subtracted 5. It got 185.
Llama did the same thing. Subtracted 5. Got 185.
They didn't reason through the problem. They saw the number 5, saw a sentence that sounded like it mattered, and blindly turned it into a subtraction.
The models do not understand what subtraction means. They see a pattern that looks like subtraction and apply it. That is all.
Apple tested this across all models. They call the dataset "GSM-NoOp" — as in, the added clause is a no-operation. It does nothing. It changes nothing.
The results are catastrophic.
Phi-3-mini dropped over 65%. More than half of its "math ability" vanished from one irrelevant sentence.
GPT-4o dropped from 94.9% to 63.1%.
o1-mini dropped from 94.5% to 66.0%.
o1-preview, OpenAI's most advanced reasoning model at the time, dropped from 92.7% to 77.4%.
Even giving the models 8 examples of the exact same question beforehand, with the correct solution shown each time, barely helped. The models still fell for the irrelevant clause.
This means it's not a prompting problem. It's not a context problem. It's structural.
The Apple researchers also found that models convert words into math operations without understanding what those words mean. They see the word "discount" and multiply. They see a number near the word "smaller" and subtract. Regardless of whether it makes any sense.
The paper's exact words: "current LLMs are not capable of genuine logical reasoning; instead, they attempt to replicate the reasoning steps observed in their training data."
And: "LLMs likely perform a form of probabilistic pattern-matching and searching to find closest seen data during training without proper understanding of concepts."
They also tested what happens when you increase the number of steps in a problem. Performance didn't just decrease. The rate of decrease accelerated. Adding two extra clauses to a problem dropped Gemma2-9b from 84.4% to 41.8%. Phi-3.5-mini from 87.6% to 44.8%. The more thinking required, the more the models collapse.
A real reasoner would slow down and work through it. These models don't slow down. They pattern-match. And when the pattern becomes complex enough, they crash.
This paper was published at ICLR 2025, one of the most prestigious AI conferences in the world.
You are using AI to help you make financial decisions. To check legal documents. To solve problems at work. To help your children with homework. And Apple just proved that the AI is not thinking about any of it. It is pattern matching. And the moment something unexpected shows up in your question, it breaks. It does not tell you it broke. It just quietly gives you the wrong answer with full confidence.
@GermanusHeiko@henninghoene@fdp Es muss endlich mal Schluss damit sein, dass es heißt. Der Politiker muss Verantwortung übernehmen. Das muss alles weg und der Bürger muss wieder ermächtigt werden fad beste für sein Land zu tun.
JUST IN AND UNUSUAL
CHINA is recording the US war live!
The US attacks on Iran have become a veritable military intelligence laboratory for China.
More than 300 Jilin-1 satellites are recording every detail, second by second, from munition refueling to missile trajectories.
China is turning US war doctrine into a database, including refueling times and air defense responses.
According to experts, this data could give China a military research and development advantage that will last for decades.
US war tactics are being thoroughly analyzed.
@MarcusFaber@FabioDeMasi Man ist dann ja auch ein Putin Lakaie. Europa fährt gerade die Produktion hoch während russische Depots sich nun geleert haben.
@unfassbar70@BrennpunktUA Die US müssen nur jedes Mal wenn Demonstranten beschossen werden 100 Köpfe der Garden exekutieren. Die US müssen das Regime nicht ändern, sie müssen nur die Unterdrückung der Revolution abschwächen.
@larsklingbeil wir brauchen eine EU-Verfassung der Willigen, mit gewählten Institutionen, die auch die Aufgaben übernehmen, für die sie gewählt sind. Nicht alle, aber bestimmte. Handel, Verteidigung Harmonisierung.
Hallo ihr Lieben, ich darf ein neues Stück meiner Arbeit vorstellen https://t.co/ER8gG12cSv Hier haben wir Konzepte aus biologischen neuronalen Netzwerken auf künstliche übertragen.
@ProfRieck wie kommen sie auf die Idee, dass es eine positive Auszahlung gibt, wenn die Ukraine den Krieg nicht gewinnt? Das bedeutet mit Nichten, dass Russland den Krieg gewinnt. Die Ukrainer wissen was Völkermord bedeutet.
Heute bei der #liegenddemo in #Jena zu Gast. Um 14:00 Uhr geht’s los. Gemeinsam für mehr Aufklärung und eine bessere Versorgung von #MECFS. In #jena in #thueringen und überall!
Die Herrschaft der Angst - Eine grobe Aufzählung
1960'er Kein Öl mehr in 10 Jahren!
1970'er Neue Eiszeit in 10 Jahren!
1980'er Saurer Regen wird in 10 Jahren alle Ernten zerstören!
1990'er Die Ozonschicht wird in 10 Jahren zerstört sein!
2000'er Die Eisschollen werden in 10 Jahren verschwunden sein!
2000 Y2k „Millennium-Fehler“ wird alles zerstören!
2001 Terror & Anthrax wird uns alle töten!
2002 Der West-Nil-Virus wird uns alle töten!
2003 SARS wird uns alle töten!
2004 Tsunamis werden uns alle vernichten!
2005 Vogelgrippe wird uns alle töten!
2006 E. coli wird uns alle töten!
2008 Der Finanz-Crash wird uns alle töten!
2009 Schweinegrippe wird uns alle töten!
2010 Die Schuldenkrise wird uns alle treffen!
2011 Naturkatastrophen werden uns alle auslöschen!
2012 Der Maya-Kalender endet: Wir werden alle sterben!
2013 Nord-Korea wird den 3. Weltkrieg beginnen: Wir werden alle sterben!
2014 Ebola wird uns alle töten!
2015 Die ISIS wird uns alle töten!
2016 Zika wird uns alle töten!
2018 Erderwärmung wird uns alle töten!
2019 CO2 wird uns alle töten!
2020 Corona wird uns alle töten!
2021 Mutanten werden uns alle töten!
2022 Der Ukrainekrieg wird uns alle töten!
2023 Das Wetter wird uns alle töten!
2024 Die KI wird uns alle töten!
2025 Ein Atomkrieg wird uns alle töten!
Wer sich heute noch ernsthaft Tagesschau und co. ansieht, kann kaum noch als zurechnungsfähig betrachtet werden.
@jreichelt Wahrscheinlich nicht. Linke wollen das Verfahren aber nicht das Verbot. Die CDU will das Verbot, weil sie dann die AfD Wähler abgreifen kann.