𝗜𝗻𝘁𝗿𝗼𝗱𝘂𝗰𝗶𝗻𝗴 𝗧𝘄𝗶𝗻 — 𝘁𝗵𝗲 𝗔𝗜 𝗰𝗼𝗺𝗽𝗮𝗻𝘆 𝗯𝘂𝗶𝗹𝗱𝗲𝗿.
No setup. Secure. Infinitely scalable.
We just raised a $𝟭𝟬𝗠 𝘀𝗲𝗲𝗱.
After a beta with 𝟭𝟬𝟬,𝟬𝟬𝟬+ 𝗮𝗴𝗲𝗻𝘁𝘀 𝗱𝗲𝗽𝗹𝗼𝘆𝗲𝗱, we’re now opening to everyone.
RT and comment “Twin” — first agents on us. 👇
Today Thinking Machines Lab is launching our research blog, Connectionism. Our first blog post is “Defeating Nondeterminism in LLM Inference”
We believe that science is better when shared. Connectionism will cover topics as varied as our research is: from kernel numerics to prompt engineering. Here we share what we are working on and connect with the research community frequently and openly.
The name Connectionism is a throwback to an earlier era of AI; it was the name of the subfield in the 1980s that studied neural networks and their similarity to biological brains.
https://t.co/lrJioBmpbT
You're a 19 year old kid.
You are critically wounded and dying in the jungle somewhere in the Central Highlands of Viet Nam .
Its November 14, 1965 . LZ (landing zone) X-ray.
Your unit is outnumbered 8-1 and the enemy fire is so intense from 100 yards away, that your CO (commanding officer) has ordered the MedEvac helicopters to stop coming in.
You're lying there, listening to the enemy machine guns and you know you're not getting out.
Your family is half way around the world, 12,000 miles away, and you'll never see them again.
As the world starts to fade in and out, you know this is the day.
Then - over the machine gun noise - you faintly hear that sound of a helicopter.
You look up to see a Huey coming in. But.. It doesn't seem real because no MedEvac markings are on it.
Captain Ed Freeman is coming in for you. He's not MedEvac so it's not his job, but he heard the radio call and decided he's flying his Huey down into the machine gun fire anyway.
Even after the MedEvacs were ordered not to come. He's coming anyway. And he drops it in and sits there in the machine gun fire, as they load 3 of you at a time on board.
Then he flies you up and out through the gunfire to the doctors and nurses and safety. And, he kept coming back!! 13 more times!!
Until all the wounded were out. No one knew until the mission was over that the Captain had been hit 4 times in the legs and left arm.
He took 29 of you and your buddies out that day. Some would not have made it without the Captain and his Huey.
Medal of Honor Recipient, Captain Ed Freeman, United States Army, died at the age of 81, in Boise, Idaho.
God bless our vets!
what happened in 2010-15 man. we were seized by some kind of faux frontiersman cult for urbanites. probably the worst cultural era in history:
- stomp clap hey music. lumineers, mumford & sons, imagine dragons etc. awful. no redeeming qualities whatsoever
- stipped down exposed brick burger halls with signs saying "eat" on reclaimed barnwood. dim edison bulbs. drinks served out of a mason jar
- facial hair and flannel. bears and overdone mustaches popular with the soft urbanite types. lumberjack core
- relatedly, "manly" trinkets for dudes. beard oil. shaving kits. axes and leather aprons. "modern gentleman" brands in subscription boxes. packaged masculinity for consoomers. closely linked to reddit culture
- apothecarycore. every brand became a twee "Co." all using the same serif typeface
- heritage liquor and accoutrements. whisky as a personality trait. elaborate glassware setups
what the hell was this? why did everyone collectively lose their minds
It's worth remembering that DOGE is a failure at Every. Single. Possible. Level.
At the micro level, they misread, misled and straight up lied about many of the 'savings'. They'd announce '500m saved!' and it would be some contract where the money was 98% already spent and already wasn't being renewed.
At the macro level, they didn't impact spending at all. The government is spending more in 2025 than 2024, on basically the same trajectory but a little more.
At the institutional level, they didn't even convince the GOP that deficits are a problem worth caring about. The GOP just passed a bill that explodes the deficit by trillions of dollars.
At a personal level, Elon got run out of town with his tail between his legs and the most notable public facts about the Cracked Coders are that one of them was a mini Hitler, one was called 'big balls' and none of them bothered to learn how the government actually works before they lit it on fire.
They gutted a bunch of important institutions, fired whole departments at random, decimated medical research funding, killed a bunch of people dependent on USAID, and still failed at every possible level.
Officer Didarul Islam was one of four people killed in yesterday’s horrific shooting.
A Bangladeshi immigrant who joined the NYPD four years ago, he lived in Parkchester with his pregnant wife, their two young children, and his elderly parents.
When he joined the police department, his mother asked him why he would pursue such a dangerous job. He told her it was to leave behind a legacy that his family could be proud of.
He has done that, and more.
I pray for him, his family, and honor the legacy of service and sacrifice he leaves behind.
We got a call from @xai 24 hours ago
“We want to test Grok 4 on ARC-AGI”
We heard the rumors. We knew it would be good. We didn’t know it would become the #1 public model on ARC-AGI
Here’s the testing story and what the results mean:
Yesterday, we chatted with Jimmy from the xAI team, who wanted us to validate their Grok 4 score. They did their own testing on the ARC-AGI-1 & 2 public evaluation set
To validate their score (and measure possible overfitting), we self-tested the new model on our semi-private evaluation set
We walked them through our testing policy:
* No data retention
* Model checkpoint must be intended for public use
* Temporary increase in rate limits for burst testing
They were on board, so we got started
Initially, we ran into timeout errors with normal requests, so we switched to streaming. That resolved the issue
So, what do these results mean?
First, the facts: Grok 4 is now the top-performing publicly available model on ARC-AGI. This even outperforms purpose-built solutions submitted on Kaggle.
Second, ARC-AGI-2 is hard for current AI models. To score well, models have to learn a mini-skill from a series of training examples, then demonstrate that skill at test time.
The previous top score was ~8% (by Opus 4). Below 10% is noisy
Getting 15.9% breaks through that noise barrier, Grok 4 is showing non-zero levels of fluid intelligence
But the mission isn’t over. We need new ideas to solve ARC-AGI-2. Scale alone won’t get us there
Come work on ARC-AGI with us