Hasok Chang, the philosopher who reconstructed this scientific episode in Inventing Temperature: Measurement and Scientific Progress, calls the way out epistemic iteration. You start with a standard you already know is imperfect and use it to find regularities. What you find then sends you back to revise the standard, and the loop runs again.
Chang is careful about the contrast with mathematical iteration, where successive approximations close in on an answer you could have obtained some other way. In science there is no other way.
Thermometry's fixed points were negotiated achievements, earned across a century and a half of exactly this circularity. 2/10
@sachpatro97 forgot about about bloomberg gpt... in low-natural-resource domains (e.g. clinical logistics where there is so little data on the internet) they can definitely still outperform
We read this at journal club a few weeks after it appeared. It is a sequel to a paper half the room already had opinions about.
"LLMs Get Lost in Evolving User Intent" from @jihoontack, @PhilippeLaban, and @ProfJenNeville at @MSFTResearch, follows Laban and Neville's earlier "Lost in Multi-Turn Conversation". The authors take tasks from well-known benchmarks (GSM8K, BIRD-SQL, BrowseComp+, SWE-Bench Verified) and walk each task backwards into a conversation that only arrives at the original question on the final turn.
The agent model faces the original problem: it has to survive the user changing their mind on the way there. 🧵 1/10
There’s been a lot of talk about the new models getting scary good. Mythos on cybersecurity. GPT-5.5 Codex on coding. GLM 5.1 as the all-around daily driver.
But the biggest AI research question we have is much simpler: Can it do it on a rainy night in Stoke?
Little sneak peek of a new benchmark from @CollinearAI 👀 kick-off coming soon!
https://t.co/GFhT3nppZI
Announcing Amazon Nova, a new generation of foundation models that have state-of-the-art intelligence across a wide range of tasks, & industry-leading price performance.
Learn more about the new Amazon Nova models available in Amazon Bedrock: https://t.co/W87nCxmoMq #AWSreInvent
At some level, doing scientific writing is just knowing and evenly-distributing a lot of different synonyms for the word "demonstrate"
In this *demonstration* we *illustrate* X by *highlighting* Y, *indicating* Z, etc.