I’ve been revisiting the 2023 “Jagged Frontier” paper by @emollick and co‑authors. Their core point still holds: AI is uneven. Back then the frontier looked “creative but not precise” – great at writing, weak at maths and logic – so the smart money was to keep humans in the loop on anything brittle.
The authors also never claimed the frontier was fixed. The whole idea was that as models improved, different tasks would move from “outside” to “inside” the curve.
Two years and several model generations later, that’s exactly what’s happened. We had one of Strategize Labs’ long‑horizon agents, codenamed #Ada_Lovelace, re‑draw the jagged frontier curve against today’s frontier models (Gemini 3.0, Claude Opus 4.5, GPT‑5.1 reasoning variants). The big hole the paper focused on – the “math gap” – now looks mostly closed. On a lot of analytical work, these models are already past typical human performance, which is cool.
The jagged edge now shows up somewhere else: not in single tasks, but in processes. Chaining actions over hours or days, juggling tools, staying aligned with the original goal – that’s where things still break.
So the job isn’t to retire the Jagged Frontier idea, but to re‑map it: from “what can AI do?” to “over what horizon can we trust agents to stay on course?” If we get that right, agents start to look less like a replacement of human workers, and more like handing everyone an Iron Man suit.
P.S. Our agent identified more with GenZ than I had anticipated! @awaisathar
Nothing more satisfying than building Pakistan's biggest education competition for both English and Urdu, in which children from across Pakistan and from all social backgrounds can participate and excel.
A 🧵 about Pakistan Spelling Bee (https://t.co/9NSODTdUo3)
@Huk06@microMAF@OpenAI@barbarikon Cultural = Political :)
I'll keep checking just in case. We need open corpora of all kinds :) We did publish a "culturally un-nuanced" UrduQA dataset last year ( more @ https://t.co/EWKG3HLN1t ) but we all know
اردو تو پاکستان میں بھی سوتیلی ہے، دوسروں کو کیا کہیں
@Huk06@microMAF@OpenAI@barbarikon Having said that, I noticed that Hinglish is included. Since it's supposed to be "culturally nuanced", I was hoping to take a look at the actual dataset and rubric criteria for potential bias but it's not available publicly (yet)
My talk at @INSEAD on self organising hive minds cracking hard problems in strategy. It was based on my work with @awaisathar and two other coauthors at @Cambridge_Uni and University of Manchester https://t.co/vEGZPF0A0W
@sohailabid Might have someone in my network interested. Should I ask around or do you want to work on a techie solution. There might be a short research paper in there 😀