Someone in a meeting asks "who else uses this?" Normally that's a day of grepping, reading and asking around, and you're still not sure.
I've watched an agent answer it before the meeting moved on. Not a smarter model. It just had somewhere to look.
https://t.co/LrYo63jAA1
Still surprises me how few teams write their stories down — good teams too. Everyone nods in the meeting, and by Monday it's four different features in four heads. I've done it properly for years; it's half an hour. Wrote up how, with the examples I use:
https://t.co/w7UxXepvK5
Agents, tools, skills, MCP, plugins, connectors: seven terms, mostly branding. I sorted out which is which by building one boring invoice bot on paper — including the step where the right answer is to build nothing:
https://t.co/P2hwYahTdL
Your search endpoint is a POST that lies. It doesn't change anything, but every CDN and retry library assumes it does.
HTTP finally shipped a method that doesn't lie. It only took two attempts and eighteen years.
https://t.co/lHxhFpeAPl
In 1986, a reactor exploded because of a combination nobody on the floor had been told about. We're running a similar experiment with AI in software now — softer stakes, same gaps.
I've some field notes.
https://t.co/SkLbsWcMsz
The revolution is here. The instructions aren't.
Wrote about AI, fear, and why we will adapt — not because we're fearless, because we don't have a choice.
https://t.co/WV2ZpTIbpR
I spent 10 years writing code. Then I became an engineering manager and had to unlearn most of what made me good at it.
The hardest refactor has nothing to do with code.
https://t.co/P3CRD5Ow9A
Bugün yapay zekaya verilen en zor sınav açıklandı. En iyi model %0.37 aldı.
ARC-AGI-3 yayımlandı. Sonuçlar:
→ Gemini 3.1 Pro Preview: %0.37
→ GPT 5.4 (High): %0.26
→ Claude Opus 4.6 (Max): %0.25
→ Grok-4.20: %0.00
İnsanlar aynı testte ~%100 alıyor.
ARC-AGI, yapay zekanın ezberleme değil, gerçekten "düşünüp düşünemediğini" ölçüyor. Hiç görmediği bir problemi çözebiliyor mu?
Geçmişe bakalım:
ARC-AGI-1 (2019): Modeller yıllarca %34'te takıldı. Sonunda o3, milyonlarca dolarlık hesaplamayla %88 aldı ve insanı geçti.
ARC-AGI-2 (2025): Yeni versiyon çıktı, modeller yeniden dibe vurdu. Başlangıçta %5'i kimse geçemedi. Aylar sonra en iyi özel sistem ancak %54'e ulaşabildi.
ARC-AGI-3 (2026 bugün): Artık statik bulmaca bile yok. Model bir ortama girip hareket etmek, hipotez kurmak, deneme-yanılma yapmak zorunda. Tam olarak insanın yaptığı şey.
Sonuç: %0.37.
----------------
Acaba ne zaman kendisinin farkında olan bir model göreceğiz?
@mertxdigital Zaten aslinda ne kadar pahalı oldugunu anlatmış. Hangi sirket bir insanı 20 yil okutup, besleyip ise aliyor? Sadece maas veriyor. Insan maaşıyla karsilastiramadigi icin insanin maliyetini şişirmiş.