Breaking Math News!
Code: WTF!
Yesterday, Sabine Hossenfelder @skdh posted a comment that, "GPT 5.6 just solved one of my maths problems that GPT 5.5 insisted for 6 months is unsolvable."
https://t.co/QU4M0Lz2B0
I replied asking if she could use her chatGPT account to help resolve a conjecture I made in a paper I published in 1988, the so-called p-defect zero conjecture by running it through AI.
Overnight at least 3 different mathematicians did exactly that, finding a counter-example to the general case. For example,
"For G = PSU₃(5), p = 2, a degree-144 irreducible character has 2-defect zero, but its permutation index is n(χ) = 2." - @jonamaccarello
Thanks also to @eSlackPhysics and any others who gave it a go.
Moreover, Sabine's prompt created a full proof in the case that the group's order was divisible by at most 2 primes, which I stated at the end as a result two colleagues had shown, but I had never seen a proof.
So now they are running the code to find a "solvable" counter-example or check whether there is a counter-example for a prime other than p = 2. To dive deeper into what might be true.
All I can say is that for nearly 40 years this has been an open problem, and a few simple prompts and double-checks and chatGPT 5.6 has now solved it overnight.
Maths is in crisis -- that's for sure.
Thank you to all who worked on this, and here is a link to my original paper.
https://t.co/MQNd1UnxrF
The answer is: don't use worktrees. I haven't in months. Agents can handle conflicts. Zone your project (by domain e.g.), let an agent bottle your work into tasks and let it plan for conflict-free triaging.
If necessary replace fs tools with your ones that use write locks.
Plugin-based btw
and yes, it used graphs before.
Models were not good in multi-agent settings back then, but they are now.
Been building this since February as a distillation of ai tools I built the years before.
Might declutter the interface and release a light version
"Everything is a plugin"
Does this mean everyone can finally stop building their own ai harness and infrastructure?
GPUs must be tired of building all those custom ai harnesses and infrastructure over and over again!
guilty btw.
current state of my harness psychosis:
🧩 DeepSeek Harness v0.1 is now available in Developer Preview!
🔹 We’re opening it up to developers building agent harnesses worldwide and open-sourcing the codebase in MIT license.
🔹 Powered by the Cordis meta-framework, DeepSeek Harness is an agent harness built around one core idea: Everything is a plugin. Models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration, and UI are ALL implemented as plugins, and can be mixed, matched, replaced, and extended.
Try it now!
https://t.co/2YWSvJHhKA
Someone 'accused' me of vibe coding the LFRR simulator.
I've been coding since 1982, I've hand assembled machine code, injected fixes into running processes, worked on software that's made it to Mars and worked as a software developer for decades.
OF COURSE I USED AI CODE TOOLS
introducing fable-os,
this is not a fake bullshit os that runs in your browser - this is an agentic operating system that writes its own drivers, evolves itself, and runs on bare metal.
here it built it's own audio driver from scratch to play a sound
@Karl_Lauterbach@Stanford "dass er keineswegs dumm ist"... wenn junge gut ausgebildete Menschen auswandern, dann wegen Fremdscham für diese Ignoranz.
Als nächstes finden Sie heraus, dass Karp unter Habermas Philosophie studiert hat?
Observability driven development is the way.
It always shocks me how much secondary data is let go to waste when building. Collect it and loop it back to the agents.
This is the first step for self-healing and incremental self-improvement if done right.
people hating on europe miss that they actually have some of the top ai engineers in the world if you actually to elicit the right ones.
we're basically running the most competitive global arena for ai talent (as measured by talks and workshops), will be v interesting to see how this goes as WF enters the eval window
Now I can finally reveal why I've been quiet for so long!
We're building a radically new kind of formally verified AI for science, math, engineering, and everything else, over at @lanyon_ai. Expect more information in the coming hours and days. Link below.
Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5.
This is a 17-place jump from Kimi-k2.6 (#18 -> #1).
In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, and Content Creation Tools, landing #2 only in Gaming behind Fable 5.
The full model weights will be released by July 27.
Congrats to the @Kimi_Moonshot team on this major milestone!
Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5.
This is a 17-place jump from Kimi-k2.6 (#18 -> #1).
In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, and Content Creation Tools, landing #2 only in Gaming behind Fable 5.
The full model weights will be released by July 27.
Congrats to the @Kimi_Moonshot team on this major milestone!