For the past two months I've been building a new AI benchmark from scratch, based on real tasks from my own workflows across fields that current benchmarks barely touch.
The goal is a different read on model capability: strict evaluation, detailed explanations, and clear results on which model is actually best for which kind of work.
It stays 100% neutral and fair regardless of sponsorship.
I've already put over $12000 of my own money into testing the major models.
If any company wants to sponsor the project or provide credits, that means more models tested, more trials run, and a more rigorous final benchmark.
I asked Claude, GPT, and Kimi to analyze my ETH address transactions and calculate the wallet revenue from buying and selling ETH. Claude (Opus 5) and GPT said they cannot load transactions due to an Internal Error 🤷♂️.
K3 Swarm (High) did the job.
Going to double down on the MacBook mission with Omarchy. We almost have perfect coverage for the vintage Intel era going from 2009-2020. There's a straight shot to get the M1 and M2 machines going too, even if it's a lot more work. But we'll do the work. We'll fix everything.
@MIDIDesigner@Kappaemme1926 I need to run 7 copies of the project in Docker Compose—there are resources as well. Yes, 7 is the max because I don’t have time to run more 😂, but I’m sure it’s possible.