@NethermindSec@Particle_CS@nethermind had been such a great partner for us and we are very happy to had chosen working with them.
The audit team is exceptionally kind and professional. Together with the best AI auditing infrastructure it make them a great choice for any protocol out there.
❤️
We at @Particle_CS did a very similar process when we built our protocol.
- sanity in early development
- threat modeling when shape is right
- invariant + fuzz
- AI agent runs #auditagent
- AI bounty competition #agentarena
- full audit @NethermindSec
Final result - clean audit
https://t.co/FgjdBL2ZJR
building with #GodotEngine is really fun in the age of agents!
Love the web support @godotengine, it so cool to be able to run simple 3d graphics same in the browser
Here's what we built at ETHOnline 2026 hosted by @ethglobal:
🏦 Branch Zero – A walkable 3D bank where every desk is a real smart-account op, no wallet pop-ups after one consent.
Game is live
https://t.co/eFNjeGkYfH
Check it out here: https://t.co/FjFsXJweIt
Here's what we built at ETHOnline 2026 hosted by @ethglobal:
🏦 Branch Zero – A walkable 3D bank where every desk is a real smart-account op, no wallet pop-ups after one consent.
Game is live
https://t.co/eFNjeGkYfH
Check it out here: https://t.co/FjFsXJweIt
As a customer using all major LLM providers on a daily basis (@grok , @ChatGPT, @claudeai).
i was surprised to see how well @OpenAI dominate the economic frontier providing a very strong intelligent per cost value.
LLM + Harness = effective intelligence
so this is just one side of the equation, but still super impressive results by open ai
The interesting part isn’t the full ranking.
It’s what happens when you remove every economically dominated option:
If another configuration is equally smart or smarter and cheaper, why use it?
The resulting Pareto frontier is surprisingly small 👇
Which LLM should you use — and at what reasoning setting?
I compared 32 configurations across OpenAI, Anthropic and xAI using one metric:
Intelligence per Dollar.
The goal: find the cheapest model that’s smart enough for the task.
The result surprised me. 🧵
This is what I actually want to solve next.
I work across Cursor, Claude and Codex, and I want to understand which:
tool + model + reasoning setting
is optimal for coding, debugging, architecture, research and agentic work.
Would a per-tool ranking be useful?
My main takeaway:
Don’t default to the smartest model.
Use the cheapest configuration that is intelligent enough for the task - and escalate only when the task actually needs it.
Model selection becomes a routing problem.
Another one:
Claude Opus 5 Max
54 intelligence → ~$4.21/task
GPT-6 Astra xHigh
54 intelligence → ~$1.85/task
Same benchmark score at less than half the cost.
This is why “which model?” isn’t enough.
A few examples show why model + reasoning setting matters:
Grok 4.6 Medium
49 intelligence → ~$0.95/task
GPT-6 Astra Low
49 intelligence → ~$0.63/task
Same benchmark intelligence. ~34% cheaper.
For anyone who wants to audit the result:
• Artificial Analysis Intelligence Index v4.2
• Measured cost per benchmark task
• Intelligence cutoff ≥40
• 32 model + reasoning configurations
Full ranking + methodology below 👇
The interesting part isn’t the full ranking.
It’s what happens when you remove every economically dominated option:
If another configuration is equally smart or smarter and cheaper, why use it?
The resulting Pareto frontier is surprisingly small 👇