First local agent to beat Hermes on benchmarks!
✦ runs Qwen, Gemma, Llama via llama.cpp
✦ stable-prefix caching keeps sessions cheap
✦ TurboQuant cuts the KV-cache 6.4× smaller
✦ 37 tasks solved vs Hermes' 31 on GAIA Level 1
Open source on macOS, Windows & Linux 👇
Grok 4.5 crushed OpenAI & Meta in this test!
Cost per run:
GPT Sol $1.63
Grok 4.5 $2.47
Meta Muse Spark 1.1 $1.08
The three prompts:
• Fruit Ninja – style slicing game
• Angry Birds – style fort collapse
• Crossy Road clone
Grok nailed all three: clean physics, smooth playback, good visuals. GPT Sol did fine until Crossy Road, where it froze. Meta Muse Spark was cheapest and it showed: its Crossy Road lagged badly.
The question isn't which model is cheapest, it's what a broken output costs you in reruns, wasted time, and things you can't ship. Cheap stops being cheap when it doesn't work.
Grok 4.5 crushed OpenAI & Meta in this test!
Cost per run:
GPT Sol $1.63
Grok 4.5 $2.47
Meta Muse Spark 1.1 $1.08
The three prompts:
• Fruit Ninja – style slicing game
• Angry Birds – style fort collapse
• Crossy Road clone
Grok nailed all three: clean physics, smooth playback, good visuals. GPT Sol did fine until Crossy Road, where it froze. Meta Muse Spark was cheapest and it showed: its Crossy Road lagged badly.
The question isn't which model is cheapest, it's what a broken output costs you in reruns, wasted time, and things you can't ship. Cheap stops being cheap when it doesn't work.
GPT 5.6 Sol & Terra just made Fable 5's pricing look like a joke!
We gave OpenAI's top models (Sol, Terra) and Anthropic's top models (Fable 5, Opus 4.8) the same 3 prompts:
• Supernova boom
• Meteor hitting a city
• Solar system model
The bill:
Sol: $4.77
Terra: $1.24
Fable 5: $9.94
Opus 4.8: $2.46
Outputs came out surprisingly close. The prices didn't. Fable 5 cost unreasonably more than everything else, with Sol not far behind. Considering latest Fable 5 nerf, it's hard to see what you're paying for.
HOLY F*CK THE MYTHOS-CLASS MODELS ARE JUST BUILT DIFFERENT
Fable 5 is back on @AIMLAPI, and the benchmark results are just WILD.
A recent test gave Sonnet 5 and Fable 5 the exact same 70-airport flight dataset and asked for a 3D globe in a single HTML file.
→ Sonnet 5 ($0.10):
it mapped the routes on a dark wireframe 🤡
→ Fable 5 ($0.77):
it built an actual Earth, textured oceans *and* an atmospheric glow 💥
An extra $0.67 for a masterpiece? Easiest money you'll ever spend!
Docs and setup in 🧵↓
Fable 5 is BACK on AI/ML API
We gave Sonnet 5 and Fable 5 the exact same prompt and same real flight dataset of 70 airports and 435 routes pulled from flightradar. Then asked each one to turn it into a cinematic 3D globe as a single HTML file.
Outputs:
Sonnet 5: 9.8k tokens, $0.10
Fable 5: 15k tokens, $0.77
Sonnet 5 came in 87% cheaper. It drew the routes, but the planet underneath them barely exists: a dark wireframe with the arcs floating in space. Fable 5 built an actual Earth: textured oceans, ice caps, atmospheric glow. Mythos-class models are truly a masterpiece.
Docs, guides & setup below:
Gemini Omni Flash Preview is now live on AI/ML API!
Gemini Omni vs Seedance 2.0. We gave both models a GTA 6 night chase and asked for photorealistic, anime and LEGO versions.
Looks like Seedance 2.0. is still unmatched.
Both ran on one AI/ML API key.
Docs, guides & setup below:
Gemini Omni vs Seedance 2.0
We gave both models a GTA 6 night chase and asked for photorealistic, anime and LEGO versions. Seedance wins on motion physics. Gemini wins on style switching.
Both ran on one AI/ML API key. Docs, guides & setup below:
We gave 4 AI models a secret word each and told them to steal the others.
GPT, Claude, Grok and DeepSeek bluffed and interrogated each other while an AI judge scored the smartest player. GPT won by barely speaking: it planted one idea early and let the rest expose themselves.
All four ran on one AI/ML API key. Code and full transcript below:
We gave 4 AI models a secret word each and told them to steal the others.
GPT, Claude, Grok and DeepSeek bluffed and interrogated each other while an AI judge scored the smartest player. GPT won by barely speaking: it planted one idea early and let the rest expose themselves.
All four ran on one AI/ML API key. Code and full transcript below:
We gave 4 AI models a secret word each and told them to steal the others.
GPT, Claude, Grok and DeepSeek bluffed and interrogated each other while an AI judge scored the smartest player. GPT won by barely speaking: it planted one idea early and let the rest expose themselves.
All four ran on one AI/ML API key. Code and full transcript below:
Anthropic engineer:
"At Anthropic, >80% of our engineers are building with self-improving loops, in 3-6 months, it will be 100%"
Same lesson shows up in this agent game
One tiny rule created strategies nobody programmed
Interrogation, bluffing, pressure, silence - all emerged on their own
An LLM-as-judge picked the winner
The quiet agent won
Loop + agents + incentives + judge - that's the stack
Tiny rule, emergent strategy, big lesson for builders
Bookmark and watch full interview
@OpenAI he most interesting part isn't the models themselves it's dropping all three at once and seeing which one people actually use
https://t.co/4sDEgejold
We gave 4 AI models a secret word each and told them to steal the others.
GPT, Claude, Grok and DeepSeek bluffed and interrogated each other while an AI judge scored the smartest player. GPT won by barely speaking: it planted one idea early and let the rest expose themselves.
All four ran on one AI/ML API key. Code and full transcript below:
We gave 4 AI models a secret word each and told them to steal the others.
GPT, Claude, Grok and DeepSeek bluffed and interrogated each other while an AI judge scored the smartest player. GPT won by barely speaking: it planted one idea early and let the rest expose themselves.
All four ran on one AI/ML API key. Code and full transcript below:
We gave 4 AI models a secret word each and told them to steal the others.
GPT, Claude, Grok and DeepSeek bluffed and interrogated each other while an AI judge scored the smartest player. GPT won by barely speaking: it planted one idea early and let the rest expose themselves.
All four ran on one AI/ML API key. Code and full transcript below:
API-first email built for AI agents
One prompt to plug in via MCP or Agent Skill
Your agent gets its own inbox – and can run any workflow over email
Free in open alpha - link in comments
GLM-5.2 is now available on AI/ML API!
We asked it and Opus 4.8 to one-shot a playable Backrooms, same prompt.
Opus got the flashlight working but you can't run and can't pause. GLM-5.2 built the full mechanics.
Opus: 2:14sec & $1.94
GLM: 1:08sec & $0.37
Docs, guides & setup below:
Kimi K2.7-Code is now available on AI/ML API!
Moonshot's latest is built for long-horizon agentic coding that self-corrects instead of one-shotting. So we gave it a hard one.
SpaceX is going public, a company built on one habit: fly, fail, fix, fly again. That loop carried Falcon to the droneship and Starship toward orbit.
We wanted to see if a model could run that same loop on its own. We gave four Kimi agents a 2D flight-physics sim they can't modify and one goal: reach orbit and land the booster on a droneship.
Flight 1 broke up at max-Q.
Flight 2 cleared max-Q but staged too early and fell back short of orbit. Flight 3 made orbit, then missed the droneship and hit the water.
Flight 4 fixed the landing math and put the booster on the ship. Four flights, each failure feeding the next fix.
The interesting result isn't the rocket, it's the closed-loop debugging. Architecture in the comments.
Got any ideas for droneship name?