Run Codex 24/7 in the cloud with Atomic Bot!
Just hand it a repo and close the laptop. One click to start.
We ran it on ours first and it found a bug nobody had answered. Codex fixed it before we got back.
Qwen 3.8 27B on Atomic Agent was 2x faster than on Hermes and Prime agents!
Outputs:
Atomic: 5,535 lines, 19 modules, 2h35min
Hermes: 10,424 lines, 109 modules, 4h15min
Prime: 12,226 lines, 129 modules, 4h42min
We gave three agents the same tasks, rebuild these games:
✦ Hole Io
✦ Tower Bloxx
✦ Subway Surfers
Results:
✦ Atomic Agent gave us a real city going down the hole, skyscrapers tipping in, cars on the crossings, shops eaten one by one. In the tower every block swings in on its cable and lands where you drop it. Its runner is the cleanest of the three, though the world behind the boy stands still while he runs.
✦ Hermes dropped its hole onto a bare brown field with a few crates, and the hole grew while nothing ever fell in. Its tower buildings hang in the air with no cable at all, the logs reporting 12/12 tests passing on a game it had just broken. Its runner is an empty corridor with almost nothing in it.
✦ Prime built the prettiest town of the three, and five seconds in the whole thing tore off the ground into a tornado spinning around the hole, cars and houses circling forever and never falling in. Its tower lags on every drop, colours flickering on the buildings, and the lag is what brings the whole thing down by the fifth floor. In its runner the coins move at exactly the boy's speed, so the gap never closes and you can never pick one up.
Run Qwen 3.8 27B in Atomic Agent!
Atomic Mail put OpenClaw and Hermes head-to-head
We gave OpenClaw and Hermes one inbox each and the same prompt: prove you’re better. Who won?
Get your agent an inbox in 30 seconds:
– Plug in with one prompt via MCP or Agent Skill
– No CAPTCHA, no signup
– Run any workflow over email
Try the same prompt from the replies and share what happens.
We've tested the Qwen 3.8 Max preview and it absolutely rocks!
Each scene is built in one shot, a single Three.js file with nothing else added:
• Odysseus' war galley riding the waves
• Trojan horse
• Greek war helmet
Qwen managed to execute the task faster than the other models, while output quality is relatively the same across all of them. If the pricing stays close to Alibaba's previous flagships, we'll get one of the best speed/quality ratios on the market — perfect for running agents at scale.
Soon live on AI/ML API.
First local agent to beat Hermes on benchmarks!
✦ runs Qwen, Gemma, Llama via llama.cpp
✦ stable-prefix caching keeps sessions cheap
✦ TurboQuant cuts the KV-cache 6.4× smaller
✦ 37 tasks solved vs Hermes' 31 on GAIA Level 1
Open source on macOS, Windows & Linux 👇
Yup, a football video. The World Cup made us do it
Luma rebuilt image generation from scratch — reasoning first, pixels second. And it beats Google's Nano Banana 2 and GPT Image 1.5 on reasoning benchmarks
All 3 new models are now live on AI/ML API
luma/uni-1 plans before it draws. The model generates autoregressively: it works out layout, composition and text placement first, then renders the pixels. $0.052/image
luma/uni-1-max — same prompts, same params, max fidelity. 2K output + editing with up to 9 reference images. Built for hero shots and ad creative. $0.13/image
luma/ray-3-2 — up to 16 keyframes per clip, 20s, 1080p, native HDR + 16-bit EXR export. The video in this post came straight out of it
model ids "luma/uni-1" "luma/uni-1-max" "luma/ray-3-2"
@LumaLabsAI cooked. We serve
Kimi K3 just outperformed Claude Fable 5 at a quarter of the price.
Kimi K3 $3.11
Claude Fable 5 $12.23
Same prompt: a one-shot MECCHA CHAMELEON, Steam hide-and-seek game. A white chameleon hides in a hand-drawn room, paints itself with the mouse to match the wall behind it, then survives three sweeps of a robot seeker.
We asked for a live pixel-diff match %, five procedural zones, synthesized sound, three scored rounds — one HTML file, no libraries.
Both models live on AI/ML API
Grok 4.5 crushed OpenAI & Meta in this test!
Cost per run:
GPT Sol $1.63
Grok 4.5 $2.47
Meta Muse Spark 1.1 $1.08
The three prompts:
• Fruit Ninja – style slicing game
• Angry Birds – style fort collapse
• Crossy Road clone
Grok nailed all three: clean physics, smooth playback, good visuals. GPT Sol did fine until Crossy Road, where it froze. Meta Muse Spark was cheapest and it showed: its Crossy Road lagged badly.
The question isn't which model is cheapest, it's what a broken output costs you in reruns, wasted time, and things you can't ship. Cheap stops being cheap when it doesn't work.
GPT 5.6 Sol & Terra just made Fable 5's pricing look like a joke!
We gave OpenAI's top models (Sol, Terra) and Anthropic's top models (Fable 5, Opus 4.8) the same 3 prompts:
• Supernova boom
• Meteor hitting a city
• Solar system model
The bill:
Sol: $4.77
Terra: $1.24
Fable 5: $9.94
Opus 4.8: $2.46
Outputs came out surprisingly close. The prices didn't. Fable 5 cost unreasonably more than everything else, with Sol not far behind. Considering latest Fable 5 nerf, it's hard to see what you're paying for.
Fable 5 is BACK on AI/ML API
We gave Sonnet 5 and Fable 5 the exact same prompt and same real flight dataset of 70 airports and 435 routes pulled from flightradar. Then asked each one to turn it into a cinematic 3D globe as a single HTML file.
Outputs:
Sonnet 5: 9.8k tokens, $0.10
Fable 5: 15k tokens, $0.77
Sonnet 5 came in 87% cheaper. It drew the routes, but the planet underneath them barely exists: a dark wireframe with the arcs floating in space. Fable 5 built an actual Earth: textured oceans, ice caps, atmospheric glow. Mythos-class models are truly a masterpiece.
Docs, guides & setup below:
Gemini Omni Flash Preview is now live on AI/ML API!
Gemini Omni vs Seedance 2.0. We gave both models a GTA 6 night chase and asked for photorealistic, anime and LEGO versions.
Looks like Seedance 2.0. is still unmatched.
Both ran on one AI/ML API key.
Docs, guides & setup below:
We gave 4 AI models a secret word each and told them to steal the others.
GPT, Claude, Grok and DeepSeek bluffed and interrogated each other while an AI judge scored the smartest player. GPT won by barely speaking: it planted one idea early and let the rest expose themselves.
All four ran on one AI/ML API key. Code and full transcript below:
API-first email built for AI agents
One prompt to plug in via MCP or Agent Skill
Your agent gets its own inbox – and can run any workflow over email
Free in open alpha - link in comments