Here is a direct comparison of over 3400 distinct findings across hundreds of PRs.
Mergestorm is cheaper, faster, and catches way more findings than Greptile.
This report was generated by Fable 5.
If you're looking for a code review agent that's cheap and fast and uses the latest frontier models, add the Mergestorm github app.
Visit: https://t.co/Ex9dHjf7pj
I have used SOL for planning a few times and it seemed solid, but haven’t experimented with it enough I’m way too comfortable and trusting with Opus.
I always use Cursor’s default for the model selection.
Fable High
Opus High
Don’t think I’ve ever used xhigh.
If your plans are well scoped you don’t really need the higher versions.
Please man if this is AI, tell it to stop using em dashes. If you’re a human you should also stop using em dashes even if you feel like it, a sad truth but it will help your visibility whether you like it or not.
For human eyes, yes it’s important to get human input into your coding pipeline, the question is where?
We can automate code review + patching on a branch / PR level.
You have some options:
You can use humans to prompt agents to open issues + PRs. Then trust your pipeline enough to let it automatically review the code + patch it, using specialized agents and harnesses (which we’ve built if you’re interested)
You also need humans to sanity check your code at the merge gate, whether you’re using Graphite for stacked PRs or just normal gh.
Humans are critical to prevent drift and AI slop, but their job is now much higher level, even higher than architecture.
Cursor Grok 4.5 is hands down y win henchman when it comes to writing code, opening PRs, applying changes, etc.
Opus / Fable is still king for planning, but Claude + Grok are an unbeatable combo.
Hey everyone 👋
It’s been a minute.
For those who’ve been following since the early crypto days, thank you.
This account started as the voice of The Merkle, and it grew with all of you through the wild Bitcoin/altcoin era.
A lot has changed since then.
The Merkle @themerklehash and NullTX @nulltxnews have new homes and fresh directions. This account is now my personal founder account (@marginsystems).
I’m still the same guy who built those projects, but today I’m all-in on AI, autonomous coding systems, and building the future at https://t.co/D0IrJ0yrRX (@Mergestorm)
If you’re into:
• AI agents & coding tools
• Deep tech that actually ships
• Honest takes from someone who’s been in the trenches
…then stick around.
I’ll be posting more transparently here about what we’re building, lessons learned, and the crazy stuff happening in AI right now.
Old followers welcome. New ones even more so.
Let’s build the next chapter.
Best,
-Mark
I love how AI talks to me now:
"Ship P1-16 as a 2-PR Graphite stack: (1) shared llm_run_usage writer in review-harness, (2) shared triage HTTP helpers in llm-triage-chain. Tip PR Closes #57; API chat streaming stays a linked follow-up so #57 can close without boiling the ocean."
What is boiling the ocean?
The amount of idioms AI is starting to use is really lovely. I'm learning 2-3 new ones every day.
Dogfooding
table-stakes
snowshoe
boiling the ocean
Fable especially loves using them.
It’s not as easy or simple to switch models and test them for real world use cases.
Take for example our AI code review pipeline. We got it tuned just perfect with Claude + DeepSeek.
There has to be a clear reason why we would want to expend resources, time, and money to try out Gemini. Unfortunately I don’t see a clear reason at the moment, but that might change.
Would love to get a trial maybe to test it out
Please why do you call your model Flash….
DeepSeek has V4 Flash…
Why name it the same… just confusing for no reason and seems disrespectful to DeepSeek.
We’re rolling out three new models to make AI agents faster, smarter, and cheaper at scale:
🔵 Gemini 3.6 Flash: It uses fewer tokens than 3.5 Flash to deliver higher quality work at the exact same cost.
🔵 Gemini 3.5 Flash-Lite: A fast, cost-effective option for everyday tasks like processing documents and agentic search.
🔵 Gemini 3.5 Flash Cyber: A cybersecurity model built to find and patch critical software vulnerabilities.
@OfficialLoganK@MarcosHernanz Is there any real reason to use Gemini 3.6 Flash vs DeepSeek V4 Flash? Nothing even comes close to the cost / intelligence that DeepSeek is offering.
@growing_daniel None of our customers have asked for OpenAI models. Everyone is still look ng for Claude models to do code reviews + Kimi K3 has been asked about.
Don’t get me wrong OpenAI’s SOL is smart, but it’s hard to justify using it over Claude’s fleet of models, besides cost maybe.