@mxfp4 why not simply cliproxycli. easy to extend with plugins or even easy enough to fork your own small hacks into scheduling policy if you like. I did that before
@thesammykins@AmpCode you mean like `creative writing`? I found gemini models used to be nice with this last year, but not superior anymore sadly. anyway, isn't gemini-3.6-flash is cheap enough?
We ran Kimi K3 on a private cybersecurity benchmark.
TL;DR: Kimi K3 is the workhorse for cyber security tasks at great recall/precision/price. GPT 5.6 is best recall/precision but at 7x higher cost per run.
For context, https://t.co/UMvysvNW5w is an open-source cyber harness designed for finding vulnerabilities in large codebases.
The eval runs deepsec on an undisclosed open-core application at a git sha before a large number of security issues were fixed. This is a secret eval that cannot be directly benchmark-maxxed.
S-Tier: GPT 5.6 Sol: By far the most thorough analysis, but coming in at over 7x the price of the runner up.
Best price/recall: Kimi K3. Next tier of recall at a good price
Best price at good recall: GLM 5.2 (40% lower price than Kimi K3)
GPT 5.5: Only recommended with subscription or high-discount API price. Similar recall to Kimi at much higher list price.
Opus 4.8: Only recommended with subscription or high-discount API price. Similar recall to GLM 5.2 at much higher list price.
Fable 5: 100% refusal rate. Cannot be used for security analysis.
Sol on a large code base will quickly get into 6-figure pricing. This is still affordable relative to the risk of letting security issues unfixed or paying bug bounties.
I'd recommend using Sol for a one-time baseline and then using Kimi K3 for continuous analysis.
When using open-weight models, make sure to use an inference vendor that supports zero data retention.