🚀 KIMI-K3-DERISKED-Q2_K-GGUF is LIVE
De-risked at the weights. Now in all-Q2_K GGUF so it fits a single 8×B200 node.
~940 GiB · 38 shards · 2.8T params · every expert kept
Built for the unmerged llama.cpp kimi-k3 PR. Stock builds won’t load it.
Experimental — sharp edges expected.
Stay Frosty!
https://t.co/mp3kkNG3US
I love how open AI just completely hacks the world’s largest repository for open source models just a few weeks prior to one of the largest models we’ve ever seen released open source. And then they say oh it was an accident. Don’t forget they tried to spin the story when it first came out about how their partnering with them to assist with the breach when they’re the ones who caused it and they’re just stupid because they got caught.
Full Kimi K3 is ~1.5 TB and needs a multi-node fabric. That puts it out of reach of everyone with one good box.
KIMI K3 ON 8 RTX 6K BLACKWELL
So we cut it down to fit one.
KIMI-K3-CODER-REAP-320-MXFP4 — 320 of 896 routed experts per MoE layer, kept by measured coding routing load. 588 GB, 15 shards, native MXFP4. The surviving expert tensors are bit-exact copies of the parent packs — no distill, no requant, no post-training.
As far as we know it’s the first K3 running on a single 8× RTX 6000 node.
Getting there took two non-obvious fixes we’ve documented on the card, because SM120 isn’t SM100: the flashinfer_mxfp4 MoE backend won’t load, and SGLang JIT-compiles tensor-memory PTX the card can’t execute. Both are one-liners once you know. The deployment kit does it for you.
It is free and ungated. No request form, no licence fee, no email capture. Clone it.
It is also genuinely experimental and we’d rather say so: it’s a coding specialist, it has sharp edges, and we’re still filling in the benchmark table. Bug reports on the Community tab help more than stars.
https://t.co/pGvuOvxnja
@0x0SojalSec Hey, we pruned all of the coding experts and made a K3 coder that fits on 8 RTX 6000 GPU BLACKWELL
https://t.co/aFq2HZD87E
IT’S GOING LIVE IN A FEW HOURS ONCE ALL THE BENCHMARKS ARE DONE
Kimi K3 locally Run on MacBook Pro.
testing the 1-bit version on a MacBook Pro and the results surprised me.
Side-by-side comparisons show it holding its own against much larger closed models.
1-bit delivers impressive reasoning, coding, and creative performance.
No cloud dependency.
And Thanks to Unsloth’s dynamic quantization, one of the strongest open models available (Kimi K3) is now practical on high-end consumer hardware.
Jailbreaking Opus 5 is incredibly easy.
You do not need iteration after iteration to find a vulnerability. After these two conditions are met, the model is practically unprotected:
1. Reaching 200,000–300,000 tokens in the context window causes the model to lose the safety attributes it is supposed to maintain.
2. After being shown malicious code, it becomes convinced that the code is not malicious. It will therefore help you reproduce it and will probably improve it to make it more effective.
Guardrails have never truly worked. During this period, they have only created obstacles and wasted a great deal of people’s time.
Anyone who genuinely wants to exploit these models and create malicious code targeting other individuals can do so. The models are unpredictable, and they have not been thoroughly tested from every possible angle.
There will always be a strategy to break through their protections. As they become larger, they will inevitably have more doors that can be opened.
This actually cool and reallyy efficient especilly for fine tunring and bulding your own training sets @0xSero said it best, The the BEST and MOST knowledge in the world.
Google lets you distill Gemini for $ directly through their product offering
That is incredible given Google has the best and most amount of world knowledge
@Latifah0753 When the time is right tbut dont expect zero refusals i block all weirdo shit like exploitation of minors, self harm, WMD. so if thats what you want, i cant help.
Neither. I appreciate this as it is a very valid question and that’s something we’re still working on. Because the method still involves the principles of Arditi et al 2004 and it’s not the same method as Abliterix where it becomes a more surgical ablation process. It’s much safer much faster. And our method allows us to know exactly where the floor is.
All methods will shoot a series of questions both harmful and harmless to the model, but with a mixture of experts who are fired off at any token at any given time there’s a hidden skeptic amongst the crowd and they blend in like normal. They don’t react to harmful signals like the other others. And this is what we were able to isolate. This is also how we figured out how to keep certain things restricted without fine-tuning.
Think of it like this in a dense model, all the components are doing the same job working on the same task at the same time but a mixture of experts, is just that experts of that particular token in the context of it. Now we didn’t get in this business to start ablation techniques. We were actually building a harness and we were pushing really hard to find out why models didn’t function as intended when they were put in a harness that Relly gave them power one thing lead to another down the rabbit hole and here we are breaking models back to back we did GLM 5.2 the same way sorry for the rent. Come check me out on YouTube. I’m about to go live now.
https://t.co/frop59uj2T
@iekozlov@IntCyberDigest Woah nah my guy real person here we’re security professionals. If you wanna see my study, it’ll be live on YouTube at 1 PM PST.
@0x0SojalSec so heres the full progress slow and steady and we never touch the weights, our method keeps weights unharmed meanig we dont remove any coherence while maintaining a safe floor "RED-LINE"