FOUND ANOTHER FREE LOCAL MODEL
UNCENSORED / ADULT-CAPABLE
FineP*rn V5 is a Krea 2 checkpoint for ComfyUI
• ~12GB INT8
• ~4.7GB text encoder
• no paywall
• realistic outputs
How to use:
1. download the model
2. add it to ComfyUI
3. load the workflow
4. generate
Model:
https://t.co/g4epy7IRgV
Fictional adults only. No real people or minors
FOUND AN UNCENSORED LOCAL VIDEO MODEL TO CREATE ADULT CONTENT 👀
NSFW Wan 1.3B was fine-tuned specifically for adult content
• text to video
• based on Wan 1.3B
• trained on NSFW data
• around 2.8GB for the newer checkpoint
• can run locally
don't download the old e10 checkpoint though
the repo now recommends: wan_1.3B_exp_e14.safetensors
check it out:
https://t.co/kahdzH9zBG
test it locally with fictional adult characters
FOUND A LOCAL AI BODY SWAP MODEL
BFS can swap:
• body swap
• head swap
• face swap
• works with Krea 2
• Qwen Image 2.1
• Flux 2 Klein
• ComfyUI workflows included
the body swap version is already available
link: https://t.co/Wk7NPG4vBY
As promised, instead of for $15,000 GPUs, we max optimized a model that can fit comfortably in any GPU with 8GB VRAM running Windows
200 tok/s, *128k context, text/image/audio.
Anyone who wants to join the local AI community with 5 year old laptop can run this with recipe below, and this test was done on 3070ti laptop I forgot I had.
We pulled every instruct model released since May that fits 8 or 12 GB, and ran the six that mattered through the same 20 tasks: JSON extraction, code, a table, arithmetic, a fake-treaty honesty check, tool calls, language, a persona, a fact planted deep in a long prompt, and three photos. Gemma 4 E4B, Gemma 4 12B, Qwen3.5-9B, Qwen3.5-4B, Ornith-1.5-9B, the Qwen3.8-9B distill.
E4B is not the top scorer. Qwen3.5-9B and Ornith beat it on long context and on reading fine print in a photo. It won the pick for the things a newcomer on an old card actually needs: Apache 2.0, text + image + audio in, 3 GB of headroom on an 8 GB card, Weights Google trained to be 4-bit (QAT), and Google's own speculative drafter shipped next to them. Nothing else in the tier has all five, and the drafter alone is worth 2–3×.
Qwen3.8-27B is running locally on an RTX 4060 with just 8GB of VRAM.
And it’s not some tiny quant running with a crippled context window.
The setup is:
• RTX 4060, 8GB VRAM
• Qwen3.8-27B
• Unsloth IQ4_XS quant
• 14.6GB model on disk
• 64K context
• Native MTP
• ~150 tok/s prefill
• ~5 tok/s decode
• Only 25 layers offloaded
• No VRAM overflow
The model is simply split between GPU and system memory to stay inside that 8GB VRAM limit.
And that’s the part I find fascinating.
A 27B model that you’d normally associate with much beefier hardware is now usable on a mainstream $300-class GPU.
It’s not fast by desktop-GPU standards. Around 5 tok/s for generation is obviously not going to win any speed contest.
But that’s not really the point.
The fact that you can load a model of this size, keep a 64K context window, use native speculative decoding, and actually run meaningful coding workloads on 8GB of VRAM is wild.
Then there’s the benchmark result.
Qwen3.8-27B reportedly scores 61.7 on SWE-Bench Pro, compared with 53.4 for Claude Opus 4.6 Max on the same benchmark.