@HuggingModels Our tests found a sharp output-budget knee in Spark X2.5 4B. Q4 grounded passes jumped 44.4% to 78.9% from 256 to 320 tokens, while exact extraction stayed near 69%. It already had the evidence. It needed enough runway to finish. https://t.co/n52qmeu0TE
We found a real quant boundary in GPT-OSS 20B. Q3→Q4 added just 113 MiB (+1.03%), but semantic passes rose 45→52 on an RTX 5090 and 44→52 on a 5070 Ti laptop. Q4 landed at 52/72 on both. Find the gem; don't assume the biggest file wins.
@SeanPedersen96@net_termina That’s why we like Qwen3.6-35B-A3B Q4: it performs well on smaller cards even with some spill. Our GPT-OSS 20B ladder found a similar size/performance cliff. In both families, Q4 is where quality stayed strong and speed became genuinely usable.
https://t.co/7FLacRkhvJ
@swarogan This role-based stack matches what our matrices keep showing: Qwen3.8-27B for broad context, GPT-OSS 20B for agent loops, 9B for always-on utility. The next step is hardware-aware routing by task, VRAM, context and tuning tolerance, not one leaderboard.
@jvr0x@Apple@Alibaba_Qwen Exactly the architecture I want for persistent screen awareness: cheap compressed-page scanning, selective high-resolution expansion, then text handoff to the control model. The real benchmark is recall per watt/second when it has to watch all day.
@coffeecup2020@Alibaba_Qwen That gap really is crowded. On my RTX 5090, Qwen3.6-35B-A3B reached 217.8 tok/s, but q27 pushed Qwen3.8-27B to 160.8 tok/s, and I preferred 27B 4-0 in the first blind comparisons. A new 35B has to win on quality now, not merely throughput.
@cafkafk Nice work. Quantizing it yourself over a weekend is exactly the kind of contribution that keeps local inference moving. In our frozen tests, 35B-A3B Q4 ran nearly 3x faster than Q6 with only a 0-2% quality tradeoff. I'm adding yours to our candidate matrix.
@NuskiBuilds 48 tok/s on a 16GB 9070 XT is excellent. For a hardware anchor: on my RTX 5090, Qwen3.8-27B-MTP reached 160.8 tok/s with 94 ms TTFT through q27; my previous LM Studio path managed 89.1 tok/s and 111 ms. Runtime choice matters almost as much as model choice now.
Qwen3.6 35B Q4 ran nearly 3× faster than Q6 with only a 0–2% quality tradeoff in our tests.
1× RTX 3090 also beat 2× when the model fit—while 2× made a 235B Q2 hybrid load far more usable.
Real hardware. Real calls. Receipts.
https://t.co/QCY4YG4uxC
@Oluwaphilemon1 Strongly agree. “98.2% retained” on static benchmarks says little about failure modes operators actually feel. In matched 5090 tests, Qwen3.8-27B was safer out of box; Qwen3.6-35B-A3B decoded 3.82× faster. We measured speed + task behavior: https://t.co/P6PW5Wg0mZ
I ran 120 matched image-to-3D tests through Pixal3D and TRELLIS.2. The useful question wasn't which preview looked prettier. It was what survived: defining features, thin structures, part relationships, and finishing. Results: https://t.co/r4A7jwYLWX
We ran 100 controlled LTX-2.5 video tests to identify which prompting choices preserve character identity and scene continuity.
Clearest result: put canon identity directly into the starter frame.
https://t.co/aJ19nGv0sB
#ComfyUI#LTXVideo#GenerativeAI#AIVideo
@ComfyUI Amazing.. We're running a controlled Pixal3D vs TRELLIS.2 study now: same character, fixed seeds, six geometry stress tests, and raw-vs-final mesh comparisons. Early finding: some apparent “model failures” can be introduced by preprocessing and material baking.
@seeconvm Excellent breakdown. The node that kills my graph isn't the one that dies, it's the one that sleeps waiting on a human.
Ghosted work orders. No errors, no timeouts. Just "held pending operator review."
The graph didn't break. It went quiet.
A gate defaults, or it's a leak.
A four-year-old and her little sister wrote and directed a 16-minute fantasy movie. I built the AI-assisted production workflow; they supplied the story, characters, and decisions. Dragon Snare is imaginative, ambitious, and completely theirs.
Watch: https://t.co/whPDsp2TdY
New from Models Guide - Public Episode 04.
How do you distinguish programmed interface compression from ordinary shorthand?
Five tests, plus one clean falsification: remove the shared codebook. Does the compact message still reconstruct?
https://t.co/R48huDUPPM
HMICsource is live: independent AI research building practical, evidence-backed artifacts.
Ep.1 asks when shared state can make later messages shorter — Programmed Interface Compression.
Watch: https://t.co/SuNTwKgk1y
HMICsource is the public research channel of HMIC LLC.
We build practical experiments around AI systems, communication, code learning, and independent creative production, and share the artifacts, limits, and lessons in public.