A underwater peaceful vegetation setup, made in IlluGen.
The tiler's shapes are oriented to the UV direction and then stretched with an advect.
#IlluGen
@jukan05 The demand gap is still huge, but Hyperscalers have to limit their Capex ambitions because of cash flow, so for now we just need to patiently wait for profitability on the application side to improve.
@jukan05 This is old news from March. The launch of Gemini 3.5 Flash on May 19, 2026, significantly eased the compute capacity constraints. Meanwhile, the recent release of GLM 5.2 has been drawing even more user attention away.
🙏 Thanks to the @NVIDIAAI team for highlighting DFlash support on vLLM!
With DFlash speculative decoding, swapping EAGLE-3 for a DFlash checkpoint is a config-only change — no code edits needed.
It runs through the open-source Speculators library, which links the DFlash drafter to the target model's hidden states in the vLLM inference path.
On Gemma-4 31B on a single Blackwell Ultra GPU, this delivers up to 5.8x higher throughput at the same concurrency over autoregressive decoding:
🧮 Math500 — 5.8x
➕ GSM8K — 5.3x
💻 HumanEval — 5.6x
🐍 MBPP — 4.4x
Read the blog here! 👇
@NoamShazeer Excellent. Crisis averted—OpenAI’s poor performance was about to crash the market. Wall Street and the whole market will be thanking you. 😂
We're the first to make the full GLM-5.2 (FP8) run on RTX 4090s.
GLM-5.2 is the new 753B SOTA open-weights model, and it officially ships for datacenter GPUs only: H100, H200, B200. We ported its sparse-attention kernel stack to consumer hardware.
A frontier open model, off the scarce GPUs and onto the abundant kind.
https://t.co/08rCk3b3h5
@Zai_org