ComfyUI is the most flexible, composable, and powerful open-source media generation tool with a massive ecosystem of workflows and custom nodes.
Your Hermes Agent can now install, launch, manage, and run sophisticated @ComfyUI workflows on demand.
someone is trying to improve VAE decoding in ComfyUI. a hellish mix of LLMs attempting to make the VAE more reliable...
- 720p-4K video on 8–12GB VRAM; dynamic batching, no OOM;
- hdd offloading; works with LTX/Hunyuan.
Who wants to test?
https://t.co/qZIeIMg7Gy
Recraft V4 & V4 Pro are now available in ComfyUI.
Tuned with designers for composition, lighting, color, and material realism.
Generate SVGs from text prompts with clean, editable paths. Production-ready vectors, no tracing.
1/2 Qwen3.5 is here. The next frontier of Native Multimodal Agents is open. 🚀
We are thrilled to release Qwen3.5-397B-A17B, our flagship open-weight vision-language model. Built for the future of coding, reasoning, and seamless multimodal interaction.
Key Highlights:
Inference Efficiency: A massive 397B total parameters, but only 17B active—delivering flagship power at a fraction of the cost.
Hybrid Architecture: Innovative Gated Delta Networks (Linear Attention) + Sparse MoE for extreme speed.
True Multimodality: Exceptional performance across GUI interaction, video comprehension, and agentic workflows.
Global Scale: Qwen3.5 now supports over 200 languages.
Empowering developers and enterprises to build smarter, faster, and more versatile AI agents.
Just shipped FastFlux2 Realtime Editor.
A fully open-source real-time editing studio in your browser.
Webcam → FLUX.2-klein-4B → Single 4090 @ 5 FPS, H100 @ 10+ FPS.
Repo: https://t.co/yzWvCddtg9
New open source LoRA Multi-Angles I created with @fal for Flux.2 🔥
Flux2.-Lora-Multi-Angles: give a camera angle in degrees → get your image from that view.
To build this, I created a system that captures 72 positions from Gaussian Splatting to generate the training dataset fully automatically in a webgl viewer.
▶️ Test: https://t.co/6lWL2TJy5r
📥 Weights: https://t.co/YUV6CrA2yv
Should I open source my WebGL viewer so you can train your own from gaussian splatting ?
A new image model from another team at Alibaba has been released.
7B mmdit+2B textencoder OVIS 2.5
OVIS-IMAGE, specifically designed for poster and text layout design.
OpenAudio S1-mini 🔊 a new OPEN multilingual TTS model trained on 2M+ hours of data, by @FishAudio
https://t.co/Yl76uDmU2f
✨ Supports 14 languages
✨ 50+ emotions & tones
✨ RLHF-optimized
✨ Special effects: laughing, crying, shouting, etc.
made a node that bounces the precision down to fp16 on upscale models
The result is almost twice as fast, with virtually 0 fidelity loss in image quality.
https://t.co/iE05wwSH7J
#comfyUI