🚀 New paper: Balancing Frequencies and Pixels in Flow Matching
We tackle the low-frequency bias in pixel-space flow matching and train JiT up to 40% faster without any architectural changes.
📄 Read it here: https://t.co/VmNzR4wRzd
Google releases TIPSv2: four sizes, each with a general image-text encoder and a DPT version for dense vision tasks.📜 Apache 2.0.
🤖 https://t.co/urJ6CXwt8S
📄 https://t.co/K4SHatz8sd
🏆 SOTA on all four reported zero-shot segmentation benchmarks, with top-two results on 5/7 global image-text and 7/9 image-only evaluations.
⚡ Beats DINOv3 on 4/6 shared tasks at the ViT-L scale, despite its teacher using 6× more parameters and 15× more images.
🧠 iBOT++ lifts ADE150 zero-shot segmentation by 14.1 mIoU; Head-only EMA cuts training parameters by 42%.
Evoke
A 14B, 3-step CFG-free world model that generates endless, steerable video worlds — with persistent memory, camera control, and mid-rollout prompt changes.
ByteDance released SwanTale on Hugging Face
A unified model for multi-speaker speech and audio generation across zero-shot and instruct tasks, enabling expressive voice cloning, natural language style control, and acoustic scene synthesis.