Synthetic data promised to shatter data scarcity barriers, but self-generated samples trigger catastrophic model collapse.
We discovered the key is thinking in reverse: degradation from self-training isn't random noise—it's a powerful signal provably anti-aligned with the real-data gradient.
Neon reverses this degradation to achieve SOTA image generation—FID 1.02 on ImageNet-256—with <1% additional compute.
🦭 Come check out WaLRUS at #NeurIPS2024 today!
We introduce wavelets into the state-space model family—outperforming HiPPO-based methods on long-range representation tasks.
📍 Exhibit Hall C,D,E — Poster #2309 🕟 4:30–7:30 PM PT
(1/n) Until recently, even strong LLMs struggled with USAMO/IMO problems. This year, specific model variants from Google and OpenAI’s were reported to solve 5/6 IMO problems. In our recent work, we asked a relevant question: Can we grade proofs fairly with partial credit using LLMs?
Fair points! That's exactly why we developed Neon.
SIMS was diffusion-specific. Neon extends it to all model architectures.
More importantly, we theoretically prove that mode-seeking samplers (which nearly all modern models use) create anti-alignment between synthetic and real gradients. This guarantees improvement from synthetic data when used correctly.
You're right it needs validation on large-scale web pretraining - we focused on controlled experiments where we can measure precisely. But the theoretical mechanism is general.
Synthetic data promised to shatter data scarcity barriers, but self-generated samples trigger catastrophic model collapse.
We discovered the key is thinking in reverse: degradation from self-training isn't random noise—it's a powerful signal provably anti-aligned with the real-data gradient.
Neon reverses this degradation to achieve SOTA image generation—FID 1.02 on ImageNet-256—with <1% additional compute.
Neon shows synthetic data collapse can be harnessed, not avoided.
By reversing the degradation direction, we turn a known failure mode into practical improvement—no fresh real data, no auxiliary models, no complex training.
Just a simple parameter merge that makes models better.
📄 Paper: https://t.co/6KxgNPXM1P💻 Code: https://t.co/pSUXUSJebj
When does negative extrapolation fail?
Theory predicts Neon only works for mode-seeking samplers. With diversity-seeking samplers (temperature τ > 1, or those upweighting low-probability regions), gradient alignment flips: ⟨r_d, r_s⟩ > 0
In this rare case, you'd use w < 0 instead—interpolating toward the self-trained weights rather than extrapolating away.
We empirically verify: diversity-seeking (ζ=0.9, green) optimizes at w<0, while mode-seeking (ζ=1.1, orange) optimizes at w>0 (Neon).
Excited that our paper "Pedagogical Alignment of LLMs" is accepted at EMNLP'24 findings 🎉
Thanks to all authors - Kangqi Ni, @Sapana_007, @rbaraniuk
Read here: https://t.co/ovMkO87imj
@iliaishacked@jeremyphoward@simonw +1 to Ilia said. On another note in our recent paper we showed that model collapsed can be 100 prevented ( no increase in FID even if synthetic data was in the training set) https://t.co/5J6onAJMLE