Synthetic data promised to shatter data scarcity barriers, but self-generated samples trigger catastrophic model collapse.
We discovered the key is thinking in reverse: degradation from self-training isn't random noise—it's a powerful signal provably anti-aligned with the real-data gradient.
Neon reverses this degradation to achieve SOTA image generation—FID 1.02 on ImageNet-256—with <1% additional compute.
Professor @randall_balestr discussing some exciting research he has been working on recently, in particular around spline-based interpretability.
• Why reconstruction learning can fail for perception
• How deep nets partition space like a crystal
• Spline theory new insights
• How you can jailbreak RLHF by increasing the intrinsic subspace of the prompt
Dropping on MLST shortly
Excited that our paper "Pedagogical Alignment of LLMs" is accepted at EMNLP'24 findings 🎉
Thanks to all authors - Kangqi Ni, @Sapana_007, @rbaraniuk
Read here: https://t.co/ovMkO87imj
Does training a generative model on its own synthetic data always result in MADness/model collapse?Turns out it doesn’t! We show that a diffusion model can “self-improve” using its own synthetic data while preventing MADness/model collapse altogether! Link to paper: https://t.co/6tKJQ9IEZm
Does recursive training of generative models only lead to a loss of diversity or is there more?
We show that 'self-consuming' generative models don't only face model collapse, they can emphasize hidden artifacts of the generative model as well!
https://t.co/bQFBmQlWfv
🧵(1/N)
Our paper "A Probabilistic Framework for Pruning Transformers via a Finite Admixture of Keys" has been recognized as one of the top 3% accepted papers at ICASSP this year. Congratulations to @LongMinh210, @MinhTam160895, and Hai Do for their great work!
https://t.co/8rLaZaFqpA
How are Deep Neural Networks black-boxes if you can visualize them in an 'exact' manner?
Our new #CVPR23 paper, presents a fast and scalable PyTorch toolbox to visualize the linear regions, aka partition+decision boundary, of any DNN (red🔻)!
https://t.co/yEJepElaBa
🧵 1/N
Fasten your seatbelts as our new paper takes you on a visual tour of current challenges in multimodal language models using stable diffusion models (SDMs). We find that SDMs struggle with function words like pronouns, preposition, conjunction, etc.
https://t.co/fWVXWrxSh3
Provable control of the quality and diversity of sampling for pre-trained deep generative networks ... without any additional learning! Check us out at #CVPR2022 tomorrow at 8:30 am in Hall B1, Oral Session 3.1.1, and Poster Session 3.1
1/6 WHO adopted a global strategy for cervical cancer elimination - includes treatment of 90% of women with precancer
Our new paper https://t.co/58MsGLvDSs uses deep learning to detect cervical precancer from high-resolution in vivo optical images @Rice_BIOE@rice_dsp
Richard Baraniuk has been elected to the National Academy of Engineering in recognition of his contributions to engineering "for the development and broad dissemination of open educational resources and for foundational contributions to compressive sensing."
Congratulations, CJ Barberan on successfully defending your dissertation “NeuroView: Explainable Deep Network Decision Making” and becoming Dr. Barberan!
MaGNET is a framework that allows uniform sampling from pre-trained deep generative network manifolds. Published at #ICLR2022
Paper Link: https://t.co/CBlezdIDAN
ICLR Video: https://t.co/tsBsKF9Z9m
Colab & Codes: https://t.co/Xb3zQUgbt0
@imtiazprio@randall_balestr@rbaraniuk
🎉 We invite you to join us on April 18 at 4:30 p.m. for a virtual celebration in honor of Richard Baraniuk's election to the National Academy of Engineering. Register via the link below!
🔗 https://t.co/vyebu5l1Jo
For those of you interested in understanding the over-parametrized models, don't miss to attend the TOPML workshop!
When 📅 - April 5-6, 10.30am-6.30pm EST
Where 📍 - Register at https://t.co/qtdGYPlhQG.
Web 🖥️ - https://t.co/Yr4TzAvUIY
#DeepLearning#MachineLearning@rice_dsp
Our recent #CVPR2022 paper investigates how neural net architecture impacts decision boundaries. We find that two runs with the same architecture yield very similar boundaries. Two different architectures yield radically different boundaries (yet similar test acc).
It is not what a deep net can approximate that matters, but how it learns to approximate. Matrix perturbation theory explains why the loss surface of a ResNet or DenseNet is less erratic and eccentric than a ConvNet and hence easier to optimize under SGD https://t.co/BypTblnPZt