Sungho Lee, Marco Martínez-Ramírez, Junghyun Koo, Wei-Hsiang Liao, Kyogu Lee, Yuki Mitsufuji, "Exploring the Design Space of Representation Learning for Audio Transformations" https://t.co/hIHpDQObxc
SCRAPL: Scattering Transform with Random Paths for Machine Learning
I’m excited to share our #ICLR2026 paper on SCRAPL: an algorithm that makes wavelet scattering transforms usable as differentiable loss functions!
paper: https://t.co/D0X4v80P9J
web: https://t.co/l6hc0zos5e
I just released PhilTorch v0.1!
https://t.co/KYgGx1OiXV
Features:
- Fast, differentiable, parameter-varying, scipy-like filters.
- Extremely fast biquads on GPU.
- State-space models.
- Forward-mode AD.
- Comb filters.
- More to come.
It's on PyPI. Try it. DDSP 4ever.
@randall_balestr Thanks for the brilliant work. I have a small question: have you tried SIGReg + I-JEPA? My thought was that even if EMA/detach heuristics become unnecessary with SIGReg, masked latent prediction could still have some advantages over the DINO-like invariance learning.
Sungho Lee, Marco Mart\'inez-Ram\'irez, Wei-Hsiang Liao, Stefan Uhlich, Giorgio Fabbro, Kyogu Lee, Yuki Mitsufuji, "Reverse Engineering of Music Mixing Graphs with Differentiable Processors and Iterative Pruning," https://t.co/0bMuAWCtNX
``Score Distillation Sampling for Audio: Source Separation, Synthesis, and Beyond,'' Jessie Richter-Powell, Antonio Torralba, Jonathan Lorraine, https://t.co/P5iuWrYMeL
``Improving Inference-Time Optimisation for Vocal Effects Style Transfer with a Gaussian Prior,'' Chin-Yun Yu, Marco A. Mart\'inez-Ram\'irez, Junghyun Koo, Wei-Hsiang Liao, Yuki Mitsufuji, Gy\"orgy Fazekas, https://t.co/SYj37minFg
I'm pleased to share my internship project @SonyAI_global, DiffVox, a differentiable effects chain for vocals with hundreds of presets.
arXiv: https://t.co/7M0o4ma81C
code: https://t.co/3vSj4Pv2un
🤗: https://t.co/ySu0i77oAm
``DiffVox: A Differentiable Model for Capturing and Analysing Professional Effects Distributions,'' Chin-Yun Yu, Marco A. Mart\'inez-Ram\'irez, Junghyun Koo, Ben Hayes, Wei-Hsiang Liao, Gy\"orgy Fazekas, Yuki Mitsufuji, https://t.co/R3v6tpbG07
``Towards Generalizability to Tone and Content Variations in the Transcription of Amplifier Rendered Electric Guitar Audio,'' Yu-Hua Chen, Yuan-Chiao Cheng, Yen-Tung Yeh, Jui-Te Wu, Jyh-Shing Roger Jang, Yi-Hsuan Yang, https://t.co/qRzoWbsm4U
1.5 yrs ago, we set out to answer a seemingly simple question: what are we *actually* getting out of RL in fine-tuning? I'm thrilled to share a pearl we found on the deepest dive of my PhD: the value of RL in RLHF seems to come from *generation-verification gaps*. Get ready to🤿!
``Differentiable Black-box and Gray-box Modeling of Nonlinear Audio Effects,'' Marco Comunit\`a, Christian J. Steinmetz, Joshua D. Reiss, https://t.co/UAYshlxS79
``Differentiable Black-box and Gray-box Modeling of Nonlinear Audio Effects,'' Marco Comunit\`a, Christian J. Steinmetz, Joshua D. Reiss, https://t.co/UAYshlxS79
``Retrieval-Augmented Neural Field for HRTF Upsampling and Personalization,'' Yoshiki Masuyama, Gordon Wichern, Fran\c{c}ois G. Germain, Christopher Ick, Jonathan Le Roux, https://t.co/00bh8DdeYz
Recurring question: So what? It can't be used to train Deep NNs, let alone LLMs.
That's true. ES doesn't work well in high dimensional spaces. As it maintains a population, memory requirements (if you are using ES to train Deep NNs) are prohibitive and sample efficiency is not competitive with Adam. Moreover, Choromanska et al. (2015). The Loss Surfaces of Multilayer Networks. arXiv. https://t.co/QJvM5GTiVh shows that as dimensionality increases, most local minima are "good" in the sense that they generalize well and are close in value to the global minimum. So most of the time we don't care about finding the global optimum.
That said, many successful recent in-context learning papers can be understood under the lens of ES, for example, Wang et al. (2023). Voyager: An Open-Ended Embodied Agent with Large Language Models. https://t.co/1499CMYhQX, FunSearch by Romera-Paredes et al. (2024). Mathematical discoveries from program search with large language models. Nature, 625(7995), 468–475. https://t.co/E0pvrs1LgI, Fernando et al. (2024). Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution. ICML. https://t.co/PUqg6uwDQy), Samvelyan et al. (2024). Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts. NeurIPS. https://t.co/bA3HDiJiaS), Zhang et al. (2023). OMNI: Open-endedness via Models of human Notions of Interestingness. ICML. https://t.co/i5kReptzsn, Faldor et al. (2024). OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code. arXiv. https://t.co/u7dSC8J9eS, Lu et al. (2024). The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery. arXiv. https://t.co/DXfAdbRgTF). All of these works use foundation models to either vary or select artifacts (or both). All of them implement an evolutionary strategy in some way, but not on the NN parameter level. Most of them keep an archive of novel and interesting stepping stones that become the basis for more successful solutions later.