(5/5) We have one workshop paper on GenAI in Finance https://t.co/OFd9EEhD6z which uses LLMs for genetical programming for effective and efficient alpha mining. The assets selected by our alpha factors have better long-term returns than traditional ML in US/HK/CN market.
(4/5) Our work in NeurIPS about data selection for LLM fine-tuning https://t.co/81VT9czQ1Q designs how to select data based on robustness for efficient instruction tuning. On several datasets, we can compress the training set to its 30% without performance loss after finetuning.
(3/5) Our work in NeurIPS about fast sparse adversarial training https://t.co/axOIf6ZAl5 studies the loss landscape of fast AT against sparse perturbations. We propose smoothing techniques to improve the algorithm performance and stability.
(2/5) Our work in NeurIPS Dualoptim https://t.co/9PyHS77UzS greatly improves the stability of machine unlearning with a simple trick: separate optimizers to the forget and the retain. It is effective for popular MU methods and models: classification/diffusion models/LLMs.
Our work about mitigating architecture overfitting in dataset distillation is updated on https://t.co/cVfIsc2NxT and will be published in IEEE TNNLS. It discusses different techniques in architecture, algorithm and optimizer to solve this problem by comprehensive case studies.
Our JMLR paper is online (https://t.co/jKLIrpd7RI) after almost three years.. it is about theoretical analyses of overfitting issue in adversarial training. (I finally found my account passwords to tweet..
Change the descent geometry and got rid of coordinate descent can greatly mitigate this issue and make more efficient training possible. Cmp with existing methods, our simple algorithm has no more hyper params, better efficiency, almost no memory overhead and better performance.
At #ICML, we will present a work about how to efficiently and stably obtain robust deep neural net for l1 bounded perturbations, a different story from l2 and linf cases. paper: https://t.co/7ICABcUtTe. Code: https://t.co/XdN9Ac75hm. Look forward to the events in Hawaii
key take away: coordinate descent used in existing methods for l1 robustness tend to generate sparse perturbations and leads to more frequent catastrophic overfitting in l1 case, even when training against multi step attacks.
I'll give a talk about our ICML'22 paper https://t.co/F63P1KkdCw tmrw (3pm CET) at the ELLIS Mathematics of Deep Learning reading group.
The zoom link is available here: https://t.co/X2tjSJwm04. Feel free to drop by if you want to chat a bit about sharpness and related topics :)
Excited to share our paper in @Nature: We revisit the 50+ year-old maths problem with AI: how efficiently can we multiply two matrices? Surprisingly, the answer is still not known - even for 3x3 matrices! With AI, we discover many new efficient and exact algorithms. 1/13