Final-year AI PhD. Student at @Westlake_Uni and @ZJU_China || Advisor: Stan Z. Li || Research interest: AIGC, UMM, Network architecture, optimization, and AI4S.
#AAAI2026 Oral 16:30-17:30 at Tourmaline 207: MergeDNA. Join us at Singapore Expo with research on DNA foundation models. Feel free to find us today (01/23, 12:00-14:00) at Hall 2~
- MergeDNA (https://t.co/0Sd4cTvGK9) Board #248
- TrinityDNA (https://t.co/l8PbVVfJdi) Board #1209
Introducing DINOv3: a state-of-the-art computer vision model trained with self-supervised learning (SSL) that produces powerful, high-resolution image features. For the first time, a single frozen vision backbone outperforms specialized solutions on multiple long-standing dense prediction tasks.
Learn more about DINOv3 here: https://t.co/lQpKhJLTZQ
Rep-MTL is an ICCV 2025 Highlight!
A new framework for Multi-Task Learning that uses representation-level task saliency to share knowledge without changing optimizers or architectures.
Our LLM optimizer wrapper, SGG (scaling with gradient grouping)💡, is published in #ACL2025, yielding consistantly improvement to LLM/MLLMs.
Welcome to view Virtual Presentations 1 on July 28 at 11:00 AM (UTC+2).
📄Paper: https://t.co/J6Dvkpla39
⚙️Project: https://t.co/z5LpOn3XmG
We got three papers accepted by #CVPR2025🥳View our works on Generation (MergeVQ), MLLM benchmark (https://t.co/mOnacFHhZF), and 3D Editing🚀 Thanks my co-authors for their efforts, @ZLYChangYu. @ZedongWangAI, @JuanxiTian, @chengtan9907.
View our #CVPR2025 paper MergeVQ🚀that resolves the trade-off between image generation and representation learning for visual tokenizers and AR generator.
📄 Paper: https://t.co/PYWvE0HWYQ
🌐 Project: https://t.co/rr0q95iK7o
🤗 HF (Daily paper top1): https://t.co/hKha8oH9V1
@3scorciav @ZedongWangAI I agree the professionalism of ACs. But it's a huge burden for ACs and reviewers to handle massive submissions nowadays. If the responsibility of the final decision is entirely placed on ACs, their workload would be overwhelming to provide reliable decisions in limited time.😅
Big shoutout to #CVPR2025 PCs & ACs for their dedication and incredible work!🫡 I would like to raise a concern about the "borderline" option in final ratings:
BA (borderline accept) / BR (borderline reject) share the same score (3) in #CVPR2025 but may carry quite opposite intent (accept vs. reject). IMHO, the "borderline" option for final ratings creates a procedural middle ground that allows reviewers to (relatively) disengage from for the authors' rebuttal and their post-rebuttal decisions.
As both an author and a reviewer, I suggest removing the "borderline" (3) option from final ratings in future flagship conferences. I believe this would help better clarify accept/reject stances, encourage reviewers to stay more engaged, and thereby assist ACs in making the final recommendations. @CVPR@ICCVConference@eccvconf@iclr_conf
@ZedongWangAI @jbhuang0604 We found that AdamW variants are likely to achieve high performances with great generalizability and no BOCB. But they still have limitations with pre-training and fine-tuning paradigm, e.g., LAMB pre-trained models requires the similar optimizers (e.g., AdamW) rather than SGD.
@JuanxiTian@jbhuang0604 Decades battle of deep learning foundational techniques that you should know🤩 Welcome to check our investigation of the coupling bias of vision networks and optimizers!
@JuanxiTian Check our new free lunch Switch EMA for DL optimization through a simple modification of EMA🤩
- Switching the EMA model to the online model
- Performance gains and speeding up without costs
- Pluggable to any DL optimizers
- Generalizable to various DL tasks
🌟[1/3] We propose the Switch Exponential Moving Average (SEMA) method, and through visualization of the loss landscape and decision boundary experiments, we demonstrate its effectiveness in improving model performance across various scenarios.
https://t.co/hULQPOvDJW
Through comprehensive benchmarks of 20 typical vision backbones vs. 20 popular optimizers, we explore the BOCB phenomena with four case studies and reach out empirical conclusions of network designs and the choice of optimizers. Check our online project (https://t.co/bfNpGHbRGf).
Why we should use a certain optimizer (like AdamW) to train Transformers or ConvNeXt, rather than classical optimizer like SGD? View our new benchmark and empirical analysis of the coupling bias of vision backbones and optimizers! 🤔
Checking details of our four papers
(1) CHALA for long sequence, https://t.co/mxudKA8TSh
(2) VQDNA for genome (first-authored), https://t.co/o5YSSuuEd5
(3) MMM for PPI, https://t.co/LGSc3AzGkn
(4) ReDock for molecular
docking, https://t.co/AZJ3d8NcTM
Happy to attend #ICML2024 to share four papers in #Vienna, which is my second trip to Vienna. Feel free to discuss with us at Hall C 4-9 today, especially my first-authored AI4Gene work VQDNA at #1312.
View our #ICLR2024 poster SemiReward😄, which conducts semi-supervised learning (SSL) like LLMs with rewarding techniques, making it a pluggable, efficient, and general SSL method.
SemiReward: 05.09, 16:30 p.m. - 18:30 p.m., Hall B
Poster: https://t.co/HUcOHMKvvg