We are excited to share that SlideFormer (An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU), publicly posted on arXiv on March 17, 2026, has been accepted to DAC 2026.
More importantly, we believe that single-GPU full-parameter fine-tuning of 100B+ models should not be viewed merely as a headline number. It is fundamentally a systems problem — and an extreme but meaningful benchmark for testing the roofline and design limits of a training framework under heterogeneous memory constraints.
SlideFormer combines:
a lightweight asynchronous layer-sliding engine,
efficient heterogeneous memory management,
integrated advanced I/O,
and optimized Triton kernels.
In our evaluation, SlideFormer enables fine-tuning of 123B+ models on a single RTX 4090, supports up to 8× larger batch sizes and 6× larger model sizes, improves throughput by 1.40×–6.27× over baselines, roughly halves CPU/GPU memory usage, and sustains >95% peak performance on both NVIDIA and AMD GPUs. On a high-end PC with 256 GB host memory, models up to 24B can be fine-tuned at >95% peak performance.
We see this work not as chasing a single eye-catching number, but as pushing single-GPU full-parameter fine-tuning closer to its practical system limit — and helping make advanced LLM adaptation more accessible to individual researchers and small labs.
Paper: 📎https://t.co/LgUoZzP6Nx
Code release: planned for May 2026
Beyond the paper, we compare it with additional offloading baselines and evaluate long-context fine-tuning performance.
SlideFormer still performs well.
We are excited to release SlideFormer.
Can a single commodity GPU perform full-parameter fine-tuning of 100B+ LLMs?
SlideFormer explores this issue as a problem of heterogeneous co-design.
Paper: https://t.co/OIRtqIlDOe (DAC '26)
Code: https://t.co/wTgCcE7KI6
We are excited to share that SlideFormer (An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU), publicly posted on arXiv on March 17, 2026, has been accepted to DAC 2026.
More importantly, we believe that single-GPU full-parameter fine-tuning of 100B+ models should not be viewed merely as a headline number. It is fundamentally a systems problem — and an extreme but meaningful benchmark for testing the roofline and design limits of a training framework under heterogeneous memory constraints.
SlideFormer combines:
a lightweight asynchronous layer-sliding engine,
efficient heterogeneous memory management,
integrated advanced I/O,
and optimized Triton kernels.
In our evaluation, SlideFormer enables fine-tuning of 123B+ models on a single RTX 4090, supports up to 8× larger batch sizes and 6× larger model sizes, improves throughput by 1.40×–6.27× over baselines, roughly halves CPU/GPU memory usage, and sustains >95% peak performance on both NVIDIA and AMD GPUs. On a high-end PC with 256 GB host memory, models up to 24B can be fine-tuned at >95% peak performance.
We see this work not as chasing a single eye-catching number, but as pushing single-GPU full-parameter fine-tuning closer to its practical system limit — and helping make advanced LLM adaptation more accessible to individual researchers and small labs.
Paper: 📎https://t.co/LgUoZzP6Nx
Code release: planned for May 2026
In practical terms, for GPUs such as the **RTX 4090 / RTX 5090 / RTX Pro 6000**, the more realistic sweet spot for efficient fine-tuning is often closer to the **3B–14B** range, where turnaround time is much more suitable for real-world iteration.
We found one last signed jersey from the 2025 squad and we’re giving it away 🎁
The last @Twistzz signed 25th Anniversary we’ll give away on socials — RT & Reply for a chance to win! Good luck 🙏