[9/9] Importantly, because AC-GPT can be implemented without any architectural modification, it can be directly applied to fine-tuning large models, unlike all previous methods that required modifying the model architecture. Experimental results show that doing LoRA fine-tuning with AC-GPT on open-weight large models can get a clear improvement on arbitrary-conditional modeling tasks, without noticeably hurting unconditional generation ability.
I'm excited to share my latest work, which focuses on how to give causal transformers the ability to model arbitrary conditionals. This work was done with a bunch of excellent folks @EricElmoznino@g_lajoie_@le0gagn0n@sarthmit@tejaskasetty . https://t.co/paKZUnvZfZ
@EricElmoznino@g_lajoie_@le0gagn0n@sarthmit@tejaskasetty [8/9] Experiments show that AC-GPT and σ-GPT-temporal are far stronger than all baselines that can model arbitrary conditionals, including the traditional σ-GPT, on arbitrary-conditional modeling tasks, and are very close to vanilla GPT on pure left-to-right generation tasks.