Why exactly does fine-tuning make language models worse at some tasks? Can this change how we use models like ChatGPT? In recent work with @AdtRaghunathan and @jacspringer, we find that “catastrophic forgetting” might not be as catastrophic as expected.
https://t.co/N9FwXMah0e