Instruction-following Evaluation through Verbalizer Manipulation
paper page: https://t.co/Aeas20H5Js
While instruction-tuned models have shown remarkable success in various natural language processing tasks, accurately evaluating their ability to follow instructions remains challenging. Existing benchmarks primarily focus on common instructions that align well with what the model learned during training. However, proficiency in responding to these instructions does not necessarily imply strong ability in instruction following. In this paper, we propose a novel instruction-following evaluation protocol called verbalizer manipulation. It instructs the model to verbalize the task label with words aligning with model priors to different extents, adopting verbalizers from highly aligned (e.g., outputting ``postive'' for positive sentiment), to minimally aligned (e.g., outputting ``negative'' for positive sentiment). Verbalizer manipulation can be seamlessly integrated with any classification benchmark to examine the model's reliance on priors and its ability to override them to accurately follow the instructions. We conduct a comprehensive evaluation of four major model families across nine datasets, employing twelve sets of verbalizers for each of them. We observe that the instruction-following abilities of models, across different families and scales, are significantly distinguished by their performance on less natural verbalizers. Even the strongest GPT-4 model struggles to perform better than random guessing on the most challenging verbalizer, emphasizing the need for continued advancements to improve their instruction-following abilities.
AlpaGasus: Training A Better Alpaca with Fewer Data
Significantly outperforms the original Alpaca and reaches 90% of davinci-003 w/ 5.7x faster training.
proj: https://t.co/Gy8JMehYka
abs: https://t.co/byFO28kI17
AlpaGasus: Training A Better Alpaca with Fewer Data
Significantly outperforms the original Alpaca and reaches 90% of davinci-003 w/ 5.7x faster training.
proj: https://t.co/Gy8JMehYka
abs: https://t.co/byFO28kI17
AlpaGasus: Training A Better Alpaca with Fewer Data
paper page: https://t.co/WylXf7Ng4d
Large language models~(LLMs) obtain instruction-following capability through instruction-finetuning (IFT) on supervised instruction/response data. However, widely used IFT datasets (e.g., Alpaca's 52k data) surprisingly contain many low-quality instances with incorrect or irrelevant responses, which are misleading and detrimental to IFT. In this paper, we propose a simple and effective data selection strategy that automatically identifies and removes low-quality data using a strong LLM (e.g., ChatGPT). To this end, we introduce AlpaGasus, which is finetuned on only 9k high-quality data filtered from the 52k Alpaca data. AlpaGasus significantly outperforms the original Alpaca as evaluated by GPT-4 on multiple test sets and its 13B variant matches >90% performance of its teacher LLM (i.e., Text-Davinci-003) on test tasks. It also provides 5.7x faster training, reducing the training time for a 7B variant from 80 minutes (for Alpaca) to 14 minutes We apply IFT for the same number of epochs as Alpaca(7B) but on fewer data, using 4timesNVIDIA A100 (80GB) GPUs and following the original Alpaca setting and hyperparameters.. Overall, AlpaGasus demonstrates a novel data-centric IFT paradigm that can be generally applied to instruction-tuning data, leading to faster training and better instruction-following models.
@amitness@Thom_Wolf @colinraffel @PatrickPlaten @GoogleAI Hi Amit, do you find anything to finish this? I didn't find it. I want to do something like a circle translation, en-->fr-->en
Family first
Keep in touch with friends
You are not your job
Do not respond to negativity
Fight against entitlement
Be honest
Be kind
Forgive first
Drink water
Eat less sugar
Be humble
Learn to dance
Know when to leave
Tip well
Don't nitpick
Learn to learn
Laugh loudly
Write more
As academics, we’re surrounded by people, yet our jobs can be extremely isolating. Put people in your life with whom you can share your struggles, celebrate your victories, and support each other through both. Emotionally healthy people are happier and more effective people.
"Failing" is too often portrayed as something negative. Failures, and not successes, have been a higher drive for me both personally and professionally (and there have been plenty!). If you never fail, you are doing it wrong : )