🎬 Ever wondered how to turn dense research into a 60-sec short-form explainer?
Meet SciTalk—a creator-inspired, multi-agent framework that iteratively refines scientific videos via prompt-driven feedback loops.
🔗 https://t.co/zEjcpCNSIk
🤖Excited to share @MinnesotaNlp Labs latest work on Data Efficient Instruction Tuning: "SelectLLM: Can LLMs Select Important Instructions to Annotate?". Leveraging LLMs to select high quality instructions. Arxiv: https://t.co/M0OEF5XINm #LLM#NLP#AI#ArtificialIntelligence
In our benchmark, we highlight potential flaws in employing LLMs as automatic evaluators, finding that most models are affected by various cognitive biases when making evaluations.
Project Page: https://t.co/eYa0Io5TmB
Codebase: https://t.co/OCgqse9DdV
Our benchmark (CoBBLEr: Cognitive Bias Benchmark for LLMs as Evaluators) studies how LLMs-as-evaluators are impacted by certain artifacts in prompts (e.g., prompting format, prompt information) that negatively influence their ability to perform unbiased evaluations.
Analyzing ~630k total samples, we investigated 6 different cognitive biases from 15 language models instruction tuned or trained on human feedback and found that almost all models show, on average, 40% of responses being cognitively biased.
LLMs have proven to outperform humans on a multitude of tasks. Does this also mean they are more biased too? In our work, we benchmark several different LLMs as automatic evaluators for various cognitive biases.
https://t.co/nqoyBazdcN