Ever feel like it's too hard to keep track of what LLMs cannot do as well as humans? We're making your life easier over at: https://t.co/dESLuXAmET
We're compiling a list of papers testing the abilities of LLMs against humans. Check it out! And you can help contribute too!
Excited to kickstart this effort with @glnmario and a star-studded advisory team! 🚀
Also, we're hiring for 2 MTS and 1 Ops manager! Come build the future of white-box audits with us: https://t.co/82Xp4vUak9
2. Positional Overload: Positional Debiasing and Context Window Extension for Large Language Models using Set Encoding (PDF coming soon!). Lukas Kinder at KIT did all the hard work, showing that we can use shared position IDs across texts to extend context length and debias MCQA
Happy to announce two papers at #ACL2025!
1. We made an extension to our CUTE benchmark, EXECUTE! We find that LLMs actually *do* have the ability to do character-level manipulation, but language bias, and tokenization too, get in the way. https://t.co/DnsaRvheOb
1/2
@bminixhofer@licwu@PontiEdoardo@huggingface Cool work! I'd love to test these byte models on our CUTE benchmark, I suspect they will do quite well. Any plans on uploading a safetensors version?
I'm happy to announce that our LLM benchmark for measuring orthographic understanding, CUTE, was accepted to EMNLP main! Looking forward to seeing you all in Miami! #NLProc
Arxiv: https://t.co/1gkdoxJUKf
HF Leaderboard: https://t.co/sR0SqVFvBo
Github: https://t.co/3ioe8pb81n
@vansteenkiste_s@jowenpetty We compared char-level to subword-level for NMT, and found the char models are generally on-par or better, and show word-level processing capabilities when needed. Indeed, the major issue seems to be that inference is ~5x slower.
Link: https://t.co/fVJfSDygJ4