HELM has a new leaderboard: HELM Capabilities v1.0! We curated 5 challenging datasets (MMLU-Pro, GPQA, IFEval, WildBench, Omni-MATH) and evaluated 22 top language models:
🎮 Check out the live demo for LM-Steer “Word Embeddings Are Steers for LMs” (Outstanding Paper ACL 2024) at https://t.co/inhnGZo9Xv, which can:
1.🕹️Steer model generation
2.🔬Discover word embedding dimensions
3.📊Profile sentences & identify keywords
#ACL2024#LLMs#Science4LM
🎖Excited that "LM-Steer: Word Embeddings Are Steers for Language Models" became my another 1st-authored Outstanding Paper #ACL2024 (besides LM-Infinite)
We revealed steering roles of word embeddings for continuous, compositional, efficient, interpretable& transferrable control!
🎖Excited that "LM-Steer: Word Embeddings Are Steers for Language Models" became my another 1st-authored Outstanding Paper #ACL2024 (besides LM-Infinite)
We revealed steering roles of word embeddings for continuous, compositional, efficient, interpretable& transferrable control!
Reposting this gem from Michael. Sad to see that Silicon Valley and Stanford lost a legend who is wise, kind, and funny, with the amazing talent to always put smile on people’s faces. Check out this post if you are looking for an awesome car!
✨Attention, #SiliconValley!✨
As my time at @Stanford approaches its end, I’m selling this factory-grade, car-sized car!🤯
Ready for a yourself-driving experience? Meet my 2015 Subaru Impreza (51k miles young👀). It features a limited-edition interface called a 'steering wheel' that seemlessly connects to your hands (dual-wield mode) for real-time, user-driven navigation!
This manual car is so responsive, it feels like it’s reading your mind (but really, it’s just your hands on the wheel and the stick).
No terms and conditions to scroll through. No black-box algorithms, just pure, unadulterated mechanical feedback.
Say goodbye to software updates and hello to hardware that listens to you! Experience the thrill of making decisions at every single turn! This car is perfect for those who prefer to be in control rather than be controlled.
Put the 'I' back in driving, and hit me with a DM!
Multimodal LLMs (MLLMs) may have memorized the web, but do they really leverage that knowledge?
In this new preprint, we find that:
1. Reverse Image Retrieval (RIR) can distinctly boost even latest models like #GPT4o in knowledge-intensive tasks
-->e.g. RIR improves #GPT4V by ~40% on infoseek
2. Surprisingly, RIR can help by stimulating the MLLM's own parametric world knowledge (which may not get leveraged in regular visual query, see below)
In short: MLLMs do know more than they reveal in VQA!
Paper: https://t.co/ezNS5Y6V7m
Code: https://t.co/9qDaXxYNur
Grateful for this fun collab with @liamjxu* and @jure!
Some more teasers in short 🧵 below:
3 OVAL projects are awarded 2024-2025 Magic Grants!
“African History from the Bottom Up with LLM-Augmented Agents”, @sina_semnani et al.
“Cross-Lingual Multi-Perspective News”, @liamjxu et al.
“DataTalk: All Documents and Data, All at Once, All Verified”, @ShichengGLiu et al.
Steering LLMs towards desired stances and intentions:
Toxic → Non-toxic
Negative → Positive
Biased → Unbiased
And more...
How do we achieve this? Our research proves the role of word embeddings as effective steers.
The best part? These steers are compositional!
Check out our #ACL24 paper: https://t.co/YGgRxXbTwf!
Introducing SUQL (Structured and Unstructured Query Language), a novel method to power assistants on hybrid data corpus. It combines SQL relational operators with free text primitives based on retrievers and LLMs, complete with key optimizations
Code available on GitHub and PyPI
Excited that LM-Infinite has been accepted into #NAACL2024 ! It is the first-of-its-kind zero-shot length generalizations for language models, with 200M length inference and downstream (Retrieval, Qasper) improvements! Great thanks to all my collaborators!
https://t.co/T6MSXbtWpv
🚀 Unveiling SeqBoat 🚤 - A groundbreaking neural architecture accepted by #NeurIPS2023!
It eclipses Transformers with a 28.38% average accuracy boost on LRA benchmarks, 10.4x faster training, and 95% reduced memory usage for 4k input length!
📖 Paper: https://t.co/rJz7zLDq3m
Should we choose to use LLMs or smaller finetuned models in practical use cases? Take a look at our survey https://t.co/IO4kDGnDTn , which covers NLU tasks, Generation tasks, Knowledge-intensive tasks, abilities regarding scaling, some miscellaneous and real-world tasks.
Many unsolved problems exist in ML safety which are not solved by closed-source GPT models. As LLMs become more prevalent, it becomes increasingly important to build safe and reliable systems. Some key research areas:
🧵