Proud to announce that our group has 3 papers accepted to the main conference of #ACL2025! Our contributions span from filtering structured supervision for task-specific pre-training to automatic long-context mastering. Looking forward to share ideas with you at Vienna! 🇦🇹🏰
🚀 New Preprint Alert: LLMs can be radically efficient when their pre-training objectives are tailored for specific paradigms! 🌱💡
Introducing Cuckoo 🐦: a 0.3B LM specialized for Information Extraction (IE), pre-trained solely on extractive patterns of language modeling.
🔓 Cuckoo unlocks the scaling potential of IE models by shifting from vocabulary prediction to token tagging - enabling lightweight, task-specific performance that scales effortlessly with trillions of LLM's pre-training data.
🔍 Key Insights:
🪶 Tiny Yet Mighty: 96% fewer parameters and 20X faster than standard 8B LLMs.
🎯 Paradigm-Tuned: Tags tokens instead of predicting them, aligning perfectly with IE tasks.
🚀 Free-Riding Evolution: Cuckoo leverages LLM advancements without extra cost.
🔥 Performance:
- Outperforms an 8B LLaMA on few-shot IE tasks (entities, relations, machine reading comprehension (MRC), & instruction-following).
- Beats IE pre-training on human-labeled web links, GPT-4-generated data, & MRC datasets.
📄 Paper: https://t.co/HGpw9YoC82
🛠️ Code: https://t.co/L3oUFyql5s
Time to let your Cuckoo fly! 🐦💥 #LLM #InformationExtraction #NLP
✨My New ICLR2025 pub: Some knowledge is memorized by LLMs🤖, but cannot be sampled by even prompting 1,000,000 times😵💫 - we provide a simple way to find it. Here's why and how.
Why a prompt cannot hold all potential answers? Because answer (next token) representations are grouped in clusters. One example is a character cluster (A, B, C, ...). Whenever the LLM wants to predict one of those characters (e.g., What is a white and black animal?➡️P➡️Panda🐼), it cannot avoid predicting other characters (e.g., What is a white and black animal?➡️Q, V, X, ...➡️❌) with high probability (rank). On the other hand, some correct answers in other clusters are lost because a cluster dominates the probability.
We find that LLMs memorize out-of-cluster knowledge even if it's almost impossible to sample it from the initial prompt. We get such knowledge by prompting "What is a white and black animal other than Panda?" This eliminates the probability of "P" (and the character cluster) and navigates the LLM to search in another cluster for the answer. The navigation successfully finds the correct knowledge that is missed by the initial prompt.
This phenomenon also reminds us of recent efforts to sample reasoning paths for reinforcement learning. Sometimes, we fail to sample any correct thinking path from the initial prompt. Maybe we need a smarter navigation strategy to explore all potential thoughts hidden by the spurious correlation between vocabularies.
➡️Dive deeper into our work, which includes the application of navigation to math reasoning and data synthesis: https://t.co/GP8DSlqKPy
➡️View the spurious correlation: https://t.co/cP8jP1fgie
➡️Code: https://t.co/ZtBPESY7l1
New preprint: Do you know LLMs 🤖 learn knowledge composition in their parameters? We attribute a certain kind of compositional generalization to the linear correlation between answers from LLMs for certain prompts.
That is to say, if LLM believes X lives in the city of Paris🗼, then LLM will also believe X lives in the country of France 🇫🇷. The correlation is also resilient to gradients, thus updating X’s residential city to NYC 🗽 will also update X’s residential country to USA 🇺🇸.
However, some correlations learned by LLM are imprecise. For example, Indianapolis 🏙️ is correlated with India 🇮🇳 rather than the USA 🇺🇸 (Ground-truth). We find that learning that X lives in Indianapolis 🏙️ will cause hallucinations as if X lived in India 🇮🇳.
➡️ Dive deeper into the connections of correlation, generalization, and hallucination: https://t.co/SKIFczOJQn
➡️ Code: https://t.co/zpqC7vBfUV
We're thrilled to announce that 4 groundbreaking papers by our brilliant students have been accepted to #ICLR2025! 🎉 Focused on core representations in #LLM, our research covers both their limitations and applications. Join us in Singapore 🇸🇬 for these exciting insights! 🚀
Thrilled to share that our brilliant team at #SDLab has published 6 papers (4 Main, 2 Findings) at #EMNLP2024! From incubating small models by LLMs to addressing risks in LLM benchmarking, our research pushes the boundaries of #LLM and #NLP. Huge congrats to the team! 🚀💡🌴☀🐊
(1/3) Tired of Reflection-70B controversy? We’d like to share our recent work on LLM Benchmark Cheating and Detecting. (Reflection-70B included in Fig. 3!)
Title: “Data Contamination Can Cross Language Barriers”
Paper: https://t.co/6vgs7LQCUv
Code: https://t.co/0dX8VFYQNf
Congratulations to the genius students of our SDLab for publishing 9 papers (4 Main + 5 Findings) at #ACL2024! Our research focuses on large language models integrated with extremely weak supervision, personalization, and innovative tools #LLM#NLP. Great work, team! 🌟📚🎉👩🔬👨🔬
Interested in tabular understanding with LLMs? Check out our Chain-of-Table at #ICLR2024 ! The poster will be at Hall B #73 10:45-12:45 am. Feel free to ask questions and chat 😃
Can LLMs debug programs like human developers? 🚀 Launching 🛠️LDB, a debugging framework with LLMs🧠! Paper: https://t.co/oWsClrM3rv
LDB mimics how devs debug—breaking down codes into basic blocks & tracking variables step-by-step via the runtime information, enabling LLMs to zoom in on errors more precisely.
Embrace a future where coding is smarter & debugging is a breeze. 🌟💻 #LLM #DebuggingInnovation
Demo: https://t.co/GlgJHOJDn8
Code: https://t.co/nb0YDLV85o
As a boss, you want to cluster topics that people care about but your embedder does it on opinions?
To analyze your texts smartly, check our new instruction-following embedder: InBedder!
⭐Star our repo to support us: https://t.co/zo2fH1e8fn
📄Paper: https://t.co/ix4J2rQ2L9
Transformer as a decision tree algorithm!?
Thrilled to share our latest work MetaTree🌳 trained from classical algorithms (greedy CART/optimal GOSDT), that can produce strong decision tree models.
Paper: https://t.co/o63WxXuIMJ
Code: https://t.co/9Zy99ofpPB
(1/n)
Check out our latest work: Chain-of-Table for tabular reasoning with LLMs! 🌟
arXiv: https://t.co/feaYJU9HFV
Chain-of-Table guides the LLM to generate a series of tabular operations step by step, provideing a more structured and clear representation of the reasoning process.
Complex tables are transformed into simpler and more manageable segments, leading to in-depth understanding & analysis.
#LLM #AI #TableBasedReasoning
Happy to share the acceptance of our paper, "Towards Few-shot Entity Recognition in Document Images: A Graph Neural Network Approach Robust to Image Manipulation" at LREC-COLING 2024! 🎉
@LrecColing#NLProc
(1/4)