🚨 How do we build culturally grounded data for #LLMs?
Excited to share that ✨From National Curricula to Cultural Awareness✨ was accepted to #EMNLP2026 Main 🎉
TL;DR
We use national curricula to build supervision for cultural alignment.
📄 https://t.co/9455vxRr1P
(1/🧵)
Google DeepMind argues RAG is broken.
They published a paper that proved vectors databases are the dead end.
For the last three years, the default engineering response to any AI memory or data problem has been identical: "Just build a RAG pipeline."
Chunk the data, push it into a vector database, and let embeddings handle the rest.
Every company scaling enterprise AI assumes that if an embedding model fails, it's just a matter of time. Better training data, larger models, more parameters—throw compute at it, and the search gets smarter.
This paper proves that assumption is completely false.
They mathematically demonstrated that single-vector embeddings have a hard, uncrossable limit.
Here is the core flaw:
An embedding compresses an entire document or a complex query down into a single fixed-length vector of numbers.
When you run a search, the model takes the dot product of those vectors to measure similarity.
The math reveals a brutal constraint. The number of distinct document combinations a model can possibly retrieve for different queries is strictly bounded by the dimension of its embedding space.
It is a hard mathematical ceiling dictated by geometry and communication complexity.
No amount of data scaling can fix it. No amount of fine-tuning will punch through it.
Even if you give an embedding model infinite, unconstrained training freedom on the test set, it still hits the wall.
DeepMind built a stress-test dataset called LIMIT to prove it.
They threw state-of-the-art embedding models at it, models with thousands of dimensions.
The models completely failed. Even on simple, structured queries, the single-vector bottleneck forced the system to drop critical context and hallucinate irrelevant results.
Why? Because a single vector cannot capture complex, multi-faceted relationships between documents.
When you ask an AI to reason, follow complex instructions, or handle nuanced cross-document dependencies, the vector space simply runs out of room.
It collapses.
This changes everything for software architecture.
If your AI agent's memory relies on standard single-vector retrieval, it is structurally blind to complex logic. It is missing pieces of your data right now, and no prompt tweak can save it.
If we want AI that actually understands enterprise knowledge, we have to throw out the single vector.
And invent something entirely new.
TMLR has been facing an significant uptick in the number of submissions since the start of the year. This is placing an extreme burden on our amazing team of reviewers and action editors.
To ease this burden, TMLR will be implementing submission quotas, effective July 1. 1/n
Attended conference for an in-person presentation, after 2 yr.s since ACL 2024! A bit early field trip for 8-months-old baby though 😅
Presented a joint work with Dr. Geunhye Kim of HUFS, on the speech act of Korean wedding invitation
Recently, there's been complaints on low-quality AI reviews at conferences and journals. What if we put the frontier LMs into an agent harness?
With the right setup, on 82 Nature-family papers, 45 expert scientists judged that AI reviewers outperform the best human reviewer!
🤗 https://t.co/PI9WwCTzRq
Days in a work week: 5
Days in a month: 30
Total new submissions to arXiv in September: 26,646
arXiv editorial and user support staff: 7
someone who is good at science please help me with this. our team isn't sleeping.
#openaccess#preprints
we all quote Firth for "You shall know a word by the company it keeps!" perfectly fine, as he indeed said so in his book.
but, how many of us knew which word he had in his mind when he wrote this sentence?
🤣
LLM-as-a-judge has become a norm, but how can we be sure that it will really agree with human annotators? 🤔In our new paper, we introduce a principled approach to provide LLM judges with provable guarantees of human agreement⚡
#LLM#LLM_as_a_judge#reliable_evaluation 🧵[1/n]
Thank you for an intersting talk and for sharing your expertise with SKKU student, @warnikchow
"Lifeless comp linguists annotate furiously - 무기력한 내 일상 속 작은 데이터 어노테이션"
Currently at PolyU HK for #PACLIC 2023 and will present a poster on Korean dataset studies in the afternoon :)
It's a meta-research -- a research on research -- for practitioners, and you may find other interesting studies in the conference. Please drop by!
About a week ago -- attended #EAAMO2023 and presented our PaperCard paper with @ejcho95 and @kchonyc
Truely great people and community :) Thanks for all the constructive comments, discussions, and engaging interactions!
arXiv version available in:
https://t.co/GlhEdwQlyi
📢IT'S OFFICIAL!
🇧🇷The ACM Conference on Fairness, Accountability, and Transparency #FAccT2024 will be held Monday, June 3rd through Thursday, June 6th, 2024 in Rio de Janeiro, Brazil!
https://t.co/uIhXHmp0dY
📅Stay tuned for the CFP
I get asked a lot: why stay in academia, all the excitement in AI is happening in industry with massive compute. And I am seeing some profs leaving academia, but also seeing lots of researchers in industry looking to go back to academia, especially those who don’t work on LLMs.
I have spent some time and actively working with industry, and it has been a great experience, but for me the answer is simple:
1. Students: Nothing can replace working with really smart & amazing students.
2. Academic freedom: In general, industry will one way or another dictate what you should work on. Today it is LLMs, tomorrow it is something else. The good old days of "here is a cool idea, let me investigate and publish" are pretty much over. But as an academic, I can work on whatever I want: I can start working on "black holes" tomorrow if I choose to.
3. And my favourite one: No reorgs -- big tech really loves reorgs that happen every few months.
Finally, as an academic, I can always take some time off, work in industry or start-up, and come back if I want to. Going into industry is usually a one-way street: It is far easier to go to industry from academia and much harder the other way around.
Today we're sharing details on AudioCraft, a new family of generative AI models built for generating high-quality, realistic audio & music from text. AudioCraft is a single code base that works for music, sound, compression & generation — all in the same place.
More details ⬇️
I have almost 30k citations and millions in research funding.
But almost 20 years ago, I just wrote my first research paper.
I had no writing support during grad school & I want you to have it better.
Here are secret actionable writing tips for 6 sections of your paper. ↓