Got 3 papers accepted at #EMNLP2026🥳!
📄 HALvest-Contrastive: Retrieval-like Authorship Attribution with Patch-level Late Interaction (Main)
📄 Language-Switching Triggers Take a Latent Detour Through Language Models (Main)
📄 LLM Forensics: Where Do Backdoors Hide? (Findings)
🧵New paper: "Lost in Backpropagation: The LM Head is a Gradient Bottleneck"
The output layer of LLMs destroys 95-99% of your training signal during backpropagation, and this significantly slows down pretraining 👇
Thrilled to announce that I have received the @AtalaOrg Best PhD Dissertation Prize at CORIA-TALN 2025! 🇫🇷
For the bravest, my dissertation is now available online: https://t.co/ZCTYo8mp5G
Big thanks to the jury and to my great supervisors @bensagot and @DeVillemonte !
Excited to introduce 𝗕𝗶𝗼𝗺𝗲𝗱-𝗘𝗻𝗿𝗶𝗰𝗵𝗲𝗱 🎉, a new annotated biomedical dataset designed to tackle the scarcity of clinical data for NLP research!
133M paragraphs from PMC-OA annotated for type, domain, and educational quality and publicly available on @huggingface👇🧵
ModernBERT or DeBERTaV3?
What's driving performance: architecture or data?
To find out we pretrained ModernBERT on the same dataset as CamemBERTaV2 (a DeBERTaV3 model) to isolate architecture effects.
Here are our findings:
llm's are surprisingly bad at copying writing style.
i have been trying to get claude to sound like me and it's no where close EVEN with ~30 examples.
The one to solve this in 2025 will be a billionaire
CamemBERT 2.0: A Smarter French 🇫🇷 Language Model Aged to Perfection 👌
We release a much-needed update for the previous. SOTA French encoder LM.
We introduce two new models CamemBERTa-v2 and CamemBERT-v2, based on the DeBERTaV3 and RoBERTa recipe.
So what's new?
[1/8]
CamemBERT 2.0: A Smarter French 🇫🇷 Language Model Aged to Perfection 👌
We release a much-needed update for the previous. SOTA French encoder LM.
We introduce two new models CamemBERTa-v2 and CamemBERT-v2, based on the DeBERTaV3 and RoBERTa recipe.
So what's new?
[1/8]
Announcing mOSCAR, multilingual interleaved text-image corpus as part of @oscarnlp project.
Paper: https://t.co/1hhnyYCyI3
Dataset: https://t.co/hiEUJ1Q3iJ
Doc: https://t.co/KsnT5wVee2
1/6
🤏 Why do small Language Models underperform?
We prove empirically and theoretically that the LM head on top of language models can limit performance through the softmax bottleneck phenomenon, especially when the hidden dimension <1000.
📄Paper: https://t.co/YkdQttDDSK
(1/10)
Excited to share our latest research paper: "From Text to Source: Results in Detecting Large Language Model-Generated Content"
We research cross-model detection and model attribution, covering a wide range of LLM sizes and families.
Paper: https://t.co/WKCqUANUg0
A thread🧵
🔥 New pretraining method 🔥
Our models learn to recover masked input embeddings instead of predicting masked input tokens.
It leads to substantial improvements and efficiency gains for encoder AND decoder models.
🗞️ Paper: https://t.co/ByummSfMe9
🧵 below
📝 New paper accepted at the Student Research workshop of @aclmeeting !
We found out that anisotropy affects Transformers-based models not only for NLP but also character-level, speech and vision modalities !
Pre-print: https://t.co/zaRobzcg0s
🧵
We are proud to announce CamemBERT-bio, a state-of-the-art French biomedical language model 🎉
We release everything on @huggingface !
📜Pre-print: https://t.co/fHmqsDV4Jb
🤗Corpus: https://t.co/svtnTA8wAW
🤗Model: https://t.co/WM8acHILhb
Logo by @Alix_Tz