Revolutionary work from Huggingface.
Nothing less.
---
FineWeb 🍷 | TL;DR :
0. Pretraining is far less intuitive than instruct finetune
1. Unclear what data to include to boost performance
2. HF's team simply tried different filters [3] -> trained small models on the filtered dataset -> measured eval scores.
Result:
The first version of FineWeb 🍷: A pre-training dataset that is was methodically selected for the exact purpose of making the models you train on it get the best eval scores as possible.
---
FineWeb-Edu 📚 | TL;DR:
0. Filtered the dataset even more to include only "high educational value" texts. [1]
2. Used Llama-3-70B-Instruct [2] to extract "educational score" for 500K texts.
3. Trained a classifier to classify the rest of the 15T tokens.
Result:
FineWeb-Edu 📚 outperform every other open pre-training dataset by a huge margin.
---
Other treasures in the blog
In the blog there is a ton of useful material: how to filter at scale? how to run LLMs at scale? breakdown of every single filter contribution to the score..
To the best of my knowledge this is the first large scale empirical study of pre-training data trying to improve an underlying model's performance.
Extremely high quality.
Seriously, hats off.
-
[0] On the screenshot: Harsh truth.
[1] As proposed on "Textsbooks are all you need" (phi-1): https://t.co/zrxijCsBgM
[2] Language filter -> Gopher filter (from it's paper) -> MinHash -> C4 filter (simple rules in the paper) -> PII filter.
[3] Prompt from https://t.co/BZ7HDFaXnK
Interesting experiment from Google that creates an NPR-like discussion about any academic paper.
It definitely suggests some cool possibilities for science communication. And the voices, pauses, and breaths really scream public radio. Listen to at least the first 30 seconds.
🤖Building an AI Agent With Memory Using MongoDB, Fireworks AI, and LangChain
This tutorial provides a step-by-step guide on building an AI research assistant agent
Has insights into equipping agents with memory systems and knowledge management
https://t.co/Q1jmOdROa0
🤖Building an AI Agent With Memory Using MongoDB, Fireworks AI, and LangChain
This tutorial provides a step-by-step guide on building an AI research assistant agent
Has insights into equipping agents with memory systems and knowledge management
https://t.co/Q1jmOdROa0
🧞PDFs -> Knowledge Graphs
The PDF to knowledge graph conversion may be one of the trickiest things to get right
It's hard to parse PDFs, it's hard to construct knowledge graphs
Luckily our friends from @neo4j have a great video to get you started!
https://t.co/GGjd9NYlvx
🎥 Webinar Recording: Memory for Autonomous Agents 🧠
The recording of our recent webinar on memary - a fully open-source reference implementation for long-term memory in autonomous agents - is now available.
In this session, the authors of memary, @JulianSaks, Kevin Li, Seyeong Han provide a deep dive into the project and engaged in a thought-provoking discussion and Q&A session around memory, including a multi-agent future with co-learning.
In case you haven’t seen it, memary introduces a novel architecture that combines sequential memory with knowledge graphs. This helps provide personalization in the context window but also maintains a larger KG you can query.
If you missed the live webinar or want to revisit it, you can now watch the recording here: https://t.co/5GvBvBCsrR
Check out the source repo too: https://t.co/W6wTGn7hAe
Los lingüistas sabemos cómo funcionan las lenguas, pero son sus hablantes los que nos enseñan por qué no hay nada más importante que ellas.
Este es nuestro homenaje para todos los que nos habéis compartido vuestro tesoro. ¡Gracias!
#DíaDeLaLenguaMaterna
🔗https://t.co/3ydTXfNXug
¡Divulgación y transferencia de conocimiento científico a través de #Wikipedia! Una gran iniciativa y experiencia 👏👏👏
.
Esperamos seguir trabajando y promoviendo el conocimiento libre💚.
#CampañaWMES | Súmate a #1bib1ref donde podrás contribuir a verificar fuentes en Wikipedia 🌐.
.
➡️ Tienes hasta el 5 de junio! 📚.
👉 https://t.co/yRL0Ir7znC
Maiatzean wikilariek 947.000 hitz gehitu dituzte euskarazko Wikipedian. Inprimatu bagenu, Lur Hiztegi Entziklopedikoaren 328 liburu hartuko lituzke. Azken hilabetean 4 liburu idatzi ditugu.
https://t.co/gu0ihDGsRa
Twenty-four individuals from thirteen countries across the world gathered in Cambridge this April to take part in the seventh biannual Cultural Heritage Data School at Cambridge Digital Humanities.
Read what participants had to say ⬇️
https://t.co/E2peIz0nlN
@mentxuwiki La inteligencia artificial ha dado pasos de gigante en el último lustro. A mi me tiene fascinado y lo que hay que hacer es establecer pautas de convivencia y buen uso.
#dhum1727
🟡 Estamos muy contentas de haber sido seleccionadas para participar en el próximo #XEncuentroCulturayCiudadanía junto a otro proyectos que conocemos y admiramos. ¡Nos veremos en septiembre en Santiago de Compostela!
#JuntasSomosMásVisibles
Thrilled to announce that I joined OpenAI a few months ago! I'm back in SF and working out of the headquarters.
A special thank-you to my management chain and the many friends that made this happen. If you and I have worked together in the past and you'd like a referral, please reach out!
A #mindmap of the paper "The Monsters Within: Gothic Monstrosity in Dracula, Frankenstein, and Strange Case of Dr. Jekyll and Mr. Hyde and its Role in the Nineteenth Century English Society" by Rubén Ortiz Trueba. #dhum1727