Last night I taught nanochat d32 how to count 'r' in strawberry (or similar variations). I thought this would be a good/fun example of how to add capabilities to nanochat and I wrote up a full guide here:
https://t.co/fz1AMI5kqk
This is done via a new synthetic task `SpellingBee` that generates examples of a user asking for this kind of a problem, and an ideal solution from an assistant. We then midtrain/SFT finetune on these to endow the LLM with the capability, or further train with RL to make it more robust. There are many details to get right especially at smaller model sizes and the guide steps through them. As a brief overview:
- You have to ensure diversity in user prompts/queries
- For small models like nanochat especially, you have to be really careful with the tokenization details to make the task easy for an LLM. In particular, you have to be careful with whitespace, and then you have to spread the reasoning computation across many tokens of partial solution: first we standardize the word into quotes, then we spell it out (to break up tokens), then we iterate and keep an explicit counter, etc.
- I am encouraging the model to solve the model in two separate ways: a manual way (mental arithmetic in its head) and also via tool use of the Python interpreter that nanochat has access to. This is a bit "smoke and mirrors" because every solution atm is "clean", with no mistakes. One could either adjust the task to simulate mistakes and demonstrate recoveries by example, or run RL. Most likely, a combination of both works best, where the former acts as the prior for the RL and gives it things to work with.
If nanochat was a much bigger model, you'd expect or hope for this capability to more easily "pop out" at some point. But because nanochat d32 "brain" is the size of a ~honeybee, if we want it to count r's in strawberry, we have to do it by over-representing it in the data, to encourage the model to learn it earlier. But it works! :)
Grâce à l’application Saytù Hémophilie, un chatbot en wolof, la sensibilisation progresse. Baamtu est fier d’accompagner l’Université de Genève et le CNTS dans cette initiative qui facilite l'accès à l'information pour les patients :
https://t.co/05mc9z9IR3
Our Llama Stack Distribution is a huge step forward in how we support developers with a single endpoint. We are now sharing with the community a simplified and consistent experience that will enable them to work with Llama models in multiple environments, including on-prem, cloud, single-node, and on-device
Hello Twitter,
Vous souhaiteriez vous lancer dans le monde de l'Analytics, ou découvrir ce métier en pleine expansion?
Rejoignez nous pour un webinaire et découvrez les compétences clés et les opportunités dans le domaine.
Lien d'inscription : https://t.co/lg0WVssZLX
🌟 2023 chez BAAMTU: Révolution data, soutien aux start-ups, avancées en IA.
Merci à nos clients et partenaires pour la confiance.
🎉 Pour une Année 2024 innovante et impactante !
Bonne et Heureuse Année !
#BAAMTUTechnologies#NouvelAn2024#Innovation
Do you want to expand your knowledge of the latest techniques in Retrieval Augmented Generation?
Join our latest course, built in collaboration with @truera_ai and @llama_index, and obtain the tools to build production-ready RAG applications.
Join now: https://t.co/yNQboGVeRp
Nous sommes en Aix-en-Provence pour parler IA et développement dans le cadre du #EmergingValley2023.
Un grand merci à @expertisefrance pour l'invitation 😊
We’re excited to introduce RAGs v2 - build, customize, and use multiple ChatGPTs over your data, all with natural language 💬
A huge upgrade vs. the initial launch:
💫 Easily create multiple RAG pipelines and save them
💫 Easily swap between and customize each one (e.g. over different data, or w/ different system prompts)
💫 Delete unused RAG pipelines
💫 (dev quality) added much-needed linting/CI
Check out the video 🎥 for details. It’s super easy to setup and use. Some additional features:
🧠 Supports a lot of LLMs both for building RAG and within each RAG pipeline
🌐 Supports loading load files or web pages.
Check out our repo here: https://t.co/838BDVOEbA
Salam je viens de perdre mon porte monnaie à Thiès.
Il contenait carte grise, carte étudiant (Moustapha KACHAB), carte bancaire et une somme d’argent.
Prière d’appeler au 785395220.
Rt appréciés 🚨
Depuis 2016, nous avons audité et conseillé 15+ startups africaines qui ont levé 300M$! Découvrez notre retours d'exp ici : https://t.co/CwgvwnsVae. Merci à
@orangestartupst , @free_senegal , @Adepme , @ITCnews et @tidjanedeme pour votre participation. Plus de vidéos à venir !
#JobAlert: Nous recrutons un Sénior Data Scientist.
Passionné d'IA, NLP, OCR ? Bac+5 en info ou domaine connexe, 5 ans d'expérience minimum ? Rejoignez notre équipe dynamique! Envoyez votre CV à [email protected].
Plus de détails : https://t.co/0Q0R0ZHeSr