Reinforcement Learning in the era of LLMs requires scalable, distributed systems to push the boundaries of reasoning and alignment.
Today - we release Atropos - our RL environments framework.
https://t.co/gJGzN7MS5T
Atropos is a rollout framework for reinforcement learning with foundation models that supports complex and diverse environments for advancing the capabilities of foundation models.
In Greek mythology, Atropos was the eldest of the three Fates. While her sisters spun and measured the threads of mortal lives, Atropos alone held the shears that would cut these threads, determining the final destiny of each soul. Just as Atropos guided souls to their ultimate fate, this system guides language models toward their optimal potential through reinforcement learning.
The work on Atropos was led by @dmayhem93 and built alongside @teknium, @rogershijin, @max_paperclips, @nullvaluetensor, @JSupa15, @artemsya and @karan4d
Glad to see async LLM RL getting more recognition! Our Towards System 2 paper highlights its importance under infrastructure: https://t.co/uuHa32iFGh. Also, NeoX already supports this (https://t.co/BcKK1JquRr) with shared memory, allowing you to save vram by using the same memory for training and inference, making in-flight weight updates instant
Generative Reward Models
Presents GenRM, an iterative algorithm that trains an LLM on self-generated reasoning traces, leading to synthetic preference labels matching human preference judgments
https://t.co/f7oS828PwS
My favorite bit in this paper:
I and @bbrabbasi wrote an appendix formalizing what is done evaluating models with loglikelihood multiple choice and perplexity evals.
afaik, none of this has been written up in one place in most papers and just been tacitly assumed before!
@TeraflopAI@Geronimo_AI@dmayhem93@StabilityAI Yup! We use the FLAN sources from https://t.co/E9hTYr6mBP collected by TeraflopAI's very own @EnricoShippole. It's mentioned in our 1.6B technical report and tagged in the metadata of all Stable LM 2 base model cards.
We are releasing trillions of high-quality, copyright-free, permissively licensed tokens and multimodal data. Be sure to follow our releases @TeraflopAI.
Stable LM 2 12B is a pair of powerful 12 billion parameter language models trained on multilingual data in English, Spanish, German, Italian, French, Portuguese, and Dutch, featuring a base and instruction-tuned model.
You can now try the model here: https://t.co/jLzZdYBUkJ (1/3)
Stable LM 2 12B is a pair of powerful 12 billion parameter language models trained on multilingual data in English, Spanish, German, Italian, French, Portuguese, and Dutch, featuring a base and instruction-tuned model.
You can now try the model here: https://t.co/jLzZdYBUkJ (1/3)
How to define Diversity in the context of CodeLMs and Programming Languages ?
1. Diversity is positively correlated with Performance in solving a problem.
2. Shortcomings of diversity in small codeLMs.
3. Code Embedding models don't capture semantics.
https://t.co/aNGhSpxz48
@TeraflopAI is excited to help support the @caselawaccess and @HarvardLIL, in the release of over 6.6 million state and federal court decisions published throughout U.S. history 🥳
Releasing Yarn-Llama-2-13b-128k, a Llama-2 model, trained for 128k context length using YaRN scaling. The model was trained in collaboration with u/bloc97 and @theemozilla of @NousResearch and @Void13950782 of @AiEleuther.
Releasing a new PaLM 2.1b model trained at a context length of 8k on C4. This model release is a continuation of the previously released 150m, 410m, and 1b models.
https://t.co/0L7V7m7F6n
Introducing three new open-source PaLM models trained at a context length of 8k on C4. Open-sourcing LLMs is a necessity for the fair and equitable democratization of AI. The models of sizes 150m, 410m, and 1b are available to download and use here: https://t.co/g1J6E22yBf