Are you wondering how large language models like ChatGPT and InstructGPT actually work?
One of the secret ingredients is RLHF - Reinforcement Learning from Human Feedback.
Let's dive into how RLHF works in 8 tweets!
I've been comparing a lot of transformer variants on large models (400M params): Post/Pre-LN, DeepNet, NormFormers, Swin v2, GLU variants, RMSNorm, Sandwich LN, with GELU, Swish, SmeLU…
More than 2,000h of total training time on TPU v3's 😯
Here are my findings 🤓
“In the Same Breath,” which is now streaming on HBO, delivers empathetic and often harrowing scenes from the homes, hospitals, and streets of Wuhan. Making the film felt “therapeutic and cathartic,” the director said. https://t.co/qHHqCEV20w
📣 We need your help. This is the post we hoped to never write, but today marks a huge turning point in The Strand's history. Our revenue has dropped nearly 70% compared to last year, and the loans and cash reserves that have kept us afloat these past months are depleted.
Pessimism and a sense of powerlessness about the future are a near-perfect cocktail for misery. In "How to Build a Life," @arthurbrooks offers advice on how to cope: https://t.co/Mp9CTtTgUf