LLM model size competition is intensifying… backwards!
My bet is that we'll see models that "think" very well and reliably that are very very small. There is most likely a setting even of GPT-2 parameters for which most people will consider GPT-2 "smart". The reason current models are so large is because we're still being very wasteful during training - we're asking them to memorize the internet and, remarkably, they do and can e.g. recite SHA hashes of common numbers, or recall really esoteric facts. (Actually LLMs are really good at memorization, qualitatively a lot better than humans, sometimes needing just a single update to remember a lot of detail for a long time). But imagine if you were going to be tested, closed book, on reciting arbitrary passages of the internet given the first few words. This is the standard (pre)training objective for models today. The reason doing better is hard is because demonstrations of thinking are "entangled" with knowledge, in the training data.
Therefore, the models have to first get larger before they can get smaller, because we need their (automated) help to refactor and mold the training data into ideal, synthetic formats.
It's a staircase of improvement - of one model helping to generate the training data for next, until we're left with "perfect training set". When you train GPT-2 on it, it will be a really strong / smart model by today's standards. Maybe the MMLU will be a bit lower because it won't remember all of its chemistry perfectly. Maybe it needs to look something up once in a while to make sure.
Here are the most FASCINATING facts I could find about Mexico:
1. Mexico is so huge it’s hard to comprehend. You can fit 30 European countries in Mexico and still have room to spare
(Bonus)/ LlamaFS
Local LLM-powered hard drive file organizer. Automatically rename and categorize messy files and directories with multi-modal AI
Everyone was asking me nonstop what we built, so here it is :)
@swayingoak@iyajainfinity @AshwinHegd28838
We build AI to empower people, including journalists.
Our position on the @nytimes lawsuit:
• Training is fair use, but we provide an opt-out
• "Regurgitation" is a rare bug we're driving to zero
• The New York Times is not telling the full story
https://t.co/S6fSaDsfKb
A reason we need beneficial AGI:
After five years of pain across many systems in her body (a broken foot from stepping off a curb, debilitating migraines, fatigue, joint pain and instability, etc), my wife was recently diagnosed with a genetic disorder called Hypermobile Ehlers-Danlos Syndrome (hEDS).
Because the medical system is designed for individual specialties while hEDS affects every system in her body (orthopedics, cardiology, neurology, gastroenterology, dermatology, etc), we spent five years seeing more doctors and specialists than in her whole life prior. Most doctors would only focus on what was relevant to their own specialty. We were lucky that her allergist (!!) put together the pieces after observing and hearing her full set of symptoms and issues.
As human medicine has progressed, it seems like we increase doctors’ depth at the expense of breadth. We need better tools to be able to deliver depth and breadth simultaneously to patients. This is one promise of AGI if built right — reliable, individualized, affordable healthcare in your pocket, like a panel of today’s top doctors across every speciality working together in concert to keep you healthy (and without you needing to fax forms between them).
There’s still a long way to go on the technology and on learning how to deploy it beneficially along with appropriate professional human oversight in high-stakes areas like medicine, but the promise is getting increasingly clear. Thoughtfully approached by technology developers, healthcare providers, governments, and society, there’s hope for much better care for every member of all of our families (including our non-human furry ones).