Look at this guy.
He’s a millionaire who went homeless as a social experiment.
His goal? Prove that anyone can make $1M in 12 months with just a phone.
The story of Mike Black and what happened next:
Fine-Tuning Is Typically A Bad Idea
There is a lot of misinformation in the market, and companies often ask for a "custom model," their own LLM, or want to fine-tune an LLM
Unless you need a very specific type of output/task - for example, a particular type of equation or some example of XML, fine-tuning makes everything worse in the medium to long run.
Here is why
- Your fine-tuned LLM will likely be dead on arrival. Every day, we have a new LLM release. The chances that your fine-tuned LLM will be SOTA after 3 months is ZERO
- RAG + extended context is a killer combination, and fine-tuning ONLY makes sense if this killer combination does not work
- LLMs are still in their infancy, and swapping them out should be dead easy. Otherwise, your app will soon start rotting. Ideally, you want to use the LLM for reasoning and not for information retrieval
- Post-processing - You are better off post-processing an LLM response at times vs. fine-tuning. Again, this allows you to swap out the LLM at any time.
Here's my conversation with Sam Altman (@sama), his 2nd time on the podcast. We talk about the board saga, Elon lawsuit, Ilya, Sora, GPT-5, $7 trillion in compute, open source, and AGI. This was a truly fascinating conversation.
It's here on X in full, and is up on YouTube, Spotify, and everywhere else. Links in comment.
Timestamps:
0:00 - Introduction
1:05 - OpenAI board saga
18:31 - Ilya Sutskever
24:40 - Elon Musk lawsuit
34:32 - Sora
44:23 - GPT-4
55:32 - Memory & privacy
1:02:36 - Q*
1:06:12 - GPT-5
1:09:27 - $7 trillion of compute
1:17:35 - Google and Gemini
1:28:40 - Leap to GPT-5
1:32:24 - AGI
1:50:57 - Aliens
A recent survey on Retrieval-Augmented Generation (RAG) mentions an evolving paradigm:
Modular RAG.
Modular RAG is comprised of various functional modules. Thus, modular RAG is not standalone. Instead, different RAG patterns are composed of different modules.
For example, the following animation shows:
🥚 The original naive RAG paradigm consists of the “Retrieval”, "Augmentation," and "Generation" modules.
🐣 After naive RAG has shown some limitations, advanced RAG has emerged as a new paradigm. A typical pattern of Advanced RAG builds upon the foundation of Naive RAG by adding “Rewrite” and “Rerank” modules.
🐓 Different RAG patterns, such as DSP, can be composed of entirely different modules.
The modular RAG paradigm is slowly becoming the norm in the RAG domain due to its versatility and flexibility, allowing:
- the adaption of modules within the RAG process to suit your specific problem,
- for a serialized pipeline or an end-to-end training approach across multiple modules.
I definitely recommend checking out the full survey if you want to catch up on recent advancements in the RAG domain:
🔗 https://t.co/dWkf0Uc587
Self-attention is at the core of transformer-based LLMs. It comes in several flavors. In my new article (https://t.co/VfgBeGGRtB) I cover the various variants and code them in PyTorch from the ground up:
1. Self-attention: The fundamental building block of transformer-based LLMs that allows models to weigh the importance of different parts of the input data.
2. Multi-head attention: An extension of self-attention that allows the model to simultaneously focus on information from different representation subspaces.
3. Cross-attention: A variant that enables the model to attend to two different sequences, which makes it useful in tasks like translation or summarization.
4. Causal-attention: A variant to ensure that the prediction for each token depends only on the preceding tokens, which is important for text generation, where each prediction should be based only on the prior context.
In my opinion, coding algorithms, models, and techniques from scratch is not only fun but an excellent way to learn!
My recommendations on where to focus on 2024 (for those who care about AI):
I've been thinking a lot about maximizing your chances of finding work that pays well next year. Based on what I'm hearing, the market for people with the following skills will continue to heat up:
• Ability to build applications leveraging Large Language Models.
• Retrieval Augmented Generation (RAG) workflows
• Fine-tuning open-source models
• Deploying open-source models
• Ability to build things that work (Engineering skills)
2023 was the year when we built powerful AI models. 2024 will be the year when we'll integrate these models everywhere.
Some wonder why learning prompt engineering and integrating with GPT-4 is not enough. GPT-4 is a great model, but most conversations I've had are about integrating open-source models.
GPT-4 will get you better performance out of the box. An open-source model will give you privacy, flexibility, lower long-term costs, reliability, and transparency. Some open-source models will even match or beat GPT-4 at different tasks. You can't beat that package.
Closed-source models are great for prototyping and getting the first version out there, but you'll eventually need to migrate to an open-source model.
If you want to start building AI applications to sell your time and skills, look into some of the leading open-source models in the market right now. Here are five of them:
• Meta's Llama-2
• HuggingFace's Zephyr Beta
• Mistral AI's Mistral
• Stability AI's Stable-LM Alpha
• Technology Innovation Institute's Falcon
These open-source models work great for general use cases but lack context awareness for domain-specific tasks. For example, they need help writing an email for a patient or understanding non-popular programming languages.
How do you overcome this? You want to answer these two questions:
• How do I fine-tune an open-source model?
• How do I deploy the fine-tuned version of the model?
Fine-tuning an open-source model for many tasks will give you better results than using GPT-4 or Gemini out of the box. Deploying the model will give you immediate access to build applications using it.
I've worked a ton with @monsterapis. They recently launched their platform, where you can fine-tune and deploy your open-source models.
It takes one click, and you get the following:
• An optimized, high-throughput version of your model.
• A GPU infrastructure setup you don't need to worry about.
• An inexpensive API endpoint to access your model.
• An integrated platform to deploy your fine-tuned models.
I partnered with the team, and they gave me 10,000 free credits to give to anyone using the code "SANTIAGO" in the https://t.co/Ih2HxRfK8X dashboard. You can use these credits to access, fine-tune, and deploy these open-source models.
If you want to read their latest updates, get free credits and special offers, join their Discord server: https://t.co/oecoxEec14
Microsoft quietly built the best alternative to ChatGPT, and it's completely free.
It has all the premium ChatGPT features, including GPT-4, DALL-E 3, Code Interpreter, Multimodality, and SO much more.
Here’s the deep dive of everything you need to know:
Imagine this isometric miniature as a "crime scene".
You watch a 30-second animation where someone is shot or something is stolen, and you need to figure out what happened to "solve the riddle".
Keeping in mind that all the scenes, assets, and people figures would be AI-generated, the number of possible scenarios is unlimited.
I think an entire YouTube channel could be built just on these types of videos. How cool would that be?
Another idea - making this a looping visual (let's say 3 hours long), then generating Lofi music with AI, and launching a YouTube channel. People could listen to chill music, vibe along while admiring all the tiny details this kind of animation has.
The potential is vast
A bacterial flagellum is driven by a rotary engine made up of protein, and is powered by the flow of protons across the bacterial cell membrane. It rotates at ~1,000 rpm, but the rotor alone can rotate at up to 17,000 rpm
[📹 https://t.co/sT1spXjNgG]
I can't believe I've just fine-tuned a 33B-parameter LLM on Google Colab in a few hours.😱
Insane announcement for any of you using open-source LLMs on normal GPUs! 🤯
A new paper has been released, QLoRA, which is nothing short of game-changing for the ability to train and fine-tune LLMs on consumers' GPUs.
In a few words:
QLoRA reduces the memory usage of LLM fine-tuning without any performance tradeoffs compared to standard 16-bit model fine-tuning.
This method enables 33B model fine-tuning on a single 24GB GPU and 65B model fine-tuning on a single 46GB GPU. This is incredible! 😍
More specifically, QLoRA uses 4-bit quantization to compress a pre-trained language model. The LM parameters are then frozen, and a relatively small number of trainable parameters are added to the model in the form of Low-Rank Adapters.
During finetuning, QLoRA backpropagates gradients through the frozen 4-bit quantized pretrained language model into the Low-Rank Adapters. The LoRA layers are the only parameters being updated during training. Read more about LoRA in the original LoRA paper (https://t.co/54Lsb4DBFi). 🤓
QLoRA has one storage data type (usually 4-bit NormalFloat) for the base model weights and a computation data type (16-bit BrainFloat) used to perform computations. QLoRA dequantizes weights from the storage data type to the computation data type to perform the forward and backward passes, but only computes weight gradients for the LoRA parameters, which use 16-bit bfloat. The weights are decompressed only when they are needed, therefore the memory usage stays low during training and inference. Beautiful!😱
QLoRA tuning is shown to match 16-bit finetuning methods in a wide range of experiments. In addition, the Guanaco models, which use QLoRA finetuning for LLaMA models on the OpenAssistant dataset (OASST1), are state-of-the-art chatbot systems and are close to ChatGPT on the Vicuna benchmark. This is an additional demonstration of the power of QLoRA tuning.
Their Guanaco models are reaching 99.3% of the performance level of ChatGPT while only requiring 24 hours of fine-tuning on a single GPU. You can actually do it in Google Colab.
📚 Links-
QLoRA Paper - https://t.co/srUrDq4PS6
Colab for inference - https://t.co/iBtmLleCLf
Colab for fine-tuning - https://t.co/8eZQLxKozM
GitHub Repository- https://t.co/Jo50Ltn1Xw
Use it with HuggingFace - https://t.co/TWT0xPfr2g
Zoom in on the Andromeda Galaxy with Hubble. What looks like a smudge of light to the unaided eye is actually a vast galaxy containing an estimated 1 trillion stars.
Credit: ESA/Hubble & NASA
Full HD version on Youtube: https://t.co/uE6XQiyJ5e