For more on this, see my full breakdown: https://t.co/evSIfOXLqy
Other links:
Paper: https://t.co/rgJOHjnp3E
Code+Demo: https://t.co/Gv5E0WqgKf
Samples: https://t.co/04CDVakORV
Colab: https://t.co/xnCKJrXSph
Open-source is emerging as an even bigger threat to closed-source. Many of these closed-source models haven't even considered using LoRA fine-tuning, and instead prefer to train from scratch. The pace is real.
Producing highly capable, state of the art models no longer requires expensive compute for fine-tuning. You can do it with minimal commercial resources or on a RTX 3090 now. Everyone can be their own mad scientist.
The results are noteworthy: The 65B and 33B Guanaco variants consistently matched ChatGPT-3.5's performance. While the Vicuna benchmarking is imperfect (the researchers note that extensively), it's nonetheless significant and newsworthy.
They quantize the quantization constants. This is akin to compressing their compression formula as well. And memory spikes typical in fine-tuning are optimized, which reduces max memory load required.
A special 4-bit NormalFloat data type is efficient at being precise, versus the 16-bit floats and integers which are memory-intensive. Best way to think about this is that it's like compression (but not exactly the same).
A commercial GPU with 48GB of memory is now able to produce the same fine-tuned results as the same 16-bit tuning requiring 780GB of memory. This is a massive decrease in resources.
It's so efficient that researchers were able to fine-tune a 33B parameter model on a 24GB consumer GPU (RTX 3090, etc.) in 12 hours, which scored 97.8% in a benchmark against GPT-3.5.
QLoRA is an even more efficient way of fine-tuning which truly democratizes access to fine-tuning (no longer requiring expensive GPU power). It's a 4-bit approach that generates performance similar to 16-bit finetuning methods.
Fine-tuning an existing model is already a popular and cost-effective way to enhance an existing LLMs capabilities. It's cheaper than training from scratch, but it can still cost $$$.