What's is LLM fine-tuning?
Classical LLM fine-tuning is impractical (billions of parameters, hundreds of gigabytes).
Since such computing resources aren't available to everyone, parameter-efficient finetuning techniques (PEFT) have emerged.
Before diving into each technique, here's some background to help you better understand how they work:
LLM weights are matrices of numbers that are adjusted during fine-tuning.
Most PEFT techniques boil down to finding a low-rank adaptation of these matrices, i.e., a matrix of a smaller dimension that can still represent the information from the original one.
Now that you have a basic understanding of matrix rank, let's dive into the different fine-tuning techniques.
Top 5 Techniques
1) LoRA
Two trainable low-rank matrices A and B are added next to the weight matrices.
Instead of fine-tuning W, the updates to these low-rank matrices are adjusted.
Even for the largest LLMs, LoRA matrices only take up a few megabytes of memory.
2) LoRA-FA
While LoRA greatly reduces the number of trainable parameters, updating the low-rank weights still requires a significant amount of activation memory.
LoRA-FA (FA stands for Frozen-A) freezes matrix A and only updates matrix B.
3) VeRA
In LoRA, the low-rank matrices A and B are unique to each layer.
In VeRA, A and B are frozen, random, and shared across all layers.
Instead, layer-specific scaling VECTORS b and d are trained.
4) Delta-LoRA
Here, too, the W matrix is adjusted, but not in the classical way.
The difference (delta) between the product of matrices A and B at two consecutive training steps is added to W.
5) LoRA+
In regular LoRA, both matrices A and B are updated with the same learning rate.
The authors of LoRA+ found that a higher learning rate for matrix B leads to better convergence
••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
