TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 583 subscribers
Post #1898 189
LLM FINE-TUNING

What's is LLM fine-tuning?

Classical LLM fine-tuning is impractical (billions of parameters, hundreds of gigabytes).

Since such computing resources aren't available to everyone, parameter-efficient finetuning techniques (PEFT) have emerged.

Before diving into each technique, here's some background to help you better understand how they work:

LLM weights are matrices of numbers that are adjusted during fine-tuning.

Most PEFT techniques boil down to finding a low-rank adaptation of these matrices, i.e., a matrix of a smaller dimension that can still represent the information from the original one.

Now that you have a basic understanding of matrix rank, let's dive into the different fine-tuning techniques.


Top 5 Techniques
1) LoRA

Two trainable low-rank matrices A and B are added next to the weight matrices.

Instead of fine-tuning W, the updates to these low-rank matrices are adjusted.

Even for the largest LLMs, LoRA matrices only take up a few megabytes of memory.

2) LoRA-FA

While LoRA greatly reduces the number of trainable parameters, updating the low-rank weights still requires a significant amount of activation memory.

LoRA-FA (FA stands for Frozen-A) freezes matrix A and only updates matrix B.

3) VeRA

In LoRA, the low-rank matrices A and B are unique to each layer.

In VeRA, A and B are frozen, random, and shared across all layers.

Instead, layer-specific scaling VECTORS b and d are trained.

4) Delta-LoRA

Here, too, the W matrix is adjusted, but not in the classical way.

The difference (delta) between the product of matrices A and B at two consecutive training steps is added to W.

5) LoRA+

In regular LoRA, both matrices A and B are updated with the same learning rate.

The authors of LoRA+ found that a higher learning rate for matrix B leads to better convergence


••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
More from @dataxplore
  1. Oct 1, 2026Monitoring and debugging "silent" token drift in LLM pipelines: frequency analysis of embe…
  2. Sep 30, 2026Online detection of feature collisions in TDA transformation When using Topological Data A…
  3. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  4. Sep 14, 2026Post #2188
  5. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  6. Aug 22, 2026Post #2185
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →