The low-rank decomposition, precisely
For a pretrained weight , full fine-tuning would learn an unconstrained update — as many free parameters as itself. LoRA constrains it to rank :
Layer 4 · Foundations
The math of low-rank decomposition and its gradient scaling, QLoRA's NF4 quantization, and why fine-tuning hardware requirements differ so sharply from pretraining's.
For a pretrained weight , full fine-tuning would learn an unconstrained update — as many free parameters as itself. LoRA constrains it to rank :