LoRA Fine-Tuning: Counting Trainable Parameters
Fine-tuning a large model need not update every weight. LoRA restricts the update of a linear layer to the product of two small matrices. The base weights stay frozen while the increment is learned. This reduces trainable parameters; it does not shrink every parameter in the original model.
Contents
How two matrices change the output
Let W₀ have shape d×k and x contain k components. Write the update as sBA, with A of shape r×k, B of shape d×r and a scale s commonly set to α/r. The forward expression is y = W₀x + sB(Ax). The chosen r bounds the update rank; the actual product can have a lower rank.
For an invented 2×3 layer, take A=[1,2,0], B=[0.5,−1]ᵀ, x=[2,3,4] and s=0.25. Then Ax=8 and B(Ax)=[4,−8]. After scaling, add this to the base output [2,3] to obtain [3,1].
W₀x = [2, 3]
sB(Ax) = [1, −2]
y = [3, 1]
The update is not a fixed output bias. Changing the input changes the increment produced by A and B.
Count parameters per layer
For one linear layer without bias, the full weight has dk parameters and the LoRA path has r(k+d) trainable parameters. With d=k=4096 and r=8, these counts are 16,777,216 and 65,536, a ratio of 256 to one. This is shape arithmetic, not a whole-model VRAM saving.
| Object | Count | Still required? |
|---|---|---|
| Base weight W₀ | dk | Stored and used in the forward pass |
| A and B | r(k+d) | Trainable update parameters |
| Activations and temporary tensors | Network-dependent | Not determined by rank alone |
Do not estimate whether training fits from adapter file size alone. Record base-weight precision, target layers, sequence length, batch size and peak memory separately, then evaluate quality on the task’s data.
Verify weight merging directly
For this ordinary linear layer, precomputing W=W₀+sBA makes Wx equivalent to the two-path calculation. The following Python example was executed to check both results. It invokes no training framework and measures no GPU performance.
def mv(m, x):
return [sum(a*b for a,b in zip(row,x)) for row in m]
W = [[1,0,0], [0,1,0]]
A = [[1,2,0]]
B = [[0.5], [-1]]
x = [2,3,4]
s = 0.25 # alpha / r in this toy example
base, update = mv(W,x), mv(B,mv(A,x))
y = [a+s*b for a,b in zip(base,update)]
merged = [[W[i][j]+s*B[i][0]*A[0][j]
for j in range(3)] for i in range(2)]
assert y == [3.0,1.0]
assert mv(merged,x) == y
print(y)
Floating-point implementations generally need a comparison tolerance. This example uses exactly representable values, so direct equality is sufficient. Quantization or different precisions may alter results; a real-number identity does not promise bitwise equivalence in every implementation.
Choose rank by validation
A smaller rank restricts the trainable space. Whether it is sufficient is a validation question. Keep the data split and target layers fixed, compare several ranks, and record cost and error types. Training loss alone is not a sound selection criterion.
Microsoft’s implementation documents training LoRA parameters and saving adapters, with weight merging in the relevant modes. Check the deployed library’s behavior to avoid applying an already merged update twice. This article covers ordinary LoRA algebra rather than every variant.
The original work was submitted to arXiv in 2021. This explanation uses its principle and an original small example, without reproducing paper benchmarks or claiming low-rank updates universally beat full fine-tuning. When missing evidence is the problem, also compare RAG retrieval: learning an update and retrieving source material address different needs.
Sources
Original LoRA preprint (2021); authors’ implementation. No benchmark speedup is reported here.
Publication note: actually published as a catch-up on October 10, 2026 (Beijing time), retaining the originally planned article date.
Support
If this article helped you, you can support this site.
Click a QR code to enlarge. More options: support page。


