LoRA Fine-Tuning: Counting Trainable Parameters

黎 浩然/ 8 10 月, 2026/ 大语言模型/LARGELANGUAGEMODEL/LLM, 机器学习/MACHINELEARNING, 研究生/POSTGRADUATE, 计算机/COMPUTER/ 0 comments

Fine-tuning a large model need not update every weight. LoRA restricts the update of a linear layer to the product of two small matrices. The base weights stay frozen while the increment is learned. This reduces trainable parameters; it does not shrink every parameter in the original model.

Contents
  1. How two matrices change the output
  2. Count parameters per layer
  3. Verify weight merging directly
  4. Choose rank by validation
  5. Sources

How two matrices change the output

Let W₀ have shape d×k and x contain k components. Write the update as sBA, with A of shape r×k, B of shape d×r and a scale s commonly set to α/r. The forward expression is y = W₀x + sB(Ax). The chosen r bounds the update rank; the actual product can have a lower rank.

For an invented 2×3 layer, take A=[1,2,0], B=[0.5,−1]ᵀ, x=[2,3,4] and s=0.25. Then Ax=8 and B(Ax)=[4,−8]. After scaling, add this to the base output [2,3] to obtain [3,1].

Base path
W₀x = [2, 3]
Update path
sB(Ax) = [1, −2]
Elementwise sum
y = [3, 1]
Original two-path diagram. Values are constructed for teaching, not learned in training.

The update is not a fixed output bias. Changing the input changes the increment produced by A and B.

Count parameters per layer

For one linear layer without bias, the full weight has dk parameters and the LoRA path has r(k+d) trainable parameters. With d=k=4096 and r=8, these counts are 16,777,216 and 65,536, a ratio of 256 to one. This is shape arithmetic, not a whole-model VRAM saving.

Object Count Still required?
Base weight W₀ dk Stored and used in the forward pass
A and B r(k+d) Trainable update parameters
Activations and temporary tensors Network-dependent Not determined by rank alone

Do not estimate whether training fits from adapter file size alone. Record base-weight precision, target layers, sequence length, batch size and peak memory separately, then evaluate quality on the task’s data.

Verify weight merging directly

For this ordinary linear layer, precomputing W=W₀+sBA makes Wx equivalent to the two-path calculation. The following Python example was executed to check both results. It invokes no training framework and measures no GPU performance.

def mv(m, x):
    return [sum(a*b for a,b in zip(row,x)) for row in m]

W = [[1,0,0], [0,1,0]]
A = [[1,2,0]]
B = [[0.5], [-1]]
x = [2,3,4]
s = 0.25  # alpha / r in this toy example
base, update = mv(W,x), mv(B,mv(A,x))
y = [a+s*b for a,b in zip(base,update)]
merged = [[W[i][j]+s*B[i][0]*A[0][j]
           for j in range(3)] for i in range(2)]
assert y == [3.0,1.0]
assert mv(merged,x) == y
print(y)

Floating-point implementations generally need a comparison tolerance. This example uses exactly representable values, so direct equality is sufficient. Quantization or different precisions may alter results; a real-number identity does not promise bitwise equivalence in every implementation.

Choose rank by validation

A smaller rank restricts the trainable space. Whether it is sufficient is a validation question. Keep the data split and target layers fixed, compare several ranks, and record cost and error types. Training loss alone is not a sound selection criterion.

Microsoft’s implementation documents training LoRA parameters and saving adapters, with weight merging in the relevant modes. Check the deployed library’s behavior to avoid applying an already merged update twice. This article covers ordinary LoRA algebra rather than every variant.

The original work was submitted to arXiv in 2021. This explanation uses its principle and an original small example, without reproducing paper benchmarks or claiming low-rank updates universally beat full fine-tuning. When missing evidence is the problem, also compare RAG retrieval: learning an update and retrieving source material address different needs.

中文版

Sources

Original LoRA preprint (2021); authors’ implementation. No benchmark speedup is reported here.

Publication note: actually published as a catch-up on October 10, 2026 (Beijing time), retaining the originally planned article date.

Support

If this article helped you, you can support this site.

WeChat support QR code; click to enlarge
WeChat
Alipay support QR code; click to enlarge
Alipay
Buy Me a Coffee; support this site
Buy Me a Coffee

Click a QR code to enlarge. More options: support page。

Share this Post

Leave a Comment

您的邮箱地址不会被公开。 必填项已用 * 标注

*
*