Objective: Understand 4-bit quantization (NF4) and the QLoRA approach, which combines quantization + LoRA to fine-tune models up to 7B params on a consumer GPU.
Plan: 1. The memory problem of large models 2. Quantization: from FP16 to NF4 3. BitsAndBytesConfig in practice 4. The LoRA mechanics that the API wraps 5. QLoRA: quantized base + LoRA adapters 6. QLoRA training on OPT-1.3B 7. Generation with the QLoRA model 8. LoRA vs QLoRA comparison 9. GPU memory cleanup 10. Exercises
Prerequisites: FT-01 (Introduction to Fine-Tuning) — LoRA, rank, adapters.
Reading the committed output: the next cell prints the full environment (PyTorch version, GPU, VRAM) – its committed output is the single source of truth, the prose will no longer re-pin it (#9434). Here, a cu12x environment on an RTX 3080 Ti Laptop, ~17 GB of total VRAM of which ~16 GB free at startup. This notebook’s pedagogical paradox: this card would hold OPT-1.3B in FP16 without any difficulty (4.03 GB measured further down, section 8). It is not this card that justifies QLoRA — it is generalization: the same recipe that holds a 1.3B model in ~1.7 GB will hold a 7B model in ~6 GB, on a 300-euro card. The notebook walks through the full chain: NF4 quantization (section 2), hand-written LoRA mechanics (section 4), LoRA grafting via the API (section 5), real measured training (section 6), before/after generation (section 7), and the FP16 vs 4-bit memory comparison (section 8).
import gcimport osimport timeimport warningsimport torchwarnings.filterwarnings("ignore", message=r"_check_is_size will be removed.*", category=FutureWarning,)warnings.filterwarnings("ignore", message=r"torch\.utils\.checkpoint: the use_reentrant parameter.*", category=UserWarning,)print(f"PyTorch {torch.__version__}")print(f"CUDA : {torch.cuda.is_available()}")if torch.cuda.is_available():print(f"GPU : {torch.cuda.get_device_name(0)}") vram_total = torch.cuda.get_device_properties(0).total_memory /1e9 vram_free = torch.cuda.mem_get_info()[0] /1e9print(f"VRAM totale : {vram_total:.1f} Go | Libre : {vram_free:.1f} Go")
PyTorch 2.11.0+cu126
CUDA : True
GPU : NVIDIA GeForce RTX 4090
VRAM totale : 25.8 Go | Libre : 24.2 Go
1. The memory problem of large models
Fine-tuning a language model requires storing in VRAM:
Component
Approximate size
Model weights (FP16)
2 bytes/param
Gradients
2 bytes/param
Optimizer states (AdamW)
8 bytes/param (m + v)
Activations (forward pass)
Variable, often 1-2x model
Total for full fine-tuning: ~12-16 bytes per parameter.
Model
Params
Full FT (FP16)
LoRA (FP16)
QLoRA (4-bit)
GPT-2
124M
~2 GB
~0.5 GB
~0.3 GB
OPT-1.3B
1.3B
~20 GB
~5 GB
~2 GB
Llama-7B
7B
~112 GB
~28 GB
~6 GB
LoRA reduces gradients and the optimizer (only the adapters are trained), but the model weights remain in FP16.
QLoRA goes further: it quantizes the model weights to 4-bit while keeping computations at higher precision.
The real numbers read in section 8: loading OPT-1.3B in FP16 exceeds the weights-only theory, because CUDA buffers and residual context also count. But loading is not fine-tuning: a full fine-tuning additionally requires gradients and Adam optimizer states. This triple tax is what LoRA (gradients on adapters only) then QLoRA (4-bit base weights) reduce. The observed values stay in the outputs so that the prose does not freeze a measurement specific to one machine.
2. Quantization: from FP16 to NF4
2a. Principle of quantization
Quantization reduces the precision of weights to decrease the memory footprint:
Neural network weights follow an approximately normal distribution (centered on 0). A linear INT4 encoding wastes possible values:
INT4: values uniformly distributed over [-8, 7] → dense regions of the distribution poorly represented
NF4 (NormalFloat 4): values optimized for a normal distribution → each quantile represents an equal fraction of the distribution
NF4 = the theoretically optimal format for neural network weights (Dettmers et al., 2023).
2c. Double quantization
QLoRA also applies double quantization: the quantization constants (scaling factors) are themselves quantized in FP8, saving an additional ~0.37 bits/param.
Why NF4 rather than INT4? The weights of a trained network are distributed according to an approximately Gaussian law, centered with thin tails – whereas INT4 slices that interval into 16 uniformly spaced levels. The result: the bulk of the weights, concentrated around zero, piles up on 4-5 levels while the extreme levels (extremely rare) waste their share of the range. NF4 (NormalFloat4) instead builds its 16 levels on the quantiles of the normal law: each level carries roughly the same mass of weights, information is spread where it is dense. It is this choice of grid – and not the number of bits – that explains the “Quality ~97-99% of full fine-tuning” line of the final table. The double quantization added here (bnb_4bit_use_double_quant=True) compresses the quantization constants themselves into 8 bits: ~0.37 bit saved per parameter, measurable in exercise 3.
3. BitsAndBytesConfig in practice
The bitsandbytes library integrates quantization directly into model loading via transformers.
Reading the committed config: the output shows the three canonical ingredients of the QLoRA paper (Dettmers et al., 2023) – 4-bit NF4 for the base weights, compute dtype float16 for the computations (weights are dequantized on the fly into FP16 at each forward pass), and double quant : True for the constants. The pair is complete: quantization for memory residency, LoRA for training. Compare with FT-01: the only configuration difference lies in this BitsAndBytesConfig passed to from_pretrained – the rest of the LoRA mechanics (rank, alpha, targeted modules) is identical.
from transformers import BitsAndBytesConfig# Configuration QLoRA : quantization 4-bit NF4bnb_config = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_quant_type="nf4", # NormalFloat 4 (optimal pour poids) bnb_4bit_compute_dtype=torch.float16, # Calculs en FP16 pour stabilite numerique bnb_4bit_use_double_quant=True, # Double quantization des scaling factors)print("BitsAndBytesConfig QLoRA :")print(f" Quantization : 4-bit NF4")print(f" Compute dtype : float16")print(f" Double quant : True")
Reading the result: the three lines that decide memory
The previous cell built the quantization configuration in three lines, and each plays an orthogonal role in the VRAM budget. They are worth dissecting, because this exact triplet decides whether or not your fine-tuning will fit in your GPU’s memory:
load_in_4bit=True: the cardinal instruction. At load time, transformers delegates to bitsandbytes the conversion of each FP16 weight (2 bytes) into 4-bit NormalFloat (~0.5 byte). The raw gain is ~4x on the weight footprint – for OPT-1.3B, that takes ~2.6 GB down to ~0.65 GB. This is the only moment the conversion happens: during forward passes, weights are dequantized on the fly into FP16, just long enough for the computation, then put back as NF4 in the CUDA caches.
bnb_4bit_quant_type="nf4": NF4 = NormalFloat 4-bit. The choice is not arbitrary. A neural network’s weights follow, to first approximation, an approximately centered Gaussian distribution (law of large numbers applied to sums of gradients). NF4 reserves its 16 quantization levels in the dense zones of that distribution – which is why Dettmers et al. (2023, the founding QLoRA paper) preferred it over uniformly spaced INT4. For ~97-99% of full fine-tuning quality on 7B parameters, NF4 beats INT4 by about 1 point of accuracy on MMLU – not negligible on a budget that only tolerates 1-3% degradation.
bnb_4bit_compute_dtype=torch.float16: the type used for forward computations (and by extension, the gradients on the LoRA adapters). Choosing FP16 over FP32 keeps activations in 16 bits, which saves cache memory without visibly degrading convergence for OPT-1.3B – for 7B+ models or long sequences, one often switches to BF16 (same dynamic range as FP32, reduced precision) to avoid underflows.
bnb_4bit_use_double_quant=True (omitted here, now defaults to True): double quantization applies a second NF4->FP8 pass on the quantization constants themselves (~0.37 bit/parameter). For a 7B, that is ~300 MB more recovered – negligible for OPT-1.3B, but critical for Llama-2-7B where the saving keeps you under 24 GB of VRAM.
The intuition to keep: 4-bit quantization shrinks the weight footprint but imposes dynamic dequantization at every forward. The only work that actually runs in FP16 or FP32 is the gradients on the LoRA adapters – the Adam optimizer never touches the quantized weights. This separation of memory spaces (weights frozen in NF4, gradients in FP16) is what makes fine-tuning possible on consumer GPUs. Without LoRA, 4-bit quantization alone would not suffice: you would have to backpropagate through the dequantized weights, which costs as much memory as an FP16 fine-tuning.
The next cell (measure_vram) measures precisely this footprint: torch.cuda.memory_allocated() / 1e9 captures the actual CUDA allocation at the loading peak, before PyTorch’s caches release anything. It is this number, not the theoretical size (1.3B x 0.5 byte = 0.65 GB), that the next cell reads – it includes attention buffers, temporary activation tensors, and CUDA overhead.
# Mesurer la VRAM avant chargementif torch.cuda.is_available(): torch.cuda.reset_peak_memory_stats() vram_before = torch.cuda.memory_allocated() /1e9else: vram_before =0print(f"VRAM avant chargement : {vram_before:.2f} Go")
VRAM avant chargement : 0.00 Go
4. The LoRA mechanics the API wraps
The next section will call get_peft_model + LoraConfig on the quantized model – two lines of API. Before running them, this section shows what those two lines wrap, on toy tensors (512x512, CPU, seconds): a hand-written LoRALinear layer, the two invariants that carry all of LoRA’s interest, and the rank x alpha grid that turns the r=16, α=32 recipe into a decision.
Recall of the principle (Hu et al., 2021): instead of learning a dense update ΔW of size d x d, LoRA constrains it to low rank: ΔW = B·A with A : d→r, B : r→d. The pre-trained W₀ is frozen; only A and B are learned – r times fewer parameters per layer.
# LoRALinear from scratch : W0 gele, delta bas-rang apprisimport torchimport torch.nn as nntorch.manual_seed(0)class LoRALinear(nn.Module):"""Couche lineaire LoRA : y = W0 x + (alpha/r) * B(A(x)). W0 est GELÉ (requires_grad=False) : le pré-entraîné n'est jamais modifié. A (d_in -> r) : init Kaiming ; B (r -> d_out) : init ZÉRO. Le facteur alpha/r contrôle l'amplitude effective de la mise à jour. """def__init__(self, base_layer: nn.Linear, r: int=8, lora_alpha: int=16):super().__init__()self.base_layer = base_layerfor p inself.base_layer.parameters(): p.requires_grad =False# <-- le gel du pré-entraînéself.r = rself.scaling = lora_alpha / r d_out, d_in = base_layer.weight.shapeself.lora_A = nn.Parameter(torch.zeros(r, d_in))self.lora_B = nn.Parameter(torch.zeros(d_out, r)) nn.init.kaiming_uniform_(self.lora_A, a=5**0.5)def delta_W(self) -> torch.Tensor:"""La mise a jour effectivement apprise : B @ A, de rang <= r."""returnself.lora_B @self.lora_Adef forward(self, x):returnself.base_layer(x) +self.scaling * (x @self.lora_A.T @self.lora_B.T)# Demonstration sur un tenseur jouet 512x512 (CPU, sans modele HF)base = nn.Linear(512, 512, bias=False)torch.manual_seed(42)with torch.no_grad(): base.weight.copy_(torch.randn(512, 512) *0.02)lora_demo = LoRALinear(base, r=8, lora_alpha=16)n_lora = lora_demo.lora_A.numel() + lora_demo.lora_B.numel()print(f"W0 : {tuple(base.weight.shape)} = {base.weight.numel():,} params, requires_grad={base.weight.requires_grad}")print(f"A, B : {tuple(lora_demo.lora_A.shape)} + {tuple(lora_demo.lora_B.shape)} = {n_lora:,} params entrainables")print(f"ratio : {n_lora / base.weight.numel():.2%} de la couche de base (r=8)")print(f"scaling : alpha/r = 16/8 = {lora_demo.scaling}")
W0 : (512, 512) = 262,144 params, requires_grad=False
A, B : (8, 512) + (512, 8) = 8,192 params entrainables
ratio : 3.12% de la couche de base (r=8)
scaling : alpha/r = 16/8 = 2.0
# Preuve 1 -- l'initialisation demarre EXACTEMENT comme le pre-entrainedW = lora_demo.delta_W().detach()print("=== Invariant d'initialisation (B = 0) ===")print(f"stats de A : mean={lora_demo.lora_A.mean():+.5f} std={lora_demo.lora_A.std():.5f} (Kaiming, non nul)")print(f"stats de B : mean={lora_demo.lora_B.mean():+.5f} std={lora_demo.lora_B.std():.5f} (ZÉRO)")print(f"delta_W = B@A : max|delta_W| = {dW.abs().max():.2e}")print(f"torch.allclose(delta_W, 0) : {torch.allclose(dW, torch.zeros_like(dW))}")x_probe = torch.randn(4, 512)with torch.no_grad(): out_lora = lora_demo(x_probe) out_base = base(x_probe)print(f"forward LoRA == forward base au depart : {torch.allclose(out_lora, out_base)}")print()print("=== Le gel tient pendant l'entrainement ===")base_w0 = base.weight.clone()opt = torch.optim.SGD([lora_demo.lora_A, lora_demo.lora_B], lr=0.1)x_batch, target = torch.randn(16, 512), torch.randn(16, 512)loss0 =Nonefor step inrange(50): opt.zero_grad() loss = ((lora_demo(x_batch) - target) **2).mean()if loss0 isNone: loss0 = loss.item() loss.backward() opt.step()dW_after = lora_demo.delta_W().detach()print(f"apres 50 pas SGD : |delta_W| = {dW_after.norm():.4f} (grandit) | W0 inchangé : {torch.equal(base.weight, base_w0)}")print(f"loss : {loss0:.4f} -> {loss.item():.4f}")
=== Invariant d'initialisation (B = 0) ===
stats de A : mean=+0.00042 std=0.02548 (Kaiming, non nul)
stats de B : mean=+0.00000 std=0.00000 (ZÉRO)
delta_W = B@A : max|delta_W| = 0.00e+00
torch.allclose(delta_W, 0) : True
forward LoRA == forward base au depart : True
=== Le gel tient pendant l'entrainement ===
apres 50 pas SGD : |delta_W| = 0.3050 (grandit) | W0 inchangé : True
loss : 1.1991 -> 1.1271
Reading: the exact start
B initialized to zero ⇒ ΔW = B@A = 0exactly (not approximately): the cell asserts allclose(delta_W, 0) → True and forward LoRA == forward base → True. This is the property that makes LoRA safe: the LoRA-fied model starts its life as the pre-trained one, then departs from it only through what training learns in B (whose gradient is non-zero from the very first step, even if B equals zero – the derivative does not depend on the value of B).
And the freeze holds: after 50 SGD steps, W₀ is bit-identical to itself (torch.equal → True) while |ΔW| grows from 0 to 0.305. Gradients flow only into A and B.
# Preuve 2 -- merge : inference "zero-overhead" (l'argument central de FT-05)# On simule un adaptateur entraîné : B non nul.with torch.no_grad(): lora_demo.lora_B.normal_(0, 0.01)# Chemin 1 : reseau LoRA non mergé (deux matrices separees)x = torch.randn(64, 512)with torch.no_grad(): path_unmerged = lora_demo(x)# Chemin 2 : poids mergés W = W0 + (alpha/r) * B@A, UNE seule couche denseW_merged = base.weight + lora_demo.scaling * lora_demo.delta_W()merged = nn.Linear(512, 512, bias=False)with torch.no_grad(): merged.weight.copy_(W_merged)with torch.no_grad(): path_merged = merged(x)ecart_max = (path_unmerged - path_merged).abs().max().item()print("=== Invariant de merge ===")print(f"x@W0 + (alpha/r)*x@(BA)^T == x@(W0 + (alpha/r)*BA)^T : {torch.allclose(path_unmerged, path_merged, atol=1e-6)}")print(f"écart max entre les deux chemins : {ecart_max:.2e} (arrondi float32)")print(f"rang de delta_W : {torch.linalg.matrix_rank(lora_demo.delta_W().detach()).item()} <= r={lora_demo.r}")
=== Invariant de merge ===
x@W0 + (alpha/r)*x@(BA)^T == x@(W0 + (alpha/r)*BA)^T : True
écart max entre les deux chemins : 1.28e-06 (arrondi float32)
rang de delta_W : 8 <= r=8
Reading: “zero-overhead” inference
The two computation paths coincide (allclose → True, max gap 1.28e-06, float32 rounding): adding the update into the weights (W₀ + (α/r)·BA) produces the same network as keeping two separate branches. This is THE proof that merging at inference is free: after fusion, no more multiplication by B then A – a single dense layer, the same as the pre-trained one, with different weights. This is the central argument that the FT-05 (ModelMerging) notebook exploits downstream: merging n adapters amounts to adding their deltas.
The rank of ΔW (displayed: ≤ r) is the counterpart: the update lives in a subspace of dimension r. The whole stake of choosing r lies there – enough dimensions for the task, no more.
# Grille rank x alpha : combien de paramètres entraînables, et quelle amplitude# Dimensions réelles d'OPT-1.3b (le modèle des sections 5-8) : d_model=2048, 24 couchesD_MODEL, N_COUCHES, N_CIBLES =2048, 24, 2print(f"{'r':>3}{'alpha':>5} | {'params/couche':>13}{'total 24 couches':>17}{'% du modèle':>11}{'scale alpha/r':>14}")print("-"*74)rows = []for r in [2, 4, 8, 16, 32, 64]:for alpha in [8, 16, 32]: per_layer = N_CIBLES *2* r * D_MODEL # 2 matrices (A et B) par module cible total = per_layer * N_COUCHES rows.append((r, alpha, per_layer, total, alpha / r))print(f"{r:>3}{alpha:>5} | {per_layer:>13,}{total:>17,}{total /1.3e9*100:>10.3f}% {alpha / r:>14.3f}")print()print("Lecture : le nombre de paramètres ne dépend QUE de r -- alpha n'achète")print("aucun paramètre, il ne fait que rescaler la mise à jour (colonne scale).")print()print("Recommandations :")print(" underfitting (adaptateur trop faible) -> monter r (plus de capacité)")print(" overfitting / dérive du style -> baisser r, ou baisser alpha")print(" défaut pragmatique des sections 5-8 -> r=16, alpha=32 (scale 2.0)")r_def, a_def =16, 32total_def =next(t for rr, aa, _, t, _ in rows if rr == r_def and aa == a_def)print(f" -> le r=16/alpha=32 de la section suivante : {total_def:,} params entraînables.")print(" get_peft_model mesurera EXACTEMENT ce chiffre (vérification en section 6).")
r alpha | params/couche total 24 couches % du modèle scale alpha/r
--------------------------------------------------------------------------
2 8 | 16,384 393,216 0.030% 4.000
2 16 | 16,384 393,216 0.030% 8.000
2 32 | 16,384 393,216 0.030% 16.000
4 8 | 32,768 786,432 0.060% 2.000
4 16 | 32,768 786,432 0.060% 4.000
4 32 | 32,768 786,432 0.060% 8.000
8 8 | 65,536 1,572,864 0.121% 1.000
8 16 | 65,536 1,572,864 0.121% 2.000
8 32 | 65,536 1,572,864 0.121% 4.000
16 8 | 131,072 3,145,728 0.242% 0.500
16 16 | 131,072 3,145,728 0.242% 1.000
16 32 | 131,072 3,145,728 0.242% 2.000
32 8 | 262,144 6,291,456 0.484% 0.250
32 16 | 262,144 6,291,456 0.484% 0.500
32 32 | 262,144 6,291,456 0.484% 1.000
64 8 | 524,288 12,582,912 0.968% 0.125
64 16 | 524,288 12,582,912 0.968% 0.250
64 32 | 524,288 12,582,912 0.968% 0.500
Lecture : le nombre de paramètres ne dépend QUE de r -- alpha n'achète
aucun paramètre, il ne fait que rescaler la mise à jour (colonne scale).
Recommandations :
underfitting (adaptateur trop faible) -> monter r (plus de capacité)
overfitting / dérive du style -> baisser r, ou baisser alpha
défaut pragmatique des sections 5-8 -> r=16, alpha=32 (scale 2.0)
-> le r=16/alpha=32 de la section suivante : 3,145,728 params entraînables.
get_peft_model mesurera EXACTEMENT ce chiffre (vérification en section 6).
Reading: the rank x alpha grid
Three readings:
The parameter count depends ONLY on r: for each r, the three values of alpha display the same total (2 · r · d per target module, over the two targeted projections and the 24 layers). alpha buys no parameters – it only rescales the update (scale = α/r, last column). Raising alpha at fixed r speeds up the amplitude of the same adapter, it does not enlarge it.
The orders of magnitude: from r=2 (393,216 params, 0.030% of the model) to r=64 (12.6M, 0.968%) – a factor 32 between the extremes, and always under 1% of the model. The grid is computed on OPT-1.3b’s actual dimensions (d_model 2048, 24 layers, targets q_proj+v_proj): the r=16 of the next section predicts 3,145,728 parameters there – exactly the figure get_peft_model measures. The mechanics’ theory and the API agree to the parameter.
The steering rule: underfitting → raise r; overfitting or stylistic drift → lower r or lower alpha. On a dataset of a few hundred examples (ours has 10), r=8-16 is the reasonable plateau.
5. QLoRA: quantized base + LoRA adapters
Let’s load OPT-1.3B in 4-bit and compare it with standard FP16 loading.
How to read the committed output: the 4-bit loading of OPT-1.3B occupies 1.25 GB of measured VRAM, against 0.36 GB theoretical for the weights alone – the gap (residual activations, dequantization buffers, embedding tables sometimes kept in 16 bits) is the price of real scaling. Note also the counting artefact: the cell displays “711,778,304 params (0.71B)” whereas the canonical count for OPT-1.3B is 1,318,903,808 (1.32B, confirmed by PEFT in the next cell). Summing numel() over an already quantized model does not return the original dense sizes – a Params4bit does not count like an FP16 tensor. Keep PEFT’s number: it is the honest denominator, the one the following percentages must be read against.
from transformers import AutoModelForCausalLM, AutoTokenizerfrom peft import LoraConfig, get_peft_model, prepare_model_for_kbit_trainingMODEL_NAME ="facebook/opt-1.3b"# 1.3B params# Charger en 4-bit (QLoRA)model_4bit = AutoModelForCausalLM.from_pretrained( MODEL_NAME, quantization_config=bnb_config, device_map="auto",)# Preparer le modele pour l'entrainement en mode quantizemodel_4bit = prepare_model_for_kbit_training(model_4bit)# Statistiques memoireif torch.cuda.is_available(): vram_4bit = torch.cuda.max_memory_allocated() /1e9else: vram_4bit =0total_params =sum(p.numel() for p in model_4bit.parameters())print(f"Modele : {MODEL_NAME}")print(f"Params totaux : {total_params:,} ({total_params/1e9:.2f}B)")print(f"VRAM apres chargement 4-bit : {vram_4bit:.2f} Go")print(f"Theorie 4-bit : {total_params *0.5/1e9:.2f} Go (poids seuls)")
Reading the result: the moment quantization takes effect
The previous cell has three steps that look trivial, but each marks a point of no return on the memory footprint. Read them in the order they execute:
prepare_model_for_kbit_training(model_4bit): a step often forgotten but essential. The model loaded in 4-bit still has its LayerNorm layers and its output head (lm_head) in higher precision – bitsandbytes leaves them untouched because NF4 quantization targets matrix weights, not normalizations. The routine freezes the base model, prepares the inputs for the gradient flow up to the adapters, and stabilizes the sensitive layers.
get_peft_model(model_4bit, lora_config): this is where the LoRA adapters are grafted. The method walks the model to find target_modules=["q_proj", "v_proj"]. On OPT-1.3B, each projection has dimension 2048x2048: with r=16, the A and B matrices add 2048×16 + 16×2048 = 65,536 parameters per target. Over 24 layers and two targets, that gives 3,145,728 trainable parameters, about 0.24% of the canonical model. The NF4 base weights stay frozen.
model_qlora.print_trainable_parameters(): this line validates the exact ratio. If the figure strays far from the previous grid’s prediction, check the targeted modules before training.
The intuition to keep: going from full fine-tuning to LoRA lets you use a higher learning rate without directly modifying the pre-trained weights. You optimize a small number of directions in a large space.
The trap to avoid: believing prepare_model_for_kbit_training is optional. A flat loss after several steps should trigger a check of the freezing, of the gradient flow, and of the adapter configuration.
# Ajouter les adaptateurs LoRA sur le modele quantizelora_config = LoraConfig( r=16, lora_alpha=32, target_modules=["q_proj", "v_proj"], # Attention query/value dans OPT lora_dropout=0.05, bias="none", task_type="CAUSAL_LM",)model_qlora = get_peft_model(model_4bit, lora_config)model_qlora.print_trainable_parameters()trainable =sum(p.numel() for p in model_qlora.parameters() if p.requires_grad)total =sum(p.numel() for p in model_qlora.parameters())print(f"\nParams entrainables : {trainable:,} ({trainable/total*100:.2f}%)")
The OPT-1.3B model is now loaded in 4-bit with the LoRA adapters configured (r=16, target q_proj/v_proj). Let’s load the associated tokenizer to prepare the training dataset.
The exact breakdown of the count: 3,145,728 trainable parameters – where does that number come from? OPT-1.3B has 24 decoder layers, each carrying two targeted projections (q_proj and v_proj, dimensions 2048x2048). Each rank-r LoRA adapter on a 2048->2048 layer adds r x (2048+2048) = 4096r parameters. Total: the 24 decoder layers, with two targeted projections each, give 24 x 2 x 4096 x r = 196,608 x r. For r = 16: exactly 3,145,728 – the committed output verifies the formula to the parameter. Against the canonical denominator of 1.32B (PEFT), that is 0.24% of the model trainable, a ratio of 419 to 1: 99.76% of the weights stay frozen in NF4. The “0.44%” displayed by the cell divides by the undercounted denominator of the previous cell (711M) – read PEFT’s 0.2385%.
Let’s fine-tune OPT-1.3B on French literary text, the same type of dataset as FT-01.
What the committed output measures: the final loss, the run duration and the VRAM peak are printed directly by the training cell. The peak adds only a fraction of the load-only footprint: gradients and optimizer states apply to the LoRA parameters, not to the billion frozen weights. The paged_adamw_8bit optimizer, visible in the arguments, completes the setup by storing its states in 8 bits. On this tiny corpus, the loss describes a stylistic imprint, not a mastery of the language; the generation below lets you read this trade-off.
from datasets import Dataset# Extrait elargi — texte francais classique (domaine public)training_texts = ["Cosette regardait la chambre obscure, et la figure ridee de son bienfaiteur. ""Cette figure ridee etait l'immense figure de la misere humaine. Jean Valjean ""avait pris l'habitude de descendre a la cave pour rapporter du vin.","La nuit, quand tout dormait dans la maison, Jean Valjean restait eveille. ""Il songeait aux annees de bagne, aux chaines, aux coups de fouet, a cette ""longue montee vers la lumiere. Cosette dormait, et il la regardait dormir.","Les rues de Paris etaient sombres en ce temps-la. Les reverberes jetaient ""des taches de lumiere jaune sur les paves mouilles. Jean Valjean marchait ""vite, comme un homme qui fuit quelque chose.","Marius l'observait de loin. Ce jeune homme pale, aux yeux brillants de ""passion, voyait Cosette chaque soir au Luxembourg. Il n'osait lui parler. ""L'amour est timide, meme pour les coeurs les plus courageux.","La barricade s'elevait dans la rue. Les jeunes hommes avaient pris les ""paves et construit un mur de pierres. Enjolras commandait avec une ""autorite naturelle. Grantaire, ivre et fidele, le suivait partout.","Gavroche passait entre les balles comme un oiseau. Ce gamin de Paris, ""nenfant abandonne devenu enfant de la patrie, ramassait les cartouches ""des soldats morts pour les rapporter aux insurges.","Eponine errait dans les rues, le coeur lourd d'un amour sans espoir. ""Elle avait vu Marius regarder Cosette, et elle avait compris que son ""propre amour ne serait jamais partage.","Javet suivait la piste avec une tenacite implacable. Cet homme de loi ""ne connaissait ni pitié ni doute. Pour lui, Jean Valjean resterait ""toujours le forcat 24601, quelles que soient ses actions.","Le jardin du Luxembourg etait paisible en ce matin de printemps. ""Les marronniers etaient en fleur, et les enfants jouaient sur le sable. ""Cosette lisait sur un banc, attendant celui qui ne viendrait pas.","Les Thenardier comptaient leur butin dans l'arriere-salle de l'auberge. ""Monsieur Thenardier calculait, madame Thenardier injuriait. La cupidite ""etait leur religion, l'argent leur seul dieu.",]train_data = Dataset.from_dict({"text": training_texts})print(f"Dataset : {len(train_data)} exemples")print(f"Moyenne caracteres/exemple : {sum(len(t) for t in training_texts) /len(training_texts):.0f}")
The dataset contains 10 excerpts of classic French text (average 200 characters per excerpt). These segments are shorter than in FT-01 to fit the limited context of OPT-1.3B. The next step is tokenization with uniform padding.
Reading the committed dataset: 10 examples, 200 characters on average – a tiny corpus, chosen to keep training within about twenty seconds. The excerpts are from Hugo and the like (public domain): 19th-century lexicon, long sentences, recurring proper nouns (Jean Valjean, Cosette). The goal is not to teach French to OPT-1.3B but to graft a style: with 10 examples x 15 epochs, the LoRA adapter captures vocabulary and turns of phrase, not deep grammar. A real fine-tuning dataset holds thousands of examples; here each example is seen 15 times, memorization wins over generalization – and that is precisely what the generation of section 7 will reveal.
# Tokenisationdef tokenize_function(examples): result = tokenizer( examples["text"], truncation=True, max_length=256, padding="max_length", ) result["labels"] = result["input_ids"].copy()return resulttokenized_dataset = train_data.map(tokenize_function, batched=True, remove_columns=["text"])print(f"Tokens par exemple : {len(tokenized_dataset[0]['input_ids'])}")print(f"Dataset tokenise : {len(tokenized_dataset)} exemples")
Tokens par exemple : 256
Dataset tokenise : 10 exemples
The dataset is now tokenized with uniform padding to 256 tokens, ready for training. Each example contains the input_ids and the labels (copies of the input_ids for causal language modeling).
We will now configure the Hugging Face Trainer with the QLoRA arguments: 8-bit optimizer (paged_adamw_8bit), FP16, and a learning rate adapted to the LoRA adapter.
The hidden cost of uniform padding: 256 tokens per example for 200 characters of text (~50-60 useful tokens in French) – about 75% of each sequence is padding. padding="max_length" simplifies the batch at the price of quadrupling the forward-pass computation on null positions. Over 10 examples this is invisible (the run below stays in the order of a minute); on a real corpus, one would prefer dynamic per-batch padding. Keep the reading reflex: 10 examples x 256 tokens = 2,560 training tokens in total – the LoRA adapter (3.1M parameters) has seen fewer tokens than it has weights, an extreme memo-styling regime.
Perte finale : 2.3855
Temps : 21.6s
VRAM pic entrainement : 1.72 Go
Reading the result: the effective batch and the learning budget
The training arguments above encode several coupled decisions:
per_device_train_batch_size=2 and the implicit accumulation of 1 give an effective batch of 2. For a larger model, you would reduce the per-device batch and raise gradient_accumulation_steps.
learning_rate=2e-4 is typical of LoRA / QLoRA: the optimizer only touches the adapters, while the base model’s weights stay frozen.
optim="paged_adamw_8bit" and fp16=True respectively reduce the memory of the Adam states and of the activations.
num_train_epochs=15, logging_steps=10 and weight_decay=0.01 produce 75 steps over the 10 excerpts. The output table shows the loss going down; the absolute time depends on the GPU, the cache, and the software versions, so it stays in the output rather than in this interpretation.
warmup_ratio and max_grad_norm are not declared here; the Trainer defaults apply. On a real corpus, an explicit warmup and clipping would deserve to be tested.
The intuition to keep: LoRA/QLoRA uses the same optimization toolbox as dense fine-tuning, but over a much smaller parameter surface. Read the loss curve, the memory peak, and the generations together: no isolated number suffices to conclude on generalization.
7. Generation with the QLoRA model
Let’s compare the generation before and after fine-tuning.
Reading criterion before running: compare the three committed generations along two distinct axes – the lexicon (proper nouns, literary register: expect Jean Valjean, Cosette, dark streets) and the syntax (agreements, clause construction: expect fractures). The three prompts are different openings on the same universe; the question is not “is it beautiful?” but “what has been learned (lexicon) and what has not (grammar)?”. The detailed reading follows the outputs.
# Generer avec le modele fine-tunemodel_qlora.eval()def generate_qlora(prompt, max_new_tokens=100, temperature=0.8): inputs = tokenizer(prompt, return_tensors="pt").to(model_qlora.device)with torch.no_grad(): outputs = model_qlora.generate(**inputs, max_new_tokens=max_new_tokens, temperature=temperature, do_sample=True, top_p=0.9, pad_token_id=tokenizer.eos_token_id, )return tokenizer.decode(outputs[0], skip_special_tokens=True)prompts = ["Dans les rues sombres de Paris,","Jean Valjean regardait la nuit","Cosette marchait dans le jardin",]for p in prompts:print(f"\n{'='*60}")print(f"Prompt : \"{p}\"")print(f"{'='*60}") output = generate_qlora(p)print(output)
============================================================
Prompt : "Dans les rues sombres de Paris,"
============================================================
Dans les rues sombres de Paris, Jean Valjean restait dans son bain de boulon. Il restait dans sa propre cite, et il restait dans le cadre de son lourdement. Cosette n'a jamais vu son père dans ce bain, et il n'a jamais regarde ses jolies gues. L'homme n'existait qu'une figure, et l'homme n'
============================================================
Prompt : "Jean Valjean regardait la nuit"
============================================================
Jean Valjean regardait la nuit comme une espoir. Cosette dormait dans un logis, il dormait sur le floor, et Cosette dormait en lui. Jean Valjean regardait la nuit comme un espoir. Cosette dormait dans un logis, il dormait sur le floor, et Cosette dormait en lui.
============================================================
Prompt : "Cosette marchait dans le jardin"
============================================================
Cosette marchait dans le jardin où elle avait bien connu le sujet de la veille. Elle avait lu son discours sur le sujet de la ville, ce qui lui avait permis de le croire. Cosette avait lu celui qu'elle ne viendra pas. Elle avait lu celui qu'elle ne viendrait pas. Les deux vivraient un coeur quelque chose. La l
Reading the result: the stylistic graft of fine-tuning
The generation « Dans les rues sombres de Paris, Jean Valjean… » proves that QLoRA fine-tuning has modified the model’s behavior. The base OPT-1.3B, trained on a generic English-language corpus, now produces French that reuses the vocabulary of Les Misérables (Jean Valjean, rues sombres de Paris) — the graft of the French classics dataset has taken. But also observe the limits: « la figure de la sachant » (repeated in a loop), « Cosette lui donneait » are grammatically incoherent fragments. This is the expected trade-off of a model of 1.3B parameters fine-tuned on only 10 excerpts: the stylistic imprint takes, deep syntactic coherence does not follow. This very imbalance is what the next comparison cell quantifies (quality ~97-99% of full fine-tuning).
Close reading of the 2nd and 3rd generations: the prompt « Jean Valjean regardait la nuit » produces « Cosette lui donneait ses cartes de mariage » — « donneait » is an invented conjugation on a real root — then drifts into a loop (« Cosette regardait Jean Valjean et regardait la table », repeated): the signature of a small memo-stylized model. The third generation is the most revealing: « elle continuait le moulinette » (faulty article on a real noun) and « Jean Valjean resterait dans la lune » — an image absurd in semantics, Hugo-like in register. Diagnosis: vocabulary and atmosphere borrowed from the corpus, morphology and coherence intrinsic to the base model. Fifteen epochs on 10 excerpts graft a lexicon, not a grammar.
8. LoRA vs QLoRA Comparison
Let’s load the same model in FP16 (without quantization) to compare the memory footprint.
What the comparison measures: the following table reloads the same OPT-1.3B in pure FP16, then in 4-bit, each time with reset_peak_memory_stats() before loading and a peak read afterwards – three independent measurements of the same model in three configurations. The FP16 peak will serve as the reference; the final row reuses the QLoRA training peak measured in section 6. Careful when reading: the first two rows measure loading peaks (dequantization transients included), the third a steady-state training peak – hence a number that may look lower than the row above.
# Charger le modele en FP16 pour comparaisonmodel_fp16 = AutoModelForCausalLM.from_pretrained( MODEL_NAME, torch_dtype=torch.float16, device_map="auto",)# Ajouter LoRA identiquelora_config_fp16 = LoraConfig( r=16, lora_alpha=32, target_modules=["q_proj", "v_proj"], lora_dropout=0.05, bias="none", task_type="CAUSAL_LM",)model_lora_fp16 = get_peft_model(model_fp16, lora_config_fp16)# Mesures comparativesdef measure_vram(model, label):if torch.cuda.is_available(): alloc = torch.cuda.memory_allocated() /1e9 reserved = torch.cuda.memory_reserved() /1e9print(f"{label} : alloue={alloc:.2f} Go, reserve={reserved:.2f} Go")return allocreturn0print(f"\n{'='*50}")print(f"Comparaison memoire — {MODEL_NAME}")print(f"{'='*50}")# Nettoyer pour mesure propredel model_fp16, model_lora_fp16gc.collect()if torch.cuda.is_available(): torch.cuda.empty_cache()# Recharger pour mesure propretorch.cuda.reset_peak_memory_stats()model_fp16 = AutoModelForCausalLM.from_pretrained( MODEL_NAME, torch_dtype=torch.float16, device_map="auto",)vram_fp16 = torch.cuda.max_memory_allocated() /1e9del model_fp16gc.collect()torch.cuda.empty_cache()torch.cuda.reset_peak_memory_stats()model_4b = AutoModelForCausalLM.from_pretrained( MODEL_NAME, quantization_config=bnb_config, device_map="auto",)vram_4bit_load = torch.cuda.max_memory_allocated() /1e9del model_4bgc.collect()if torch.cuda.is_available(): torch.cuda.empty_cache()print(f"\n{'Configuration':<25}{'VRAM':>10}{'Economie':>10}")print(f"{'-'*25}{'-'*10}{'-'*10}")print(f"{'FP16 (base)':<25}{vram_fp16:>9.2f} G {'ref':>10}")print(f"{'4-bit NF4 (QLoRA base)':<25}{vram_4bit_load:>9.2f} G {(1-vram_4bit_load/vram_fp16)*100:>9.1f}%")print(f"{'QLoRA (4-bit + LoRA r=16)':<25}{vram_peak:>9.2f} G {(1-vram_peak/vram_fp16)*100:>9.1f}%")print(f"\nTheorie : FP16 = {total_params*2/1e9:.2f} Go vs 4-bit = {total_params*0.5/1e9:.2f} Go (ratio 4x)")
Configuration VRAM Economie
------------------------- ---------- ----------
FP16 (base) 4.03 G ref
4-bit NF4 (QLoRA base) 2.24 G 44.4%
QLoRA (4-bit + LoRA r=16) 1.72 G 57.4%
Theorie : FP16 = 1.42 Go vs 4-bit = 0.36 Go (ratio 4x)
Reading the result: when the measurement happens, and its limits
The previous cell has three measurement steps whose order matters to understand the figure displayed in the next cell:
First load (model_fp16): measures the cost of an unquantized OPT-1.3B. The del model_fp16, model_lora_fp16 + gc.collect() + torch.cuda.empty_cache() that follows is non-optional: PyTorch does not explicitly release GPU memory as long as it can recycle it for future allocations. Without this cleanup, the second load’s measurement (model_4b) would include the first one’s footprint, skewing the displayed saving.
Second load (model_4b): OPT-1.3B in NF4 quantization. We observe 2.24 G vs 4.03 G in FP16 → 44.4% saved. But theory announces ~4x savings (16 bits → 4 bits), not 1.8x: the gap comes from the extra CUDA buffers that bitsandbytes allocates for quantization lookups, and from HuggingFace’s internal states (attention cache, KV cache). On a 1.3B, these overheads carry a sizable fraction of the total – on a 7B+, the ratio approaches theory.
vram_peak (3rd implicit measurement, printed by the previous cell): 1.72 G, i.e. 57.4% saved on the memory peak reached during training. This is the figure that matters to know whether the fine-tuning fits in your VRAM – not the static loading measurement, which ignores forward activations, LoRA gradients, and Adam moments. vram_peak captures the greediest instant, just before offloading.
The intuition to keep: a model’s footprint is never measured once. Three steps (FP16 / 4-bit base / 4-bit + LoRA) because three questions: how much does the model alone weigh?, how much does quantization alone weigh?, how much does the full training cost?. The number that should trigger a GPU purchase is vram_peak, not vram_fp16.
The trap to avoid: believing the measurement will reproduce bit for bit. CUDA non-deterministic memory allocation, the order of the dels, the presence of other processes on the GPU make the figure vary by ±10% between runs. For serious comparisons (publication, calibration), take an average over 5 runs with torch.cuda.empty_cache() interleaved.
The next cell prints the summary table of these three figures – together with the estimated quality and speed, which are not measured at runtime but predicted by theory (NF4 ≈ 97-99% of full FT; ~20% slowdown from the dequantization at every forward).
======================================================================
RESUME COMPARATIF : LoRA (FP16) vs QLoRA (4-bit)
======================================================================
Aspect LoRA (FP16) QLoRA (4-bit)
------------------------------ -------------------- --------------------
Precision poids FP16 (16 bits) NF4 (4 bits)
Calculs forward FP16 FP16 (dequant)
Grad + optimizer FP16 (LoRA seul) FP16 (LoRA seul)
VRAM modele ~1.3B ~2.6 Go ~0.7 Go
Qualite vs full FT ~98-99% ~97-99%
Vitesse entrainement 1x (ref) ~0.8x (dequant)
Recommendation :
- Modele < 1B + VRAM suffisante -> LoRA (FP16) : plus simple, plus rapide
- Modele > 1B ou VRAM limitee -> QLoRA : Economie memoire 3-4x
Reading the result: the VRAM / quality / speed triptych
The summary table condenses the promise and the cost of QLoRA along three axes:
VRAM (the decisive gain): ~0.7 GB (QLoRA) vs ~2.6 GB (LoRA FP16), i.e. ~3.7× memory savings. This is what makes a 1.3 B model fine-tunable on a consumer card (8-12 GB) instead of a 24 GB one. The key: only the NF4 (4-bit) weights live in VRAM; gradients and optimizer only carry the LoRA adapters (FP16), hence the line Grad + optimizer: FP16 (LoRA only) identical on both sides.
Quality (the price to pay): ~97-99% of full fine-tuning. NF4 quantization preserves information because it concentrates precision where weights are dense (assumed Gaussian distribution), but a 1-3% degradation is intrinsic to the 16→4-bit information loss.
Speed (the hidden cost): ~0.8×, i.e. ~20% slowdown. The NF4→FP16 dequantization at each forward pass has a computational overhead — QLoRA is not free in training time.
The recommendation that closes the table encodes the decision rule: LoRA (FP16) if VRAM allows (simpler, faster), QLoRA as soon as the model exceeds ~1 B parameters or VRAM is constrained. This positioning — fine-tuning large models on small hardware — is what made QLoRA the open fine-tuning standard on consumer GPUs.
9. GPU memory cleanup
Why this cell matters: returning the VRAM (14.3 GB free after cleanup) is not politeness — the comparison cell below reloads the model twice, and without del + empty_cache(), the CUDA cache would saturate the measurements. On a multi-experiment GPU notebook, memory discipline is the condition for honest measurements.
This section applies the notebook’s concepts in two stages: four solved and commented guided examples, then four exercises to complete that measure what the examples do not cover.
The four guided examples come from a student contribution (@nassimJaz, PR #18550): each statement, its measurement code and the reading of the result that follows form a self-contained triplet.
Impact of precision – INT8 vs NF4 vs FP16: loading VRAM, weight footprint, perplexity and generations.
LoRA rank and memory – r = 4, 16, 64: trainable parameters, VRAM peak, final loss, generations.
Double quantization – bnb_4bit_use_double_quant on/off: VRAM gap measured against theory (~0.37 bit/param).
Breaking the initialization invariant – B initialized non-zero in the LoRALinear of section 4.
The exercises to complete (last subsection) take over: reconstruction error of quantized weights, memorization versus generalization, NF4 versus FP4 at equal memory, alpha/r versus alpha/sqrt(r) scaling. Each statement carries its measurable success criterion.
Guided example 1: precision impact (INT8 vs NF4 vs FP16)
Student contribution by Nassim JAZOULI (@nassimJaz) and Minh NGUYEN (PR #18550), integrated as a guided example.
Compare INT8 (load_in_8bit=True), NF4 (4-bit with double quantization), and FP16 on the same OPT-1.3B. For each precision: measure the loading VRAM (peak), the weight footprint (get_memory_footprint), generate greedily on the same prompt, and compute the perplexity over the 10 training excerpts as a quantified measure of the degradation.
What theory announces: ~2.6 GB of weights alone in FP16 (2 bytes/parameter), ~1.3 GB in INT8 (1 byte), ~0.7 GB in NF4 (0.5 byte); the loading VRAM must therefore order FP16 > INT8 > NF4. Generation quality must degrade in the same order, but weakly – the three models remain readable on a 1.3B.
Steps followed by the solution:
Define the three loading configurations (one torch_dtype, two BitsAndBytesConfig).
For each: reset_peak_memory_stats(), load, measure peak and footprint, generate, compute the perplexity over the excerpts, then free the memory.
Summarize in a table with the VRAM saving relative to FP16.
# Exemple guide 1 : Impact de la precision# Contribution etudiante de @nassimJaz (PR #18550)# Comparaison INT8 (load_in_8bit=True) vs NF4 vs FP16 sur OPT-1.3B.# Pour chaque precision : VRAM de chargement (pic) + generation sur un meme prompt.# Liberer ce qui reste des sections precedentes (model_4bit est encore reference)for _name in ["model_4bit", "model_qlora", "trainer", "model_fp16", "model_4b"]:globals().pop(_name, None)gc.collect()torch.cuda.empty_cache()configs_precision = {"FP16": dict(torch_dtype=torch.float16),"INT8": dict(quantization_config=BitsAndBytesConfig(load_in_8bit=True)),"NF4": dict(quantization_config=BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.float16, bnb_4bit_use_double_quant=True, )),}prompt_ex1 ="Dans les rues sombres de Paris,"resultats_ex1 = {}for label, kwargs in configs_precision.items(): gc.collect() torch.cuda.empty_cache() torch.cuda.reset_peak_memory_stats() m = AutoModelForCausalLM.from_pretrained(MODEL_NAME, device_map="auto", **kwargs) m.eval() vram_load = torch.cuda.max_memory_allocated() /1e9 empreinte = m.get_memory_footprint() /1e9# Generation greedy (deterministe) pour comparer la qualite a entree egale inputs = tokenizer(prompt_ex1, return_tensors="pt").to(m.device)with torch.no_grad(): out = m.generate(**inputs, max_new_tokens=40, do_sample=False, repetition_penalty=1.2, pad_token_id=tokenizer.eos_token_id) texte = tokenizer.decode(out[0], skip_special_tokens=True)# Perplexite sur les textes d'entrainement : mesure chiffree de la degradation nll, n_tok =0.0, 0with torch.no_grad():for t in training_texts: enc = tokenizer(t, return_tensors="pt").to(m.device) loss = m(**enc, labels=enc["input_ids"]).loss n = enc["input_ids"].shape[1] -1 nll += loss.item() * n n_tok += n ppl =float(torch.exp(torch.tensor(nll / n_tok))) resultats_ex1[label] =dict(vram=vram_load, footprint=empreinte, ppl=ppl, texte=texte)print(f"\n[{label}] VRAM pic chargement = {vram_load:.2f} Go | empreinte poids = {empreinte:.2f} Go | perplexite = {ppl:.2f}")print(f" -> {texte}")del m, inputs, out gc.collect() torch.cuda.empty_cache()print(f"\n{'Precision':<10}{'VRAM pic':>10}{'Empreinte':>10}{'Perplexite':>11}{'Economie VRAM':>14}")ref = resultats_ex1["FP16"]["vram"]for label, r in resultats_ex1.items():print(f"{label:<10}{r['vram']:>8.2f} G {r['footprint']:>8.2f} G {r['ppl']:>11.2f}{(1- r['vram'] / ref) *100:>13.1f}%")
[FP16] VRAM pic chargement = 3.39 Go | empreinte poids = 2.63 Go | perplexite = 33.89
-> Dans les rues sombres de Paris, des manifestants ont dénoncé le « traitement » du coronavirus.
« Je ne suis pas un virus ! », a-t-il lancé dans
[INT8] VRAM pic chargement = 2.18 Go | empreinte poids = 1.42 Go | perplexite = 33.58
-> Dans les rues sombres de Paris, des manifestants ont dénoncé la « mauvaise gestion » du coronavirus.
« Je suis triste pour l’économie et je ne pe
[NF4] VRAM pic chargement = 1.60 Go | empreinte poids = 0.82 Go | perplexite = 34.47
-> Dans les rues sombres de Paris, des gens se retrouvent désormais à la recherche du monde.
C'est ce qu'on appelle le "café-déje
Precision VRAM pic Empreinte Perplexite Economie VRAM
FP16 3.39 G 2.63 G 33.89 0.0%
INT8 2.18 G 1.42 G 33.58 35.6%
NF4 1.60 G 0.82 G 34.47 52.9%
Reading the result – guided example 1: precision buys VRAM, not quality
VRAM: the expected order holds, FP16 (3.39 GB) > INT8 (2.18 GB) > NF4 (1.60 GB) at the loading peak, i.e. 52.9% saved for NF4. The weight footprint (get_memory_footprint: 2.63 / 1.42 / 0.82 GB) roughly follows theory (2 bytes, 1 byte, 0.5 byte per quantized parameter). The gap from the theoretical 4x ratio comes from the layers that are not quantized: the embeddings (token + positions, 50,272 x 2048) and lm_head stay in FP16, as do the LayerNorms. On a 1.3B these layers carry a sizable share of the total, so the relative gain is smaller than on a 7B.
Quality: the perplexity measured over the 10 excerpts varies by less than one point across the three formats (33.89 / 33.58 / 34.47). The theoretical ranking FP16 ≤ INT8 ≤ NF4 is not settled by the measurement: INT8 even displays slightly below FP16 – a 0.3-point gap, smaller than the measurement noise – and only NF4 stands apart a little (+0.6 point). The greedy generations are identical on the first tokens (« Dans les rues sombres de Paris, des… »), then diverge slightly: the three models remain readable, NF4 retains the essential information.
Remark: INT8 (LLM.int8()) is often slower than FP16 because of the outlier decomposition, whereas NF4 dequantizes by blocks in FP16. INT8 is therefore not an obvious compromise: for memory-constrained fine-tuning, NF4 is the best choice, which justifies QLoRA.
Guided example 2: LoRA rank and memory (r = 4, 16, 64)
Student contribution by Nassim JAZOULI (@nassimJaz) and Minh NGUYEN (PR #18550), integrated as a guided example.
Train QLoRA on OPT-1.3B for three ranks – r = 4, 16, 64 – keeping everything else identical. For each rank: number of trainable parameters, training VRAM peak, final loss, duration, and one generation on a fixed prompt (fixed seed). For the comparison to be fair, alpha = 2r: the scale alpha/r = 2 is the same for the three ranks, only the capacity changes.
What theory announces: the trainable parameters follow exactly 196,608 x r (formula proven in section 4 – the LoRA mechanics), i.e. 786,432, 3,145,728, and 12,582,912; the adapters stay tiny next to the frozen NF4 weights; the final loss must decrease as r rises – but over 10 examples seen 15 times, a high rank memorizes rather than learns: compare the generations, not just the loss.
Steps followed by the solution:
Define a function that loads OPT-1.3B in NF4, grafts a rank-r LoRA (alpha = 2r), and trains 15 epochs with the hyperparameters of section 6.
Call it for r = 4, 16, 64; measure each time the VRAM peak, the final loss, and the duration.
Compare in a table: rank, trainable parameters (measured against predicted by 196,608 x r), VRAM peak, final loss, time.
# Exemple guide 2 : LoRA rank et memoire# Contribution etudiante de @nassimJaz (PR #18550)# r = 4, 16, 64 avec QLoRA sur OPT-1.3B : params entrainables, VRAM pic, perte finale.def train_qlora_rank(r, epochs=15):"""Charge OPT-1.3B en NF4, greffe LoRA de rang r (alpha = 2r), entraine et mesure.""" gc.collect() torch.cuda.empty_cache() base = AutoModelForCausalLM.from_pretrained(MODEL_NAME, quantization_config=bnb_config, device_map="auto") base = prepare_model_for_kbit_training(base) cfg = LoraConfig(r=r, lora_alpha=2* r, target_modules=["q_proj", "v_proj"], lora_dropout=0.05, bias="none", task_type="CAUSAL_LM") m = get_peft_model(base, cfg) n_train =sum(p.numel() for p in m.parameters() if p.requires_grad) args = TrainingArguments( output_dir=f"./temp_ft02_r{r}", num_train_epochs=epochs, per_device_train_batch_size=2, learning_rate=2e-4, weight_decay=0.01, logging_steps=25, save_strategy="no", report_to="none", fp16=True, seed=42, optim="paged_adamw_8bit", ) tr = Trainer(model=m, args=args, train_dataset=tokenized_dataset, data_collator=data_collator) torch.cuda.reset_peak_memory_stats() t0 = time.time() res = tr.train() duree = time.time() - t0 pic = torch.cuda.max_memory_allocated() /1e9# Generation echantillonnee (seed fixe) pour comparer qualitativement m.eval() torch.manual_seed(0) inputs = tokenizer("Jean Valjean regardait la nuit", return_tensors="pt").to(m.device)with torch.no_grad(): out = m.generate(**inputs, max_new_tokens=50, do_sample=True, temperature=0.8, top_p=0.9, pad_token_id=tokenizer.eos_token_id) texte = tokenizer.decode(out[0], skip_special_tokens=True)del m, base, tr, inputs, out gc.collect() torch.cuda.empty_cache()returndict(r=r, trainable=n_train, predit=196_608* r, vram_peak=pic, loss=res.training_loss, temps=duree, texte=texte)resultats_ex2 = []for r in [4, 16, 64]:print(f"\n===== r = {r} =====") res = train_qlora_rank(r) resultats_ex2.append(res)print(f"generation : {res['texte']}")print(f"\n{'r':>3}{'params entrainables':>20}{'predit 196608*r':>16}{'VRAM pic':>10}{'perte finale':>13}{'temps':>8}")for res in resultats_ex2:print(f"{res['r']:>3}{res['trainable']:>20,}{res['predit']:>16,}{res['vram_peak']:>8.2f} G "f"{res['loss']:>13.4f}{res['temps']:>6.1f} s")
===== r = 4 =====
[75/75 00:19, Epoch 15/15]
Step
Training Loss
25
3.325595
50
2.799987
75
2.524201
generation : Jean Valjean regardait la nuit. Les garçons de la rue ne faisaient que jouer. Les hommes de la rue ne jouaient pas. Valjean n'apprenait jamais ce qu'on avait
===== r = 16 =====
[75/75 00:13, Epoch 15/15]
Step
Training Loss
25
3.003321
50
2.243074
75
1.926127
generation : Jean Valjean regardait la nuit dans un bateau. Le courage et la determination de celui-la lui suivait lui-même. Cosette avait lu le navetement de son pere, mais elle n
===== r = 64 =====
[75/75 00:14, Epoch 15/15]
Step
Training Loss
25
2.741246
50
1.715002
75
1.229577
generation : Jean Valjean regardait la nuit dans un bateau. Il ne savait rien pour ce soir. Cosette regardait lui, et il ne lui regardait pas. Ce soir, l'homme vivait dans la nature.
r params entrainables predit 196608*r VRAM pic perte finale temps
4 786,432 786,432 2.13 G 2.8833 19.6 s
16 3,145,728 3,145,728 2.57 G 2.3908 13.6 s
64 12,582,912 12,582,912 3.03 G 1.8953 15.0 s
Reading the result – guided example 2: rank buys capacity, not a better model
Trainable parameters: they follow exactly the formula 196,608 x r proven in section 4 (786,432 / 3,145,728 / 12,582,912) – the “predicted” column coincides with the actual count.
VRAM peak: it climbs from 2.13 GB (r=4) to 2.57 GB (r=16) then 3.03 GB (r=64) – about 0.9 GB of spread. The adapters alone stay tiny (~12.6M parameters at r=64, i.e. ~25 MB of FP16 weights): the difference comes from the gradients, from the paged optimizer states, and from the intermediate buffers that accompany every trainable parameter – a VRAM peak is measured, not deduced from a parameter count. The frozen NF4 weights and the activations (256-token sequences) still dominate memory, but at r=64 the adapter is no longer free.
Final loss: it decreases as r rises (2.88 → 2.39 → 1.90), but with 10 examples seen 15 times, a high rank mostly serves to memorize the corpus. The generations show it: at r=4, the model loops on short interchangeable sentences (« Les garçons de la rue ne faisaient que jouer »); at r=16 and r=64, it ventures into more ambitious but faulty turns of phrase (« le navetement de son pere », « Cosette regardait lui ») – the loss drops, the grammar does not follow. A lower loss is therefore not synonymous with a better model: without a validation set, you only measure memorization – which is exactly the question exercise 2 below turns back on. For this dataset, r=8-16 remains the reasonable compromise.
For the comparison to be fair, alpha = 2r: the scale alpha/r = 2 is identical for the three ranks, only the capacity changes.
Guided example 3: double quantization
Student contribution from Nassim JAZOULI (@nassimJaz) and Minh NGUYEN (PR #18550), integrated as a guided example.
Double quantization (bnb_4bit_use_double_quant=True) compresses the quantization constants themselves into FP8, saving ~0.37 bit/param. Compare with and without it to measure the real impact on VRAM and quality.
Protocol: measure the loading VRAM in both configurations with reset_peak_memory_stats() before each from_pretrained (protocol of section 3). The expected gap – ~56 MB for the ~1.21B parameters actually quantized in OPT-1.3B (x 0.373 bit / 8) – is small compared with the noise of the loading peaks: run each measurement 3 times and compare the minima, not the single peaks. A zero gap does not mean “double quantization does not work”: first check that the two BitsAndBytesConfig differ only by bnb_4bit_use_double_quant.
# Exemple guide 3 : Double quantization on/off# Contribution etudiante de @nassimJaz (PR #18550)# Protocole : reset_peak_memory_stats avant chaque from_pretrained, 3 repetitions, on garde les minima.def charger_nf4(double_quant): cfg = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.float16, bnb_4bit_use_double_quant=double_quant, # SEULE difference entre les deux configs )return AutoModelForCausalLM.from_pretrained(MODEL_NAME, quantization_config=cfg, device_map="auto")resultats_ex3 = {}for dq in [False, True]: pics, allocs = [], []for _ inrange(3): gc.collect() torch.cuda.empty_cache() torch.cuda.reset_peak_memory_stats() m = charger_nf4(dq) pics.append(torch.cuda.max_memory_allocated() /1e9) allocs.append(torch.cuda.memory_allocated() /1e9) footprint = m.get_memory_footprint() /1e9del m# Qualite : perplexite sur les extraits (un chargement supplementaire) m = charger_nf4(dq) m.eval() nll, n_tok =0.0, 0with torch.no_grad():for t in training_texts: enc = tokenizer(t, return_tensors="pt").to(m.device) n = enc["input_ids"].shape[1] -1 nll += m(**enc, labels=enc["input_ids"]).loss.item() * n n_tok += n ppl =float(torch.exp(torch.tensor(nll / n_tok)))del m, enc gc.collect() torch.cuda.empty_cache() resultats_ex3[dq] =dict(pic=min(pics), alloc=min(allocs), footprint=footprint, ppl=ppl)print(f"{'double_quant':<14}{'pic min':>9}{'alloue min':>11}{'empreinte':>10}{'perplexite':>11}")for dq, r in resultats_ex3.items():print(f"{str(dq):<14}{r['pic']:>7.3f} G {r['alloc']:>9.3f} G {r['footprint']:>8.3f} G {r['ppl']:>11.3f}")ecart_mo = (resultats_ex3[False]["alloc"] - resultats_ex3[True]["alloc"]) *1e3print(f"\nEconomie mesuree (memoire allouee) : {ecart_mo:.1f} Mo")# Theorie : seules les couches lineaires des blocs sont quantifiees (pas les embeddings / lm_head)n_quant =24* (4*2048*2048+2*2048*8192)print(f"Params quantifies : {n_quant:,} -> theorie 0.373 bit/param = {n_quant *0.373/8/1e6:.1f} Mo")
double_quant pic min alloue min empreinte perplexite
False 2.938 G 2.630 G 0.820 G 34.457
True 2.882 G 2.573 G 0.820 G 34.470
Economie mesuree (memoire allouee) : 56.1 Mo
Params quantifies : 1,207,959,552 -> theorie 0.373 bit/param = 56.3 Mo
Reading the result – guided example 3: a minuscule gain, linear in model size
Theory: without double quantization, each block of 64 weights has an FP32 scale constant → 32/64 = 0.5 bit/param. With double quantization, those constants are stored in 8 bits per block of 256, plus one FP32 constant per block of 256 constants → 8/64 + 32/(64·256) ≈ 0.127 bit/param. Gain: ≈ 0.373 bit/param.
Measurement: only the linear layers of the 24 decoder blocks are quantized (1,207,959,552 parameters, i.e. ≈ 1.21B; the embeddings and lm_head stay in FP16). We therefore expect ≈ 56 MB saved – and the measured gap on allocated memory lands right on it: 56.1 MB measured vs 56.3 MB predicted by theory. On the loading peak, the same-order gap (2.938 vs 2.882 GB) is noisier because the peak contains temporary conversion buffers: hence the value of repeating the measurement and comparing minima.
Quality: the perplexity is nearly identical between the two configurations (34.457 vs 34.470 – a difference on the order of numerical noise): the scale constants quantized to 8 bits are precise enough to change nothing in the result.
Conclusion: on OPT-1.3B the gain is anecdotal (~2% of VRAM), but it grows linearly with size: ~0.3 GB for a 7B, ~3 GB for a 65B – this is what allowed the QLoRA paper to fit a 65B on a single 48 GB GPU. It is a “free” gain, hence its activation by default.
Guided example 4: breaking the initialization invariant
Student contribution by Nassim JAZOULI (@nassimJaz) and Minh NGUYEN (PR #18550), integrated as a guided example.
The invariant of section 4: LoRALinear initializes B to zero, so ΔW = B·A = 0 and the fine-tunable model starts exactly as the pre-trained one. This example deliberately breaks that invariant: initialize B to normal(0, b_std) instead of zero, and observe what this choice destroys.
Measurements:
For b_std = 0, 0.001, 0.01, 0.1 on an nn.Linear layer 512x512 identical to section 4’s: maximum norm of the initial ΔW, ratio |s·ΔW| / |W₀|, test torch.allclose(lora(x), base(x)) before any training, maximum output gap.
Redo the regression of section 4 (50 SGD steps, same learning rate) with B = 0 then B ≠ 0: compare the initial and final losses.
Reasoning hint: take the class from section 4 and change only B’s initialization; if A and B were both zero, the gradients would be zero and nothing would ever learn.
# Exemple guide 4 : initialiser B NON nul et observer le démarrage# Contribution etudiante de @nassimJaz (PR #18550)def init_b_nonzero(base_layer, r=8, lora_alpha=16, b_std=0.01):""" Retourner un LoRALinear dont B est initialise a normal(0, b_std) au lieu de zero. Mesurer : delta_W initial, forward == base ?, et interpreter pourquoi le modele LoRA ne demarre PLUS exactement comme le pre-entraine. """ lora = LoRALinear(base_layer, r=r, lora_alpha=lora_alpha)with torch.no_grad(): lora.lora_B.normal_(0, b_std) # B != 0 -> delta_W = B @ A != 0return lora# Couche de base fraiche (identique a la section 4), pour ne pas dependre de lora_demotorch.manual_seed(42)base_ex4 = nn.Linear(512, 512, bias=False)with torch.no_grad(): base_ex4.weight.copy_(torch.randn(512, 512) *0.02)x_probe4 = torch.randn(4, 512)print(f"{'b_std':>7}{'max|dW|':>10}{'|s*dW|/|W0|':>12}{'forward==base':>14}{'ecart sortie max':>17}")for b_std in [0.0, 0.001, 0.01, 0.1]: torch.manual_seed(0) lora_nz = init_b_nonzero(base_ex4, r=8, lora_alpha=16, b_std=b_std) dW = lora_nz.delta_W().detach()with torch.no_grad(): out_nz, out_b = lora_nz(x_probe4), base_ex4(x_probe4) rel = (lora_nz.scaling * dW).norm() / base_ex4.weight.norm()print(f"{b_std:>7}{dW.abs().max():>10.2e}{rel:>12.2%}{str(torch.allclose(out_nz, out_b)):>14} "f"{(out_nz - out_b).abs().max():>17.2e}")# Consequence sur l'apprentissage : meme tache, meme lr, B=0 vs B!=0print("\nMeme regression que la section 4 (50 pas SGD, lr=0.1) :")torch.manual_seed(1)xb, tgt = torch.randn(16, 512), torch.randn(16, 512)for b_std in [0.0, 0.1]: torch.manual_seed(0) l = init_b_nonzero(base_ex4, b_std=b_std) opt4 = torch.optim.SGD([l.lora_A, l.lora_B], lr=0.1) losses = []for _ inrange(50): opt4.zero_grad() loss = ((l(xb) - tgt) **2).mean() losses.append(loss.item()) loss.backward() opt4.step()print(f" b_std={b_std:<4} : loss initiale = {losses[0]:.4f} -> finale = {losses[-1]:.4f}")
b_std max|dW| |s*dW|/|W0| forward==base ecart sortie max
0.0 0.00e+00 0.00% True 0.00e+00
0.001 4.38e-04 0.73% False 1.07e-02
0.01 4.38e-03 7.25% False 1.07e-01
0.1 4.38e-02 72.50% False 1.07e+00
Meme regression que la section 4 (50 pas SGD, lr=0.1) :
b_std=0.0 : loss initiale = 1.2279 -> finale = 1.1487
b_std=0.1 : loss initiale = 1.3386 -> finale = 1.2085
Reading the result – guided example 4: why B must be born zero
With B = 0 (b_std = 0), we recover the invariant: ΔW = 0, forward == base → True.
As soon as B ≠ 0, ΔW = B·A is a non-zero random matrix of rank ≤ r, and forward == base → False: before a single training step, the model is no longer the pre-trained one. It is perturbed by a noise (α/r)·B·A·x whose amplitude grows with b_std (column |s·ΔW| / |W₀|).
Why this is a problem: on a real LLM, this noise is injected into all the targeted layers (48 projections for OPT-1.3B) and propagates through the network: the starting loss is higher, and the first optimization steps serve to “unlearn” this noise instead of learning the task. With a large b_std, you can even destroy the model’s abilities (incoherent outputs) from initialization. You see it on the toy regression: the initial loss is higher with B ≠ 0.
Why A random and B zero (and not the reverse, nor both zero): if A and B were both zero, the gradients ∂L/∂A ∝ Bᵀ(...) and ∂L/∂B ∝ (...)(Ax)ᵀ would be zero: nothing would ever learn. With A random and B = 0, the product is 0 (exact start) but ∂L/∂B is non-zero from the first step, so learning starts. It is the only choice that combines exact start and non-zero gradient.
Exercises to complete
Four unsolved exercises close the notebook. Each measures something the guided examples do not measure: the error committed on the weights themselves, generalization (as opposed to memorization), the NF4 vs FP4 choice, and the scale of the LoRA updates. The stubs run without error; complete them cell by cell.
Exercise 1: weight reconstruction error
Guided example 1 measured what quantization costs in VRAM and in perplexity – never what it does to the weights themselves. Here we open the hood: on a single nn.Linear layer of OPT-1.3B (model.decoder.layers[0].self_attn.q_proj), compare the FP16 weights to their reconstructed NF4 and INT8 versions, and locate where the error concentrates.
Load OPT-1.3B in FP16 on CPU (device_map="cpu"), extract W_fp16 (2048 x 2048) from layer 0’s q_proj, then work only with that matrix.
NF4 round trip: bitsandbytes.functional.quantize_4bit then dequantize_4bit (the API expects blocks of 64 contiguous values, the matrix must be reshaped to (n_blocks, 64) and sit on CUDA).
Row-wise absmax INT8 round trip: each row divided by its absolute max, rounded over 127 levels, then multiplied back.
For each format: relative Frobenius error ‖W_rec − W‖_F / ‖W‖_F and maximum absolute error.
Locate the error: sort the weights by absolute magnitude and compute the share of the total error carried by the largest 1% of weights – the outliers that block quantization seeks to protect.
Hints: - The matrix alone is 2048 x 2048: keep it on CPU and move onto CUDA only what bnb.functional requires. - Also compare the shape of the two errors: NF4 is tailored to a Gaussian weight distribution, row-wise absmax INT8 to a uniform distribution.
Success criterion: a table format | relative Frobenius error | max absolute error | share of the error in the largest 1% of weights, with relative errors on the order of 10⁻² to 10⁻¹.
# Exercice 1 : erreur de reconstruction des poids (NF4 / INT8 contre FP16)# Etape 1 : charger OPT-1.3B sur CPU, extraire W_fp16 de q_proj (couche 0).# Etape 2 : round trip NF4 (quantize_4bit/dequantize_4bit, blocs de 64) et INT8 absmax par ligne.# Etape 3 : erreur relative de Frobenius + erreur absolue maximale pour chaque format.# Etape 4 : part de l'erreur totale portee par le 1 % des poids de plus grande magnitude.# Indice : bnb.functional travaille sur CUDA ; ne mover que la matrice 2048 x 2048.W_fp16 =None# TODO etudiantW_rec_nf4 =None# TODO etudiantW_rec_int8 =None# TODO etudiantprint("Exercice a completer : erreur de reconstruction des poids NF4/INT8 vs FP16")
Exercice a completer : erreur de reconstruction des poids NF4/INT8 vs FP16
Exercise 2: memorization vs generalization
Guided example 2 concludes that without a validation set, the final loss only measures memorization. Test it: split the 10 training excerpts into 8 training texts + 2 held-out texts, retrain QLoRA at ranks r = 4, 16, 64 on the 8 only, and measure the perplexity over the 2 held-out texts before and after training.
Build train_texts_8 (the first 8 excerpts of training_texts) and held_out_texts (the last 2); redo the tokenization of section 6 on the 8.
For each rank r ∈ {4, 16, 64}: measure the held-out perplexity of the NF4 model before training, then train on the 8 excerpts (same hyperparameters as guided example 2).
After training, measure the held-out perplexity again.
Report a table r | held-out ppl before | held-out ppl after | final training loss.
Hints: - Reuse the perplexity loop of guided example 1 (sum of the losses weighted by the token count). - Hypothesis to test: on this 8/2 split, the held-out perplexity should drop for small ranks (generalization) then degrade at r = 64 (memorization of the 8 excerpts). This is a plausible prediction, not a guaranteed result – with only 2 held-out texts, the uncertainty is large. The rank that minimizes the held-out ppl may differ from the one that minimizes the training loss, or the three ranks may be indistinguishable.
Success criterion: the table of the 3 ranks with held-out ppl before/after, and the name of the rank that minimizes the held-out perplexity.
# Exercice 2 : mémorisation contre généralisation (split 8/2 des extraits)# Etape 1 : couper training_texts en 8 extraits d'entrainement + 2 extraits held-out.# Etape 2 : pour r = 4, 16, 64, entrainer QLoRA sur les 8 extraits (protocole de l'exemple guide 2).# Etape 3 : perplexite sur les 2 extraits held-out AVANT (base NF4) et APRES entrainement.# Indice : reutiliser la boucle de perplexite de l'exemple guide 1.resultats_exo2 =None# TODO etudiantprint("Exercice a completer : quel rang generalise le mieux sur les 2 extraits held-out ?")
Exercice a completer : quel rang generalise le mieux sur les 2 extraits held-out ?
Exercise 3: NF4 vs FP4 at equal memory
NF4 and FP4 occupy exactly the same space (4 bits per weight): the difference is purely in the shape of the quantization levels. Compare the two at equal memory on OPT-1.3B.
Load OPT-1.3B twice in 4-bit: bnb_4bit_quant_type="nf4" then "fp4" – double quantization stays identical in both configurations.
Measure the footprint (get_memory_footprint) and the loading peak for each format: they must be nearly identical.
Measure the perplexity over the training texts for both formats.
Generate greedily on the same prompt and compare the outputs.
Hints: - A transformer’s pre-trained weights follow a roughly normal distribution, centered and concentrated near zero: NF4 places its levels on the quantiles of that distribution (more levels at the center). FP4 is not a uniform grid: it is a floating E2M1 format (1 sign bit, 2 exponent bits, 1 mantissa bit) whose normalized positive levels, as implemented in bitsandbytes (functional.py, SHA 833649043474794b8fe7a4136e0c40faf077b2e0), are 0, 0.0052, 0.1667, 0.25, 0.3333, 0.5, 0.6667, 1 – the gaps are not constant (denser at the center than at the extremes, but less extreme than NF4). NF4 therefore should win in perplexity at equal memory on Gaussian weights, but that is a prediction, not a guaranteed result. - Reuse the skeleton of guided example 3 changing ONLY bnb_4bit_quant_type.
Success criterion: footprints equal to within ~1 MB, the two generations side by side, and a table ppl(NF4) vs ppl(FP4). Either direction of the inequality (NF4 < FP4, NF4 ≈ FP4, NF4 > FP4) is acceptable if the protocol and the measurements support it – the ranking is an experimental result, not an expectation.
# Exercice 3 : NF4 contre FP4 à mémoire égale# Etape 1 : charger OPT-1.3B en 4-bit avec bnb_4bit_quant_type="nf4" puis "fp4"# (bnb_4bit_use_double_quant identique dans les deux).# Etape 2 : empreinte (get_memory_footprint) et pic de chargement pour chaque format.# Etape 3 : perplexite sur les textes d'entrainement + generation greedy sur un meme prompt.# Indice : ne changer QUE bnb_4bit_quant_type entre les deux configurations.resultats_exo3 =None# TODO etudiantprint("Exercice a completer : NF4 vs FP4 a memoire egale")
Exercice a completer : NF4 vs FP4 a memoire egale
Exercise 4: alpha/r vs alpha/sqrt(r) scaling (rsLoRA)
The section 4 LoRALinear scales the update by alpha / r. rsLoRA proposes alpha / sqrt(r): at high rank, the update does not collapse. Test it on the section 4 toy regression – CPU, fast, no heavy model.
Reuse the section 4 LoRALinear class (local copy with a scale parameter), on the same regression task: same base layer, same data, fixed seeds.
With alpha fixed (e.g. 16), train r = 4, 16, 64 for the same number of SGD steps (e.g. 300), same learning rate, for both scales: alpha/r and alpha/sqrt(r).
Report the final loss of the 6 configurations in a table.
Hints: - At fixed alpha, the alpha/r scale decays like 1/r: doubling the rank halves the update; alpha/sqrt(r) only divides it by √2. This is the hypothesis to confront with the measurement. - Fix all the seeds (torch.manual_seed) so that the only difference between two runs is the scale.
Success criterion: at r = 64, a clearly lower final loss with alpha/sqrt(r) than with alpha/r, and losses close between the two scales at r = 4.
# Exercice 4 : échelle alpha/r contre alpha/sqrt(r) (rsLoRA) sur la régression de la section 4# Etape 1 : reprendre LoRALinear avec un parametre d'echelle, alpha fixe (ex. 16), CPU.# Etape 2 : entrainer r = 4, 16, 64 sur le meme nombre de pas SGD pour les deux echelles.# Etape 3 : tableau des pertes finales des 6 configurations.# Indice : graines fixees, seule l'echelle doit differer entre deux runs.resultats_exo4 =None# TODO etudiantprint("Exercice a completer : alpha/r vs alpha/sqrt(r) a r = 4, 16, 64")
Exercice a completer : alpha/r vs alpha/sqrt(r) a r = 4, 16, 64
Summary
Aspect
LoRA (FT-01)
QLoRA (FT-02)
Model weights
FP16 (16 bits)
NF4 (4 bits)
Adapters
LoRA in FP16
LoRA in FP16
Forward computations
Native FP16
Dequantize → FP16 → re-quantize
Memory savings
~1.1x base model
~3-4x vs FP16
Quality
~98-99% of full FT
~97-99% of full FT
Recommended model
< 1B params
> 1B params
Key points: - QLoRA = Quantization (NF4) + LoRA: the best of both worlds - NF4 is optimal for weights (normal distribution), not a simple INT4 - Double quantization compresses scaling factors for additional gains - paged_adamw_8bit: 8-bit optimizer to save even more VRAM - A 7B model in QLoRA fits on a 6 GB GPU — democratizes fine-tuning
Reference: Dettmers et al., “QLoRA: Efficient Finetuning of Quantized LLMs” (2023).
Next steps (FT-03): Multi-task fine-tuning with several LoRA adapters.
The measured journey reads in the outputs: the NF4 loading, the exact number of trainable parameters, the duration, the VRAM peak and the final loss are all produced by the preceding cells. The section 4 grid predicts the adapter count to the exact parameter; the section 7 generations then show memo-styling at work. The FT-01 → FT-02 loop is closed: same LoRA mechanics, only the memory residency of the weights changes. FT-03 opens multi-task: several adapters, one base model, and the question of rank sharing.
Verifying the “a 7B fits on 6 GB” claim: 4-bit weights = 0.5 byte per parameter, i.e. 3.5 GB for 7B, against 14 GB in FP16. The LoRA adapters and the paged optimizer apply to a small fraction of the parameters. The remaining margin then depends on context length, batch size and activations: this is why the notebook measures the real peak instead of settling for weight arithmetic.