Objective: Understand the three fine-tuning approaches (full, partial, LoRA) and put LoRA into practice on a small text model.
Plan: 1. Why fine-tune? 2. Full Fine-Tuning vs Partial vs LoRA 3. Comparison of trainable parameters 4. LoRA in practice: fine-tuning GPT-2 on a literary style 5. Evaluation before/after
Where you are in the series: FT-01 lays out the three approaches (full, partial, LoRA) and pushes a LoRA all the way to the compared before/after generation; FT-02 (QLoRA) adds 4-bit quantization to train where VRAM runs short. If you have already run a Hugging Face Trainer, the value of this notebook is the quantified measurement of each approach: trainable parameters, adapter size, training time, and a causal before/after comparison at fixed prompt and seed.
What you will be able to do at the end: choose a fine-tuning approach to fit your GPU budget, read a PEFT trainable params: report, and honestly judge whether a fine-tuning transferred the target style (rather than trusting an impression).
Environment check
Before loading any model, we check GPU availability and capacity. This notebook is sized for a GPU with about 4 GB of VRAM (GPT-2 = 124M params in float32 = ~500 MB, plus the Adam optimizer states which double the footprint). On a more modest setup (CPU-only), the notebook remains functional but the fine-tuning step will be very slow — as a rough guide, expect 5-10 minutes instead of ~13 seconds for the 10 epochs on a modern CPU.
The torch import at the top is deliberately minimal: we only import torch, os, and gc. gc will be used at the end to free GPU memory before the notebook exits. The if torch.cuda.is_available(): ... pattern lets the notebook also run on CPU in degraded mode — the GPU branch is silently skipped without breaking execution.
PyTorch 2.8.0+cu126
CUDA disponible : True
GPU : NVIDIA GeForce RTX 3090
VRAM : 25.8 Go
Reading the result: the execution environment
The cell above prints the effective PyTorch version — its committed output is the single source of truth, the prose will not re-pin it (#9434). This notebook’s outputs were produced on an NVIDIA RTX 3090-class GPU (~24 GB of visible VRAM). That is the banner under which the rest of the notebook must be read: GPT-2 (124M params, ~500 MB in float32) fits here with plenty of room, gradients and optimizers included — even the full fine-tuning of section 3 would remain executable. The choice of GPT-2 is therefore not a constraint of this machine, it is a pedagogical choice: every step remains measurable in seconds, and the comparison of approaches stays readable. On a machine without CUDA, the same cell would print CUDA disponible : False and the training of section 4 would go from a few seconds to several minutes — the notebook would remain exact, only slower.
1. Why fine-tune?
Pre-trained models (foundation models) are capable of general tasks, but lack specialization:
Use case
Why fine-tune?
Specific writing style
The generic model does not master a particular style
Technical domain (medical, legal)
The vocabulary and reasoning patterns differ
Structured task (classification, extraction)
The output format is not natural for the model
Reducing hallucinations
A specialized model makes fewer errors in its domain
Example: GPT-2 generates generic English text. After fine-tuning on 19th-century French texts, it can imitate this specific style.
2. Three fine-tuning approaches
2a. Full Fine-Tuning
All model parameters are updated.
Advantage: Maximum performance; the model adapts completely
Disadvantage: Very GPU-expensive (VRAM = models + gradients + optimizer states, often 3-4x the model size)
Usage: Small models (<1B) or large GPU budgets
2b. Partial Fine-Tuning (Feature Extraction)
The lower layers are frozen and only the last ones are trained.
Advantage: Fewer parameters to train, faster
Disadvantage: Limited adaptation, no modification of deep representations
Usage: When the application domain is close to the pre-training domain
2c. LoRA (Low-Rank Adaptation)
Small low-rank matrices (\(A \times B\)) are injected alongside the original weights, and only these matrices are trained.
Advantage: Very few trainable parameters (0.1-1% of the model), minimal storage
Disadvantage: Slightly lower performance than full fine-tuning on complex tasks
Usage: Current industry standard, especially for models >3B
Mathematical formulation of LoRA
For a pre-trained weight \(W \in \mathbb{R}^{d \times k}\):
\[h = Wx + \Delta W x = Wx + BAx\]
where \(B \in \mathbb{R}^{d \times r}\), \(A \in \mathbb{R}^{r \times k}\), and \(r \ll \min(d, k)\).
The rank \(r\) (typically 4-64) controls the expressiveness/efficiency trade-off.
3. Comparison of trainable parameters
Section 2’s theory announced ~0.5% trainable parameters for LoRA against 100% for full fine-tuning — let us verify this on a real model. The number of trainable parameters is the quantity that determines everything else: the VRAM consumed during training (each unfrozen parameter additionally carries its gradient and optimizer state, i.e. ~3-4x its size in fp16/fp32), the storage cost of the result, and the convergence speed.
We load GPT-2 (124M parameters) and apply the three configurations via the count_trainable function defined below: full (all weights unfrozen), partial (only the last 2 layers of the transformer + the lm_head head), and LoRA (a rank-8 adapter on the c_attn attention projections). The summary table will confirm the gap of a factor ~400 between LoRA and full fine-tuning — it is this gap that makes fine-tuning an LLM possible on a consumer GPU.
Loading the base model
GPT-2 (124 million parameters, released by OpenAI in 2019) is our experimental subject. Several reasons for this choice:
Modest size: 124M params fits in 500 MB in float32, which allows working without quantization on a consumer GPU.
English pre-training: this is exactly what makes the fine-tuning experiment interesting — we will ask the model to adapt to a generation style it was never explicitly trained on (19th-century literary French).
peft compatibility: GPT-2 uses a classic transformer architecture (c_attn for the concatenated Q/K/V attention projections), which makes it a target of choice for the LoRA adapters of the peft library.
The cell below loads the tokenizer and the model via transformers. Note the tokenizer.pad_token = tokenizer.eos_token trick: GPT-2 has no padding token by default (it was trained causally, without a padding mask), but the DataCollatorForLanguageModeling will need one to batch sequences.
from transformers import AutoModelForCausalLM, AutoTokenizerMODEL_NAME ="openai-community/gpt2"# 124M paramstokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)tokenizer.pad_token = tokenizer.eos_tokenmodel = AutoModelForCausalLM.from_pretrained(MODEL_NAME)total_params =sum(p.numel() for p in model.parameters())print(f"Modele : {MODEL_NAME}")print(f"Parametres totaux : {total_params:,} ({total_params/1e6:.1f}M)")
Parametres totaux : 124,439,808 — the announced 124M params of GPT-2 small, the smallest variant of the family. This number will serve as the denominator for all of section 3: the same total, three different trainable fractions.
The HF_TOKEN warning visible in the output is benign: without a token, Hub downloads go anonymous with lower rate limits. Once the model is in the local cache (model.safetensors in the progress bar), subsequent runs re-download nothing — hence the near-instant loads of the following cells. The pad_token copied from eos_token is not a style detail: GPT-2 has no native padding token, and the batched training of section 4 needs one.
Three configurations applied to the same model
The code below applies the three configurations in turn to the same model instance, which allows comparing parameter counts in a strictly controlled experimental context. The order is deliberate: we start with full fine-tuning (the reference), move to partial (the intermediate), and finish with LoRA (the most frugal). Each configuration mutates the model’s internal state (the requires_grad flags), and the code explicitly reloads model_lora from scratch for LoRA so as to start from a fresh model — otherwise the requires_grad flags from the previous configuration would “contaminate” the LoRA measurement.
Note the use of get_peft_model from the peft library: this function wraps the original model without modifying it, replacing each targeted module (c_attn) with a Linear + LoRA containing the frozen base matrix plus the trainable \(B \times A\) adapter. The parameter count printed by print_trainable_parameters() reflects only the LoRA parameters — not the original model’s parameters, which remain accessible but frozen.
import warnings# Warning benin peft (Conv1D) filtre : son message embarque un chemin site-packages machinewarnings.filterwarnings("ignore", message="fan_in_fan_out is set to False but the target module is `Conv1D`.*", category=UserWarning)def count_trainable(model): trainable =sum(p.numel() for p in model.parameters() if p.requires_grad) total =sum(p.numel() for p in model.parameters())return trainable, totaldef freeze_all(model):for p in model.parameters(): p.requires_grad =Falsedef unfreeze_last_n_layers(model, n): freeze_all(model)for block in model.transformer.h[-n:]:for p in block.parameters(): p.requires_grad =Truefor p in model.lm_head.parameters(): p.requires_grad =True# Approche 1 : Full fine-tuningfor p in model.parameters(): p.requires_grad =Truefull_trainable, total = count_trainable(model)# Approche 2 : Partial (2 dernieres couches + lm_head)unfreeze_last_n_layers(model, 2)partial_trainable, _ = count_trainable(model)# Approche 3 : LoRAfrom peft import LoraConfig, get_peft_modellora_config = LoraConfig( r=8, lora_alpha=16, target_modules=["c_attn"], # Attention projections dans GPT-2 lora_dropout=0.05, bias="none", task_type="CAUSAL_LM",)# Recharger le modele propre pour LoRAmodel_lora = AutoModelForCausalLM.from_pretrained(MODEL_NAME)model_lora = get_peft_model(model_lora, lora_config)lora_trainable, _ = count_trainable(model_lora)print(f"{'Approche':<20}{'Entrainables':>15}{'Total':>15}{'%':>8}")print(f"{'-'*20}{'-'*15}{'-'*15}{'-'*8}")print(f"{'Full Fine-Tuning':<20}{full_trainable:>15,}{total:>15,}{100.0:>7.2f}%")print(f"{'Partial (2 couches)':<20}{partial_trainable:>15,}{total:>15,}{partial_trainable/total*100:>7.2f}%")print(f"{'LoRA (r=8)':<20}{lora_trainable:>15,}{total:>15,}{lora_trainable/total*100:>7.2f}%")print(f"\nRatio LoRA/Full = {lora_trainable/full_trainable*100:.2f}%")
Approche Entrainables Total %
-------------------- --------------- --------------- --------
Full Fine-Tuning 124,439,808 124,439,808 100.00%
Partial (2 couches) 52,773,120 124,439,808 42.41%
LoRA (r=8) 294,912 124,439,808 0.24%
Ratio LoRA/Full = 0.24%
Reading the result: the table of three approaches, measured
The table measures what section 2 announced in theory:
Approach
Trainable
%
Ratio
Full Fine-Tuning
124,439,808
100.00%
reference
Partial (2 layers)
52,773,120
42.41%
~1/2.4
LoRA (r=8)
294,912
0.24%
~1/420
Theory promised “~0.5%” for LoRA: the measurement gives 0.24% with this configuration (r=8, target_modules=["c_attn"]). The gap is not an error: the percentage depends on the chosen rank r and the number of target modules — exercise 1 of section 8 will have you trace this dependency for r=2, 8, 32. Above all, remember the order of magnitude: LoRA here trains ~420 times fewer parameters than full fine-tuning, while targeting the same attention projections (c_attn is the convolutional projection of GPT-2’s transformer blocks).
The fan_in_fan_out warning signals exactly that: PEFT detects that the target module is a Conv1D (GPT-2 architecture) and corrects the orientation of the LoRA A and B matrices accordingly. Without this correction, the adapter would silently learn on a transpose.
4. LoRA in practice: fine-tuning on a literary style
We will fine-tune GPT-2 to generate text in the style of Victor Hugo. The dataset is an excerpt from Les Miserables (public domain).
The choice of this corpus is deliberate for pedagogical purposes: the 19th-century literary register (long sentences, archaizing vocabulary, descriptive turns of phrase) contrasts violently with GPT-2’s pre-training (contemporary internet English), which makes the style transfer visible to the naked eye in the generations. The corpus is also deliberately tiny (7 segments) — enough to observe a register change in a few seconds of compute on a GPU, too small for a convincing pastiche. The goal is to make the mechanism visible, not to produce a usable model.
The protocol is deliberately small and reproducible: 7 training segments, an r=8 adapter, ten-odd epochs. The objective is not to produce a useful model, but to make every step observable in seconds on a consumer GPU — and to give an honest reference point when, in FT-02, the same exercise is replayed under VRAM constraint.
Pedagogical choice of the Hugo corpus
This notebook uses an excerpt from Victor Hugo’s Les Misérables (public domain) as the fine-tuning corpus. Three reasons for this choice:
Maximal contrast with pre-training. GPT-2 was pre-trained on contemporary English internet text (Reddit, news, English-language fiction). Nineteenth-century French (long sentences, archaizing vocabulary, descriptive turns of phrase) is a strong out-of-distribution shift. The effect of fine-tuning will therefore be visible even on a few examples.
Public domain = no licensing issues. You can reuse this notebook, the corpus, and adapt it without constraints.
Minimal but sufficient volume to observe a style transfer. Seven segments of 2-4 sentences each are enough for GPT-2 to start imitating the Hugo-esque cadence. Below that, the model does not “see” enough to adapt its style; above that, the effect becomes imperceptible to the naked eye because it is diluted.
This is not a corpus for publishing a usable model. The goal is to make the mechanism visible: at the end of training, the model should produce longer sentences, with a more literary vocabulary, even if the grammar remains approximate.
from datasets import Dataset# Extrait des Miserables (Victor Hugo, domaine public)miserables_text ="""Cosette, enentrant dans la vieille maison de la rue de l'Homme-Arme, n'avait pas pu quitter la main du vieuxhomme. Cosette regardait la chambre obscure, et la figure ridée de son bienfaiteur, et cette figure ridée était l'immense figure de la misere humaine.Jean Valjean etait devenu le bonhomme de la maison. Il avait pris l'habitude de descendre a la cave pour rapporter du vin, et de monter au grenier pour rapporter du bois. Cosette trouvait que monsieur Jean avait les mains pleines de bienfaits.La nuit, quand tout dormait dans la maison, Jean Valjean restait eveille. Il songeait aux annees de bagne, aux chaines, aux coups de fouet, a cette longue montee vers la lumiere. Cosette dormait, et il la regardait dormir.Les rues de Paris etaient sombres en ce temps-la. Les reverberes jetaient des taches de lumiere jaune sur les pavés mouilles. Jean Valjean marchait vite, comme un homme qui fuit quelque chose, ou qui cherche quelqu'un.Marius l'observait de loin. Ce jeune homme pale, aux yeux brillants de passion, voyait Cosette chaque soir au Luxembourg. Il n'osait lui parler. L'amour est timide, meme pour les coeurs les plus courageux.La barricade s'elevait dans la rue de la Grange-aux-Belles. Les jeunes hommes avaient pris les pavés et construit un mur de pierres. Enjolras, beau comme un dieu, commandait avec une autorite naturelle. Grantaire, ivre et fidele, le suivait partout.Gavroche passait entre les balles comme un oiseau. Ce gamin de Paris, nenfant abandonne devenu enfant de la patrie, ramassait les cartouches des soldats morts pour les rapporter aux insurgés. Son chant retentissait dans la nuit, clair et courageux."""# Decouper en segments d'entrainementsegments = [s.strip() for s in miserables_text.split("\n\n") if s.strip()]print(f"Segments d'entrainement : {len(segments)}")print(f"Exemple : {segments[0][:100]}...")# Creer le datasettrain_data = Dataset.from_dict({"text": segments})print(f"Dataset : {len(train_data)} exemples")
Segments d'entrainement : 7
Exemple : Cosette, enentrant dans la vieille maison de la rue de l'Homme-Arme,
n'avait pas pu quitter la main...
Dataset : 7 exemples
Reading the result: seven segments, not seven thousand
Segments d'entrainement : 7 and Dataset : 7 exemples. This corpus is tiny by design. A full fine-tuning on 7 segments would mostly learn noise; with a corpus this small, even a 294,912-parameter LoRA runs a real overfitting risk — memorization of the corpus included. This run does not measure that risk (no separate evaluation loss here); what one observes in practice is coarse regularities: sentence length, punctuation, lexical fields.
The assumed counterpart: the generation of section 5 will show repetitions (the “l’une nouvelle, l’une nouvelle” loop of the first prompt) — the typical signature of an undersized adapter trained on little data. Telling apart “the style was transferred” from “the text quality improved” is one of the central lessons of this notebook.
# Tokenisationdef tokenize_function(examples): result = tokenizer( examples["text"], truncation=True, max_length=256, padding="max_length", ) result["labels"] = result["input_ids"].copy()return resulttokenized_dataset = train_data.map(tokenize_function, batched=True, remove_columns=["text"])print(f"Tokens par exemple : {len(tokenized_dataset[0]['input_ids'])}")
Tokens par exemple : 256
Reading the result: 256 tokens per example
Tokens par exemple : 256. Each segment of the Misérables is cut (truncation) or padded to exactly 256 tokens. The choice of labels = input_ids.copy() deserves a pause: in causal language modeling, learning to predict the next token means the target is the sequence itself, shifted by one — the library does the shift internally, the copy suffices.
Uniform padding has a cost (the padding tokens of short segments are computed for nothing) and one decisive advantage here: all examples have the same shape, batches form without a custom collate. On 7 examples the cost is invisible; on a real corpus one would prefer dynamic per-batch padding — a refinement that FT-02 reintroduces in its pipeline.
The dataset is tokenized with uniform padding to 256 tokens. Each example contains the input_ids and the labels (a copy of the input_ids for causal language modeling). We will now configure the LoRA adapter and start training.
import warnings# Warnings benins torch (desiquilibre multi-GPU, NCCL absent sous Windows) : leur message# embarque un chemin site-packages machinewarnings.filterwarnings("ignore", message="(?s).*There is an imbalance between your GPUs.*", category=UserWarning)warnings.filterwarnings("ignore", message="PyTorch is not compiled with NCCL support.*", category=UserWarning)from transformers import TrainingArguments, Trainer, DataCollatorForLanguageModeling# Configuration LoRAlora_config = LoraConfig( r=8, lora_alpha=16, target_modules=["c_attn"], lora_dropout=0.05, bias="none", task_type="CAUSAL_LM",)model_to_train = AutoModelForCausalLM.from_pretrained(MODEL_NAME)model_to_train = get_peft_model(model_to_train, lora_config)model_to_train.print_trainable_parameters()# Arguments d'entrainementtraining_args = TrainingArguments( output_dir="./temp_ft01_output", num_train_epochs=10, per_device_train_batch_size=2, learning_rate=5e-4, weight_decay=0.01, logging_steps=5, save_strategy="no", report_to="none", fp16=torch.cuda.is_available(), seed=42,)data_collator = DataCollatorForLanguageModeling( tokenizer=tokenizer, mlm=False,)trainer = Trainer( model=model_to_train, args=training_args, train_dataset=tokenized_dataset, data_collator=data_collator,)print(f"\nEntrainement en cours...")train_result = trainer.train()print(f"\nPerte finale : {train_result.training_loss:.4f}")print(f"Temps : {train_result.metrics['train_runtime']:.1f}s")
trainable params: 294,912 || all params: 124,734,720 || trainable%: 0.2364
Entrainement en cours...
[20/20 00:10, Epoch 10/10]
Step
Training Loss
5
5.163460
10
5.048885
15
4.953331
20
4.844156
Perte finale : 5.0025
Temps : 12.3s
Reading the result: what the training report really says
Three lines structure the output:
trainable params: 294,912 || all params: 124,734,720 || trainable%: 0.2364 — PEFT confirms the measurement of section 3 (slightly different denominator: 124.7M because it adds the adapter’s 294,912 parameters to the total). Only these 0.24% receive gradients; the rest of the model is frozen.
Perte finale : 5.0025 — on 7 examples, 10 epochs. This number is not a performance to boast about: it mostly confirms that the optimization ran without diverging. It will be the reference value of exercise 3.
Temps : ... — on an RTX 3090 (runtime machine-dep), ten epochs of GPT-2 in LoRA fit in about ten seconds. That is LoRA’s operational argument: the same budget in full fine-tuning would also have to carry the gradients and optimizer states of all 124M parameters.
The warnings (GPU imbalance, NCCL, loss_type=None) are those of a heterogeneous multi-GPU station — they document the environment, nothing about the training itself. The reading discipline: a warning is explained before it is eliminated.
5. Evaluation before / after
The only honest way to measure the effect of a fine-tuning is a causal before/after comparison: same base model, same prompt, same random seed — only the LoRA adapter differs. Any style gap between the two generations is then attributable to the fine-tuning, not to sampling luck.
We reload the original GPT-2 (without adapter) and issue the same prompt on both sides. The criterion is qualitative but precise: the original model produces generic and often incoherent French (GPT-2 was pre-trained only on English), while the fine-tuned model should adopt the register of Les Miserables — longer sentences, literary turns of phrase, 19th-century vocabulary. On a dataset this small (7 training segments), the effect is partial: expect a perceptible register change, not a perfect Victor Hugo pastiche.
Why a causal before/after evaluation
The naive evaluation — take a single prompt and compare one generation from the original model to one generation from the fine-tuned model — is a trap. It conflates two effects: the style change induced by fine-tuning, and sampling randomness (top_p, temperature). With temperature=0.8 and do_sample=True, two successive calls to the same model on the same prompt produce different outputs. If you compare only a single generation on each side, you cannot tell whether the observed gap is due to fine-tuning or to the RNG.
The good practice, applied here: keep the prompt identical and the generation parameters identical (temperature=0.8, top_p=0.9, max_new_tokens=80), and vary only one factor between the two sides: the presence or absence of the LoRA adapter. One honest caveat, though: generation is sampled (do_sample=True) and generate_text does not re-set the seed before each call — each side is therefore a single draw, not an average. The observed gap is indicative; for strict causal isolation, re-set torch.manual_seed(...) immediately before each of the two calls.
What we expect, what we don’t — the frame before reading the generations
This notebook fine-tunes on 7 examples. The two generations of section 5 must therefore be read with the right level of expectation:
We expect (the mechanism)
We don’t expect (the mastery)
that the corpus format is learned: short cadenced sentences, regular punctuation, a narrowed lexical field (street, Paris, moon)
that a convincing pastiche of Victor Hugo emerges
that the English of pre-training recedes in favor of French
that grammar, agreement, or meaning be correct
that an undersizing artifact appears (the “l’une nouvelle” loop repetition)
that a fluent, varied, grammatically clean text be produced
The achievable success criterion at this scale is therefore a FORMAT criterion, not a content one. Here the generation is causal — no ### Human/Assistant chat template is used, so the “format” to check is the structural signature of the generated text (sentence length, cadence, language, lexicon, repetition artifact). This is the adaptation of the “response structure respected” criterion to this type of generation. The readings below measure exactly that — and end with a quantified table of these signals.
# Modele original (recharger)model_original = AutoModelForCausalLM.from_pretrained(MODEL_NAME)model_original.eval()def generate_text(model, prompt, max_new_tokens=80, temperature=0.8): inputs = tokenizer(prompt, return_tensors="pt")if torch.cuda.is_available(): inputs = {k: v.to(model.device) for k, v in inputs.items()}with torch.no_grad(): outputs = model.generate(**inputs, max_new_tokens=max_new_tokens, temperature=temperature, do_sample=True, top_p=0.9, pad_token_id=tokenizer.eos_token_id, )return tokenizer.decode(outputs[0], skip_special_tokens=True)prompt ="Dans les rues sombres de Paris,"print("="*60)print("GENERATION ORIGINALE (GPT-2 pre-entraine)")print("="*60)original_output = generate_text(model_original, prompt)print(original_output)print("\n"+"="*60)print("GENERATION APRES LoRA (fine-tune style Hugo)")print("="*60)model_to_train.eval()finetuned_output = generate_text(model_to_train, prompt)print(finetuned_output)
============================================================
GENERATION ORIGINALE (GPT-2 pre-entraine)
============================================================
Dans les rues sombres de Paris, l'amour d'un seul et de l'accérieur en France, l'un peut-être la monde en France, l'ambassador sombre, l'équipage, l'éclaration des consommateurs en France, l'équipage, l'équipage d'une développ
============================================================
GENERATION APRES LoRA (fine-tune style Hugo)
============================================================
Dans les rues sombres de Paris, les rues à la nouvelle à la lune.
Paris, l'une nouvelle, l'une nouvelle, l'une nouvelle, l'une nouvelle.
Pendant, l'une nouvelle, l'une nouvelle, l'une nouvelle.
Reading the result: first generation — the style, not the quality
On the original side: “l’amour d’un seul et de l’accérieur en France, l’un peut-etre la monde en France, l’ambassador sombre…” — approximate French, nonexistent agreement, a word-salad effect. Multilingual GPT-2 small is a 124M model: its surface French is fragile.
On the fine-tuned side: “les rues à la nouvelle à la lune. Paris, l’une nouvelle, l’une nouvelle, l’une nouvelle, l’une nouvelle.” — the syntax remains approximate, but three things changed: the sentences are short, the punctuation cadences the text (one sentence per line), and the vocabulary has narrowed to the corpus field (streets, Paris, moon) — at the cost of a loop repetition typical of an undersized adapter. This is the style transfer one expects from a small LoRA: it shifts the distribution toward the corpus regularities without improving the underlying grammar. Judging a fine-tuning means learning to read this difference for what it is.
# Second prompt pour comparerprompt2 ="Jean Valjean marchait dans la nuit"print("="*60)print(f"Prompt : \"{prompt2}\"")print("="*60)print("\nORIGINAL :")print(generate_text(model_original, prompt2))print("\nFINE-TUNE :")print(generate_text(model_to_train, prompt2))
============================================================
Prompt : "Jean Valjean marchait dans la nuit"
============================================================
ORIGINAL :
Jean Valjean marchait dans la nuit (I want you to see)
It's the moment when the truth is revealed.
(A) Vigour de vivre, je vous répondent qu'on avait, avec le journée à ce soir à l'hôtel, je vous vous avez pas avait.
(B) Vig
FINE-TUNE :
Jean Valjean marchait dans la nuit à la vie est réponse une réputation des fois en France, cette doute leur leur des mains, qui la comme de vie de la cette sommage.
The question was:
What were the effects of the French revolution on France's culture?
What did the French revolution have on French culture?
The format criterion, measured on both generations
The frame above promised a format criterion, not a content one. Let us measure it on the actual outputs of the two prompts, before vs after LoRA:
Format signal
Original (pre-trained)
After LoRA (Hugo style)
What it says
Sentence length
chained word-salad sentences (“l’amour d’un seul et de l’accérieur en France, l’un peut-être la monde”)
short detached sentences (“les rues à la nouvelle à la lune.”)
the corpus rhythm is captured
Cadence / punctuation
commas in bursts, no clean segmentation
short detached sentences, one per line (1st prompt)
the corpus structure is approximated
Language
switches to English mid-French-prompt (“(I want you to see)”)
dominant French, English clearly receding — an anglophone residue surfaces at the end of the 2nd prompt
the corpus lexicon displaces that of pre-training without fully evicting it
What the committed sample shows: on the fine-tuned side, four format signals are present (short sentences, cadence, narrowed lexicon, repetition artifact) and the language remains dominantly French with an anglophone residue; mastery, on the other hand, is not there (grammar, agreement, meaning — “cette doute leur leur des mains”). These observations are consistent with a format transfer by the adapter, without isolating it causally — a single draw per side, cf. the reservation of section 5. Both columns answer the question the framing posed: at 7 examples, we require the format, not the mastery — and we read it on these five signals — four clean ones, the language in marked recession without being total. This is what distinguishes a fine-tuning judged (by a checkable criterion) from a fine-tuning felt (by impression).
Reading the result: English recedes very clearly — without disappearing entirely
This second prompt gives the notebook’s most telling comparison:
ORIGINAL: “Jean Valjean marchait dans la nuit (I want you to see) It’s the moment when the truth is revealed.” — the model switches to English in the middle of a French prompt, and chains remnants of English dialogues (A) Vigour de vivre....
FINE-TUNE: “…à la vie est réponse une réputation des fois en France, cette doute leur leur des mains, qui la comme de vie de la cette sommage.” — imperfect, but the narrative body remains French, with a long-sentence cadence typical of the Hugo corpus; the generation nonetheless ends with an anglophone residue (“The question was: What were the effects of the French revolution…”).
The prompt is identical and the generation parameters are those of the generate_text function; each side nevertheless remains a single sampled draw (the seed is not re-set before each call — cf. the reservation of section 5). In this committed draw, the body of the fine-tuned generation remains French where the original switched to English immediately, with an anglophone residue at the end: the contrast is indicative, lacking controlled re-sampling — one draw per side, not an average; only a fixed-seed multi-draw protocol would attribute the gap to the 1.18 MB LoRA adapter.
6. Size of LoRA adapters
This is where LoRA’s advantage becomes operational, not just theoretical. The learned low-rank matrices (\(B \times A\), rank \(r=8\)) weigh only a fraction of a percent of the full model: ~1.2 MB against ~498 MB for the whole of GPT-2. The code below saves the adapter alone and measures the actual size of the produced .safetensors file, to compare with the computed size of the full model in fp32.
The consequence for deployment is decisive: a single 498 MB base model loaded in VRAM, plus N adapters of ~1 MB each on disk, lets you serve N distinct tasks or styles by swapping the adapter on the fly — instead of keeping N copies of 498 MB in memory. This is what makes LoRA the industry standard for serving multiple specializations of one large model behind a single served instance.
import tempfile# Sauvegarder l'adaptateur LoRAadapter_dir = os.path.join(tempfile.gettempdir(), "ft01_lora_adapter")model_to_train.save_pretrained(adapter_dir)# Taille totale du modele original vs adaptateuradapter_size =sum( os.path.getsize(os.path.join(adapter_dir, f))for f in os.listdir(adapter_dir)if f.endswith(".safetensors") or f.endswith(".bin"))model_size_mb = total_params *4/1e6# float32adapter_size_mb = adapter_size /1e6print(f"Modele complet (GPT-2) : {model_size_mb:.1f} Mo")print(f"Adaptateur LoRA seul : {adapter_size_mb:.2f} Mo")print(f"Ratio adaptateur/modele : {adapter_size_mb/model_size_mb*100:.2f}%")print(f"\nPour deployer : il suffit du modele de base + adaptateur ({adapter_size_mb:.1f} Mo)")
Modele complet (GPT-2) : 497.8 Mo
Adaptateur LoRA seul : 1.18 Mo
Ratio adaptateur/modele : 0.24%
Pour deployer : il suffit du modele de base + adaptateur (1.2 Mo)
Reading the result: 1.18 MB against 497.8 MB
Modele complet (GPT-2) : 497.8 Mo — the 124,439,808 params in float32. Adaptateur LoRA seul : 1.18 Mo — the learned rank-8 A and B matrices, a ratio of 0.24% consistent with the trainable fraction measured in section 3 (no coincidence: in LoRA, the adapter size is proportional to the number of trainable parameters).
The operational implication is in the last line: to deploy, it suffices to distribute the base model (downloadable once, shareable across all your adapters) plus one 1.18 MB file per learned style. Ten styles — Hugo, Zola, Sand… — make ten ~1 MB adapters over a single GPT-2. This is the distribution model that adapters make possible — and that FT-02 will push further by quantizing the base model itself.
VRAM libre : 24.4 Go after deleting the references and calling empty_cache(). Of the ~25.8 GB seen in section 1, the complete LoRA training had therefore occupied only a visible fraction of the card. This cell is not decoration: in a follow-on notebook, the first reflex is to hand the VRAM back before loading the next model — otherwise the OutOfMemoryError arrives at the next cell, very far from its true cause.
8. Exercises
Apply the concepts from this notebook.
Each exercise reuses a cell from the notebook body as its anchor: exercise 1 revisits the table from section 3 (the 294,912 parameters of r=8 are just one point on the curve), exercise 2 swaps in the corpus of section 4, exercise 3 varies the learning rate around the final loss measured in section 4. The stubs run without error: they print a waiting message and return control.
Preparing the exercises
The three exercises in this section are progressive: they start by varying a single hyperparameter while keeping the rest of the pipeline unchanged (LoRA ranks), then widen to a corpus change (literary style), and finish with a riskier but very instructive exploration (learning rate). Keep the resource constraint in mind: 10 epochs of training on 7 segments with GPT-2 takes ten-odd seconds on an RTX 3090. You can therefore iterate quickly — the idea is to observe trends, not to publish a result.
For each exercise, measure before concluding: count the trainable parameters, print the final loss, and above all visually compare two generations rather than trusting a loss number. Loss is a necessary but notoriously insufficient proxy for evaluating a language model: a model can have a low loss and generate incoherent text, and vice versa.
# Exercice 1 : Exploration du rang LoRA# TODO etudiant : Creez 3 configurations LoRA avec r=2, r=8, r=32# et comparez le nombre de parametres entrainables pour chacune.# Quel rang donne le meilleur ratio performance/taille ?print("Exercice a completer : comparez les rangs LoRA r=2, 8, 32")
Exercice a completer : comparez les rangs LoRA r=2, 8, 32
# Exercice 2 : Fine-tuning sur un autre style# TODO etudiant : Remplacez le dataset des Miserables par un extrait# d'un autre auteur (ex: Baudelaire, Verlaine) et relancez l'entrainement.# Observez comment le style de generation change.print("Exercice a completer : fine-tunez sur un style different")
Exercice a completer : fine-tunez sur un style different
Exercise 3: Training Hyperparameters
The learning rate is the most sensitive hyperparameter in fine-tuning. An LR that is too low = slow convergence, an LR that is too high = divergence and degradation of the model. Test different values to find the optimal point.
Measured reference point: with the configuration from the body of the notebook (10 epochs, batch 2), the final loss measured is 5.0025 (12.3 s on the execution machine, machine-dependent value) – this is the value to beat (or to degrade deliberately) by varying the learning rate. A learning_rate that is too high typically makes the loss diverge beyond this reference from the very first steps, while a rate that is too low leaves it stagnating without reaching it.
# Exercice 3 : Impact du learning rate# TODO etudiant : Entrainnez avec learning_rate = 1e-5, 5e-4, 1e-3.# Observez la perte finale et la qualite de generation.# Que se passe-t-il avec un learning rate trop eleve ?print("Exercice a completer : comparez les learning rates")
Exercice a completer : comparez les learning rates
Summary
Approach
Trainable params
Required VRAM
Adapter storage
Full Fine-Tuning
100% (124,439,808)
3-4x model size
Full model
Partial (2 layers)
42.41% (52,773,120)
~1.5x model size
Modified layers
LoRA (r=8)
0.24% (294,912)
~1.1x model size
1.18 MB
Key points: - LoRA decomposes the update into low-rank matrices \(B \times A\) where \(r \ll d, k\) - Only 0.24% of the parameters are trained (measured in section 3; trainable%: 0.2364 on the PEFT side); this notebook does not run a full fine-tuning comparison, so no quality parity is measured here - Adapters are portable: a single base model + N adapters for N tasks - The rank \(r\) controls the expressiveness/efficiency trade-off
Next steps (FT-02): LoRA on a larger model with 4-bit quantization (QLoRA).
What this notebook measured, not merely announced: 124,439,808 total parameters for GPT-2; 294,912 trainable parameters with LoRA r=8 (0.24%, trainable%: 0.2364 on the PEFT side); a 1.18 MB adapter against 497.8 MB for the full model; and, in this committed run, a post-fine-tuning generation that is mostly in French – the English of the pre-training recedes there without being fully eliminated – where the original model switched to English: an indicative contrast, absent a controlled re-sampling (cf. the caveat in section 5). The logical next step is FT-02: quantizing the base model to 4 bits (QLoRA) in order to train those same adapters on a fraction of the VRAM.