Catastrophic forgetting in continual fine-tuning: a model cube trained over three rounds loses older skill tiles as new ones arrive, beside four techniques that reduce it: a LoRA adapter, a sliding window, rehearsal data and a KL anchor to the base model
Each round of fine-tuning can erase older skills

Key Takeaways

  • If you keep fine-tuning a model on new data, it gradually loses what it could already do. That is catastrophic forgetting, and it is the main reason continual learning is harder than running SFT on a schedule.
  • It happens because new gradient updates overwrite weights that also encoded older skills. Narrow or self-generated data makes it worse: in our experiment, a model retrained only on its best runs collapsed.
  • What helps: train LoRA adapters instead of the full model, restart from the base model each round, train on a sliding window of recent data, mix in rehearsal data, or anchor the loss to the base model with a KL term.
  • None of these eliminates forgetting. Held-out evaluations of older skills, round after round, and the ability to roll back are what keep it in check.

TL;DR

You can keep fine-tuning, but not naively. Each round pulls the weights towards the newest data and away from what the model already knew. LoRA adapters, which forget less than full fine-tuning, sliding windows, rehearsal and KL anchors all reduce the damage, and NSL combines several of them. Measure older skills every round and keep the previous adapter so you can roll back.

Why Can’t You Just Keep Fine-Tuning?

Once you accept that AI agent memory isn’t learning, the obvious way to make a model learn continuously is to fine-tune it again whenever new data arrives: this week’s successful agent runs, next week’s, and so on. Each round, the model gets better at the newest data. It also gets worse at things it used to do well, including general abilities nobody was trying to change.

This is not an edge case. An empirical study of continual instruction tuning found forgetting across LLMs from 1B to 7B parameters, and it grew more severe as models grew within that range.

Training is also non-linear. A model can get worse before it gets better, as grokking shows, so every extra round is another chance to break something that worked. The less of the model each round touches, the less it can break: the principle behind most continual learning methods.

What Is Catastrophic Forgetting?

Catastrophic forgetting is the loss of previously learned abilities when a neural network is trained on new data. Gradient updates for the new task change weights that also encoded the old ones, and without anything protecting those weights, performance on earlier tasks can drop sharply rather than fade gradually.

McCloskey and Cohen described it in connectionist networks in 1989. Most later fixes take one of three routes: protect important weights, train fewer weights, or keep old data in the mix.

Why Continual Fine-Tuning Forgets

Three mechanisms drive forgetting in a continual fine-tuning loop, and a fourth point is often missed: some forgetting helps.

Every update moves shared weights

In full fine-tuning, every weight can move. The skills a model already has are spread across those same weights, so optimising for this week’s data shifts parts of the network that last month’s skills depended on.

Narrow data narrows the model

When a model trains on its own output, each round’s data reflects what the previous model already did well. In our self-training experiment, three lineages of Qwen 2.5 7B were retrained on their own successful runs. The lineage trained only on its best three runs collapsed, the one trained on all past data plateaued, and the one trained on a sliding window of recent runs kept improving.

The newest data dominates

Rows from the latest round usually outnumber anything kept from earlier rounds, and a training run treats them alike. Without a deliberate mix, the model is pulled wherever its most recent data points.

Some forgetting is useful

Forgetting everything old is the failure, but forgetting some of it is not. In the same experiment, moderate forgetting kept the model from overfitting to a narrow niche: strategies that stopped working aged out, and strategies that kept working stayed. The goal is controlled forgetting, not none.

Ways to Prevent Catastrophic Forgetting, and Their Limits

No single technique prevents it. These reduce it, and they combine:

TechniqueHow it limits forgettingCostIn NSL
LoRA adaptersThe base weights stay frozen; only a small low-rank update is trainedLearns less than full fine-tuningAlways: QLoRA on a 4-bit base model
Restart from the base each roundDrift cannot compound across roundsAn adapter only knows what its training data showsEvery training step starts a fresh adapter
Sliding window of recent dataOld rows age out gradually instead of all at onceCan drop capabilities whose data stopped appearingtraining_window_size, default 3
Rehearsal dataGeneral text in the mix keeps general abilities exercisedMore rows, more training timerehearsal_dataset, off by default
KL anchor to the base modelPenalises moving away from the base model’s predictionsExtra reference pass and memoryASFT with asft_kl_weight above 0
Weight-importance penalties (EWC)Slows learning on weights important for old tasksNeeds importance estimates per taskNot used

LoRA adapters: train less, forget less

A LoRA adapter trains a small low-rank update while the base weights stay frozen. LoRA Learns Less and Forgets Less found exactly that trade-off. LoRA underperformed full fine-tuning on the target domain but better maintained the base model’s performance outside it, and it mitigated forgetting more than weight decay or dropout.

Restarting from the base model

NSL never stacks adapters. Each training step loads the base model and attaches a fresh adapter, so a bad round cannot become the starting point for the next one. Earlier rounds influence the new adapter only through the data in its window.

Sliding windows and rehearsal

A sliding window keeps the last few rounds of data and drops older ones. In NSL, training_window_size (default 3) sets how many generations feed each training step.

Rehearsal mixes general data back in. When rehearsal_dataset names a text dataset and rehearsal_rows_per_epoch is above 0, NSL’s trainer adds rows that ask the model to continue a passage, so general writing keeps being exercised alongside the task.

Anchoring the loss to the base model

Anchored supervised fine-tuning (ASFT) adds a KL term that penalises the fine-tuned model for drifting from a reference model. In NSL the reference is the base model with the adapter switched off. The anchor only acts with a non-zero asft_kl_weight; at the default of 0, ASFT reduces to plain DFT. RL pipelines such as RLHF add a similar KL penalty towards a reference model, for the same reason.

Elastic weight consolidation takes a different route. EWC slows learning on the weights that matter most for earlier tasks, which needs an estimate of that importance for each task.

How NSL Combines These Techniques

NSL’s continuous learning loop uses several layers of protection at once. It trains QLoRA adapters only, starts each training step from the base model, and trains on a sliding window of the last three generations. Rehearsal data and the ASFT anchor are available when a task needs them.

The combination reflects the trade-off rather than solving it. The window is a choice about how much to forget; the adapter limits how far each round can move; the base-model restart stops errors from compounding.

How to Measure Catastrophic Forgetting in Your Own Loop

For AI researchers, forgetting in continual fine-tuning is an open trade-off rather than a solved problem. How much a model should forget is itself the research question. With self-generated data it also connects to the wider concern that training on generated data makes the tails of the original distribution disappear. Each technique above is a variable you can test.

What to measure after every round:

  • Old-skill retention: a fixed held-out set from earlier tasks, plus a general benchmark.
  • New-task gain: the improvement the round was meant to buy.
  • Output diversity: narrowing outputs are an early sign of collapse, and LoRA’s authors report that it helps keep generations diverse.
  • Rollbacks: how often a new adapter is worse than the previous one.

Compare history policies the way our paper did: cumulative, sliding window and best-only. Then vary one lever at a time, so each effect stays attributable: training_window_size, lora_rank (default 32), rehearsal_rows_per_epoch and asft_kl_weight.

What to Watch For

  • None of these eliminates catastrophic forgetting. Evaluate older skills on held-out tasks after every round, the same discipline that catches model drift, not just the newest task.
  • Keep every round’s adapter. In NSL each after_generation_<N> adapter stays on disk, so rolling back means serving an earlier one.
  • LoRA forgets less partly because it learns less. If the new task needs a large change, a small adapter may not capture it.
  • A window that is too short overfits recent behaviour; one that is too long lets stale data dominate. Tune it against held-out results.
  • Rehearsal rows longer than max_seq_length are dropped with only a warning, so check how many survived.

FAQs

Does LoRA prevent catastrophic forgetting?

It reduces it but doesn’t prevent it. Because the base weights stay frozen, LoRA better maintains a model’s performance outside the target domain than full fine-tuning, and it mitigates forgetting more than weight decay or dropout. The price is that it also learns less, so large changes may need more than a small adapter.

How do you prevent catastrophic forgetting in LLMs?

Combine several partial defences rather than relying on one. Train LoRA adapters instead of the full model, restart from the base model each round, keep a window or rehearsal mix of older data, and add a KL penalty towards the base model if drift is the main risk. Then evaluate older skills after every round.

Is catastrophic forgetting always bad?

No. Forgetting everything old is a failure, but some forgetting helps a model that keeps learning. In our self-training experiment, moderate forgetting from a sliding window kept the model from overfitting to a narrow niche, while training on only the best runs collapsed. The aim is controlled forgetting.

How many times can I fine-tune a model before it forgets?

There is no fixed number. It depends on how different the new data is, how much of the model each round changes, and what protects the older skills. Measure rather than guess: keep a held-out set for the abilities you care about, evaluate it after every round, and roll back when it drops.