
Key Takeaways
- A research agent that works in one lab, codebase or pipeline meets the same problems every day. Because its weights are frozen, it never gets better at them. A continual learning LoRA setup gives it a weight update every night from the day’s successful runs, with no RL pipeline or training cluster.
- Each night, a fresh LoRA adapter trains from the base model on the last few days of accepted traces, and the agent starts the next day with those successes in its weights. A better adapter solves more tasks, which gives the next night more to learn from.
- It is hacky: adapters can’t be stacked forever, each one is retrained from scratch, and in NSL a new adapter means restarting the inference server. Still, it is cheap, and in our experiment the sliding-window lineage kept improving where cumulative retraining plateaued.
- Measure the improvement after every training round on held-out tasks, and include older skills, because a window forgets as well as learns.
TL;DR
Use nightly LoRAs when a research agent repeats similar work and you can tell its good runs from its failures. They suit teams with no reward to optimise and no cluster to spare. Each night, our modified SFT trains a small adapter on a sliding window of recent successes while the base model stays frozen, so each new round of work runs on what the agent learned from the last. Continual learning LoRA is not elegant, but it is the cheapest way we know to give an agent real weight-level learning.
Why Would a Research Team Want Nightly LoRAs?
Research agents live in narrow environments: one lab’s instruments, one codebase, one data pipeline. They hit the same errors and the same conventions every day. Memory files and retrieval can carry facts forward, but the habits that caused yesterday’s failures stay in frozen weights.
A nightly adapter is worth trying when:
- the work repeats: similar tasks every day, in the same environment;
- success is checkable: a test passes, a pipeline runs, a result reproduces;
- the environment drifts: tools, APIs and data change, so last month’s examples go stale;
- compute is a GPU or two: a QLoRA run on one machine is feasible where full fine-tuning or RL is not;
- you want to know whether weights help at all: comparing memory alone with memory plus an adapter is a clean experiment.
The manual alternative is a periodic fine-tune from hand-curated logs. However, it is out of date by the time it ships, and someone has to curate it every time.
What Is Continual Learning With LoRA?
Continual learning with LoRA keeps a base model frozen and trains small low-rank adapters on new experience, round after round. As a result, the model adapts while its original weights never change. Each adapter is small enough to train overnight on one machine, so the model keeps up with the work as it changes.
Nightly LoRAs are our modified SFT for policy-free learning in its simplest operational form. Policy-free means no reward sets the direction of each night’s update: the adapter follows a moving average of what recently worked. We use it instead of RL because the SFT vs RL trade-off tips towards modified SFT when success is pass/fail and rollouts are expensive.
LoRA Continual Learning, Round by Round: How the Loop Works
“Nightly” works the way it does in a nightly build: each adapter is rebuilt from scratch, automatically, from the latest successes, and the newest one replaces the last. NSL has no scheduler, so the cadence is yours. One cycle has three phases:
- Collect (one generation). The agent works through its tasks. A pass/fail gate owned by each task keeps the successful episodes, and every LLM call in them becomes an SFT row.
- Train with modified SFT (one training step). The trainer fits a fresh LoRA adapter from the base model on the rows from the last few cycles.
training.inner_losssets the loss: plain SFT, or DFT and anchored SFT, which reweight tokens so rare, high-surprisal ones pull less. - Serve (the next generation). The next round of work runs on the new adapter. NSL applies the newest adapter automatically; a held-out check before serving is a step you add.
NSL’s training loop, scripts/run_train_loop.py, runs this cycle as alternating generations and training steps, with no clock involved. It saves every adapter as after_generation_<N>, applies the newest to the next generation through vllm.lora_adapter_path without a held-out check, and resumes an interrupted run with --run-id.
How the nightly adapter gets better
Three things let each night’s adapter improve on the last:
- The gate decides what is learned. Only episodes the environment accepted become rows, so the adapter imitates what worked in this lab or codebase.
- The loop compounds. A better adapter solves more of the next day’s tasks, so every extra success becomes extra training data for the following night.
- The model picks up strategies, not just facts. For example, in our self-training experiment, agents learned without instruction to run deliberately failing probes that draw out informative error messages. First-attempt accuracy fell while overall task success rose.
Choosing the window
training_window_size (default 3) sets how many recent generations feed each training step. A short window adapts fast and forgets fast; a long one remembers more but lets stale behaviour linger.
In our self-training experiment, three lineages of a 7B model were retrained on their own successful runs. The lineage trained on a sliding window of the last three runs kept improving without catastrophic forgetting in that setup. By contrast, retraining on all past data plateaued, and retraining on only the best runs collapsed.
The overnight training budget
Each night is one QLoRA run on a 4-bit base model. The defaults are lora_rank 32 and max_steps 50. The step count stays fixed whatever the window size, so a wider window spreads the same steps over more rows. The trainer runs as a separate process and waits for the GPUs to be freed, so training and serving can share the same machine.
Continual Learning LoRA Is Hacky: Here Is Why
The research tree on our home page lists chained LoRAs, where one adapter’s data trains the next. Its note reads “it’s hacky but it works.” The hacks are real:
- Adapters can’t be stacked forever. Each night’s adapter replaces the last rather than adding to it. Bolting on a new adapter every night would not scale.
- Every adapter starts from scratch. NSL trains each one from the base model, so it knows only what the window shows it. Earlier nights survive only through the rows still inside the window.
- Serving needs a restart. NSL serves a new adapter by restarting the inference server, not by hot-swapping it. It also launches vLLM with a maximum adapter rank of 64.
- It is a blunt instrument. An adapter changes the model’s behaviour across the board; it cannot target the specific weights a lesson should change.
For the same reasons, our post on continual learning calls LoRAs unscalable and un-beautiful. Sliding-window adapters are how we work around the scaling problem.
Why Continual Learning With LoRA Is Good Anyway
The same properties that make the approach hacky make it practical:
- It learns from the environment’s verdict. A pass/fail check is enough; there is no reward model to build or tune.
- It is cheap. One small adapter per night is a fraction of a full fine-tune: one QLoRA run of 50 steps by default.
- It adapts to drift. New days bring new rows, and the window ages out behaviour that stopped working.
- You can roll back, but that is a safety net, not the goal. The base model never changes, and every
after_generation_<N>adapter stays on disk, so a bad night costs nothing permanent.
Nightly LoRAs compared with the alternatives
For comparison, here is what each option changes every night, and what it costs:
| Approach | What changes nightly | Signal needed | Cost per night |
|---|---|---|---|
| Memory or retrieval only | Notes in the prompt | None | A write |
| Nightly LoRA adapter (modified SFT) | A small adapter | Pass/fail per run | One QLoRA run |
| Nightly full fine-tune | All weights | Accepted traces | A full training run |
| Online RL | The policy, continuously | A reward per rollout | Many rollouts per prompt, plus training |
How to Run a Continual Learning LoRA Experiment
For AI researchers, a nightly continual learning LoRA run tests a question the field has not settled: does updating weights help beyond retrieval, prompting and stored solutions? A research agent in one lab or codebase is the right setting, because the work repeats and success is checkable.
A protocol that answers it:
- Run two arms on the same agent: memory or retrieval alone, and memory plus a nightly adapter.
- Fix a held-out set: tasks from the same kind of work that are never trained on, plus a few older skills.
- Measure three things after each round: success on the held-out tasks, regressions on older skills, and how many new training rows the day produced.
- Vary one setting at a time: the window first (for example one, three and seven nights), then
lora_rank, then the training mix. - Run it for weeks, not nights, because a single night’s score is noisy.
Expect a trade-off curve rather than a single winner: short windows adapt faster and forget faster.
What to Watch For
- A window forgets. Measure older skills on a fixed held-out set after every round; the trade-offs are the subject of catastrophic forgetting in continual fine-tuning.
- A running server can ignore the new adapter. If an inference server is already answering on the target port, NSL reuses it. In that case the new adapter settings are not applied to it, so make sure the loop starts the server it configures.
- Rank above 64 trains but will not serve on vLLM, because NSL starts the server with
--max-lora-rank 64. - The gate decides what is learned. A nightly loop amplifies whatever the pass/fail check admits, good habits and bad.
- One night is noisy. Judge the approach over many cycles, not a single round’s score.
FAQs
Can you do continual learning with LoRA?
Yes. Keep the base model frozen and train a fresh LoRA adapter on recent experience at regular intervals. Then serve the newest one that passes your checks. It reduces forgetting compared with full fine-tuning but does not remove it, so older skills still need measuring after every round.
How often should you retrain a LoRA adapter?
As often as new, checkable experience accumulates and you can evaluate the result. Nightly suits agents that work every day; a slower cadence suits sparse work. Each cycle needs enough new successful runs to matter, and a held-out check before the adapter replaces the current one.
Can LoRA adapters be stacked?
Not indefinitely. Each added adapter costs serving memory, and stacked updates are hard to evaluate on their own. NSL avoids stacking. Every training step starts from the base model with a fresh adapter, so earlier rounds influence it only through the training data in the window.
What should a nightly LoRA train on?
On traces the environment confirmed as successful: runs whose tests passed or whose results held up. Failed runs are kept for inspection but not trained on. A sliding window of the last few cycles keeps the data current, and optional general text can be mixed in to protect general abilities.


