
Key Takeaways
- Your agent solved a hard task yesterday and makes the same first mistakes today, because AI agent memory (retrieval, memory files, scratchpads) carries facts, not skill. The model’s weights never change.
- There are three ways to make experience stick. External memory is cheap and reversible but only helps when the right note is retrieved. RL changes the weights but needs a reward or verifier and many rollouts. LoRA fine-tuning on the agent’s own successful attempts changes the weights from a pass/fail check.
- They are complementary: memory carries facts between sessions, while weight updates change habits. Most agent stacks today stop at memory.
- Weight updates bring their own risk, forgetting, which is the next question in this series.
TL;DR
If an agent should get better at the work it repeats, memory alone won’t do it. Retrieval puts yesterday’s notes in the prompt, but the model reading them is the same model. To change behaviour you have to change weights: through RL when you have a clean reward, or by fine-tuning a small LoRA adapter on the agent’s own successful attempts when all you have is a check that recognises success.
Why Didn’t Your Agent Learn From Solving the Problem?
Picture an agent working through a messy task. It explores, tries three approaches that fail, and the fourth works. The next day a similar task arrives. With a memory file, it may find a note saying what worked last time. Without one, it starts from scratch.
Either way, the model that made the first three mistakes is unchanged, so its instincts are unchanged too. Unless the note happens to cover this exact case, it will try the same wrong things first.
The difference is between remembering and improving. Our page on self-improving AI agents separates three levels: nothing, context, and the policy. Context means retrieval and memory files that carry facts across sessions. The policy level is where the agent updates its own weights or code from outcomes, and only that level compounds.
Large language models are built this way on purpose: trained once, then frozen. Everything an agent “learns” after deployment lives in its context window, which is bounded and has to be re-read on every call.
What Is AI Agent Memory?
AI agent memory is any mechanism that carries information from one interaction to the next outside the model’s weights: retrieval over a document or vector store, a memory file the agent reads and writes, a scratchpad, or conversation summaries. It changes what the model sees, not how the model behaves when it sees it.
That makes AI agent memory the right tool for facts and the wrong tool for skill. A note can say “the build fails unless you clear the cache first”. It cannot make the model reach for the cache first in a situation the note doesn’t describe.
AI Agent Memory vs RL vs LoRA Fine-Tuning
Three mechanisms can make an agent’s experience persist. They differ in what they change and what signal they need.
| External memory (RAG, memory files) | Reinforcement learning | LoRA fine-tuning on own successes | |
|---|---|---|---|
| What changes | The prompt | The model’s weights | A small adapter on top of frozen weights |
| Signal needed | None, beyond what to store and retrieve | A reward per rollout, ideally graded, or a verifier | A pass/fail check |
| Cost per update | A write | Many rollouts per prompt, then training | One training run on collected examples |
| What it captures | Facts, examples, notes | Behaviour, pushed towards the reward | Behaviour, pulled towards what succeeded |
| How to undo it | Delete the entry | Roll back a checkpoint | Drop the adapter |
| Main risk | Retrieval misses, context bloat | Reward hacking, contradictory signals | Forgetting, imitating weak attempts |
External memory: cheap, reversible, shallow
Retrieval-augmented generation was designed for knowledge-intensive tasks: fetch the relevant passages, put them in the prompt, and get provenance and instant updates. Agent memory files and scratchpads work the same way. Writing is cheap, and deleting a bad note undoes it completely.
The usual RAG vs fine-tuning debate is about knowledge: retrieval for facts that change, fine-tuning for stable domain knowledge. For an agent the more useful question is whether its behaviour changes. Memory only helps when the right note is retrieved for the situation at hand, and procedural know-how (“try the cheap probe before the expensive fix”) is hard to retrieve by similarity.
Memory is still worth having alongside learning. NSL’s in-process agent harness keeps a scratchpad that persists across episodes in the same worker, while the weights are updated separately.
Reinforcement learning: changes weights, needs a reward
RL updates the model’s policy so that rewarded behaviour becomes more likely. With a clean reward or a cheap verifier, as in maths or unit tests, it works very well, and RL with verifiable rewards (RLVR) is the standard approach for such tasks.
On agent tasks the requirements bite. RL mostly shifts probability between behaviours the model can already produce, and it needs a coherent signal: when different situations reward opposite behaviour, RL has no direction to converge on.
Sparse rewards add a practical cost. Group-based methods such as GRPO learn nothing from a group of rollouts that all failed, and every rollout of an agent task is a full episode with many tool calls.
LoRA fine-tuning on the agent’s own successes
The third route keeps the weight update but drops the reward. The agent attempts the task many times, the attempts that pass a check are kept, and a small LoRA adapter is fine-tuned on them and served for the next round. The base model stays frozen, so a bad adapter can simply be thrown away.
This is how Eigenform’s policy-free continuous learning loop works. In NSL the loop is scripts/run_train_loop.py, and each round trains a fresh adapter from the base model on the successful attempts of the last training_window_size generations (default 3). It needs no graded reward, only a task check that recognises success.
It also changes behaviour in ways memory could not. In our self-training experiments, agents learned without instruction to run deliberately failing probes to draw out informative error messages. First-attempt code accuracy fell across generations while overall task success rose: the model had picked up a strategy, not a fact.
AI Agent Memory or Fine-Tuning: Which Should Your Agent Use?
Choose by what you want to persist and what signal you have:
- Facts that change often (documentation, tickets, prices): external memory or RAG. Nothing else updates as cheaply.
- A clean, cheap reward or verifier (unit tests, maths answers): RL or RLVR, which can also learn from failures.
- Recurring, relatively unstructured work with a pass/fail check, where you want habits to change: LoRA fine-tuning on the agent’s own successful attempts.
- Most real agents need two of these: memory for the facts, adapters for the skill.
Without weight updates, improving an agent means improving its prompts and tools by hand, every time. That is the ceiling the continual learning approaches try to lift.
How to Test Whether Learning Beats AI Agent Memory
For AI researchers, whether weight updates add anything beyond memory is not settled, and it is cheap to test on your own tasks. Research agents are also a natural fit. They work repeatedly in one lab, codebase or data pipeline, their tasks rarely come with a clean reward or a cheap verifier, and the budget is a GPU or two rather than an RL cluster. Small LoRA adapters fit that budget, and a bad one costs nothing to discard.
A simple protocol:
- Hold out a task set the agent has not seen, drawn from the same kind of work.
- Run two arms on the same base model: memory alone, and memory plus an adapter trained on the agent’s own successful attempts. Add an RL arm only if you have a reliable reward.
- Measure more than success: task success, attempts or tool calls before the first success, and the cost of each update.
- Run several rounds and keep every adapter, so you can see whether gains compound or fade, and roll back when they fade.
Decide in advance which metric counts as improvement. The results may not move the way you expect: in our runs, first-attempt accuracy fell while task success rose, because the agents had learned to probe deliberately.
What to Watch For
- Weight updates can make a model forget what it already did well. Repeated fine-tuning on narrow, recent data erodes older skills, a problem called catastrophic forgetting.
- A pass/fail check decides what the model imitates. A check that lets weak attempts through teaches bad habits.
- Memory and learning can disagree: a stored note may contradict a habit the adapter has learned. Keep notes factual and let the weights carry strategy.
FAQs
Do AI agents learn from experience?
Not by default. Most agents carry experience forward only through memory: retrieved documents, memory files or summaries in the prompt. The model’s weights stay the same, so its habits do not change. Learning from experience in the stronger sense means updating weights, through RL or by fine-tuning on the agent’s own successful attempts.
Is RAG or fine-tuning better for AI agents?
They solve different problems. RAG is better for facts that change often, because updating it is a single write and fully reversible. Fine-tuning is better for changing how the agent behaves, such as which approach it tries first. Many agents need both: retrieval for knowledge, and a fine-tuned adapter for skill.
Can an AI agent update its own weights?
Yes, if its training loop is built for it. In NSL the agent’s successful attempts become fine-tuning data, a LoRA adapter is trained on them, and the next round of attempts runs with the new adapter. The base model stays frozen, so each adapter can be evaluated and discarded if it is worse.
Does fine-tuning an agent make it forget?
It can. Each round of fine-tuning pulls the weights towards the newest data, and older abilities can erode, which is called catastrophic forgetting. LoRA adapters, restarting from the base model each round, sliding windows of recent data and rehearsal data all reduce it, but none removes it, so keep evaluating older skills.
