TL;DR
- The AI models behind today’s chat assistants stop learning once they are released. They can’t get better at your specific problem, however long you use them.
- If your research runs on private data, you want the opposite: a model that learns your field by testing ideas against your data, without that data ever leaving your control.
- We build small models that do exactly that. They run on local hardware, improve with every round of experiments, and belong to you.
Key takeaways
- Frozen by design. Consumer models serve millions of users, so they can’t learn from any one of them.
- Real results decide what is learned. Our models only keep what actually worked when tested, so there is no score to game.
- Small, local and specialised. A compact model that has run thousands of experiments on your data can know your problem better than a general model that knows a little about everything.
Why the big models stop learning
The large models behind today’s chat assistants do not learn while you use them. Their weights, the billions of numbers that hold what the model knows, were fixed during training months earlier. Every conversation runs on that same frozen snapshot.
This is a deliberate trade-off. One model serves millions of people at once. If it learned from each conversation, one customer’s data could surface in another customer’s answers. Every change would also need fresh safety testing. Retraining at that scale costs millions of dollars, so it happens a few times a year at most.
What looks like learning is usually something else. “Memory” features save notes outside the model and paste them back into the prompt. The model reads those notes, but it does not get any better at the task. And the notes still live on someone else’s servers.
Why you want a model that learns
For everyday tasks, a frozen model is fine. Research is different.
If your work runs on proprietary data (drill logs, survey archives, lab results, internal reports), a general model has never seen anything like it. It was trained on the public internet, and your data is, by definition, not there.
You often cannot send that data to a third-party API either. Confidentiality agreements, regulation or plain competitive sense rule it out.
What you actually want is a model that learns your domain the way a new colleague does: by working with the material, testing ideas and remembering what worked. It should do this where your data already lives. And the result, the improved model itself, should belong to you.
That is test-time learning: a model that improves from its own experience while it is in use, not only in a lab before release.
How we do it
The hard part of self-improvement is deciding which lessons to keep. A model that grades its own work drifts. A model trained against a scoring rule learns to game the rule instead of solving the problem, a failure known as reward hacking.
Our answer, published in Survival is the Only Reward, is to let the environment decide. The loop is simple:
- ─ 01
Propose
The model writes code to test an idea.
- ─ 02
Execute
The code runs against real data or a real system.
- ─ 03
Measure
We record what actually changed as a result.
- ─ 04
Keep what survives
Only attempts with real, measurable results become training data for the next version of the model.
Selection rests on real consequences, not on a score, so there is nothing to game.
From experience to weights
Every pass through the loop leaves a record: what the model tried, and what happened. Those records are how the model learns.
- ─ 01
Collect
The model works through the loop until it has built up a batch of successful attempts. In our experiments that meant 4,500 records per round.
- ─ 02
Train
The batch is used to fine-tune the model, with a small sample of general material from the model's original training (500 records per round) so it keeps its broader skills.
- ─ 03
Swap
The improved model replaces the old one and picks up the same work, now spotting opportunities its predecessor would have missed.
- ─ 04
Repeat
It produces a new batch, and the cycle runs again.
Each round starts again from the original model and applies the new learning as a small add-on, rather than piling each update on top of the last. That is what lets the cycle keep running without the model forgetting what it already knew.
The training data can also be limited to the most recent rounds, so storage stays bounded. Good strategies survive because the model keeps using them, not because every old record is archived.
We also tried more elaborate reward-based training methods. The simplest option, training directly on the attempts that worked, proved the most stable.
Applied to your project
The same cycle runs on your data. The environment becomes your working material: your datasets, your models, your analysis tools. A successful attempt is one that holds up against your own evidence. The model proposes hypotheses, writes code to test them, and keeps what survives. Each round of training makes it more expert in your domain.
All of it runs on local hardware you control. Nothing is sent to an outside provider at any stage.
In our published experiments, this produced models that:
- Improve by pruning. Most gains came from dropping weak strategies and refining good ones, rather than piling up new ones. We call this negative-space learning.
- Keep improving with bounded memory. A model trained only on its three most recent batches of experience kept getting better. There is no ever-growing archive to store.
- Keep their general skills. Performance on the task rose more than fivefold, while scores on a general coding benchmark barely moved.
- Learn how to learn. Later versions started deliberately running experiments designed to fail informatively, using the error messages to explore. Nobody told them to do this.
- Get faster. The time to gather a full batch of training data fell from 5–7 days to 10–15 hours.
State-of-the-art performance, locally
Because the learning comes from the environment rather than from scale, the model does not need to be huge. Our experiments used a 7-billion-parameter open-weight model, small enough to run on a single workstation GPU such as The Box.
A general model knows a little about everything. A small model that has run thousands of experiments on your data knows your problem. In our geological experiment, 1,137 experiments produced 60 proven relationships. Training on the successful runs gave a 17.6% improvement over the base model on a separate geology benchmark.
And the expert model that comes out of the process is yours to keep.
Case study: a geological world model built on coherence
Everything above depends on the “measure” step: some honest way to tell whether an attempt worked. In a sandbox, that is easy. In mineral exploration, it is the hardest part of the job.
Ground truth is scarce and expensive: a single new data point can mean a $500,000 drill hole. So we needed a signal that tells real geological relationships from accidental ones, without drilling to check every time.
That signal comes from the geology itself. Geological observations are not independent clues. A fault, an alteration halo and an unusual elemental ratio are all consequences of one underlying history. A real relationship between them should make the rest of that history easier to explain.
Most AI for geology stops at correlation: which variables occur alongside ore? Our approach, coherence mapping, asks a harder question: does this relationship make the whole geological picture easier to explain?
The system runs hundreds or thousands of regressions across a regional dataset. Each candidate relationship is then tested against the whole world model. It stays only if it makes the rest of the geology more coherent. A strong correlation that makes everything else harder to explain gets dropped. That happens, for example, when it comes from a badly calibrated instrument or shifted coordinates.
Coherence becomes the measure in the learning loop. The model proposes a relationship, tests it against the data, and keeps it only if the world model as a whole gets more coherent. Nobody hand-labels anything. The evidence decides, just as the environment decided in our published experiments.
For an exploration team, that means:
- A voxel belief map of the subsurface, in which every retained relationship has earned its place. Explore the demo.
- Contradictions surfaced, not averaged away. When one dataset clashes with everything else, it is flagged: either the data is wrong, or the geological explanation is incomplete. Both are worth knowing.
- Fewer wasted drill holes. Targets are ranked by how well they fit the whole geological picture, not by one strong signal.
- An explainable result. Every relationship in the model can be traced back to the tests that kept it.
We have applied this on live exploration projects in Australia and Central Asia. See our case studies for more.
Work with us
Have proprietary data and a problem general models cannot touch?
We would like to hear about it. Read the full paper on arXiv, or get in touch directly.