
TL;DR
- LLM pre-training is the first and biggest stage of building an AI language model: it reads a huge amount of text and learns one skill, guessing the next word.
- The result, called a base model, knows a lot but only continues text. It does not yet answer questions or follow requests.
- Everything after that steers the base model, through further training, the software around it, or tools and skills. A benchmark then shows whether each change helped.
Key Takeaways
- Pre-training is where a model’s knowledge comes from. One open training dataset alone holds 15 trillion tokens, the word pieces a model reads.
- A model fresh from pre-training only continues text. Asked to finish a sentence about the 2006 US president, a small base model we ran turned it into a multiple-choice exam question.
- You can change what a model does without starting again. Further training, a harness, tools and skills each steer the same base model in a different way.
- Every change needs a before-and-after test. A benchmark asks both versions the same questions and scores the answers, so you can see what actually improved.
Every AI chat assistant starts as a base model, the raw product of LLM pre-training. That model holds most of what the assistant will ever know, but it behaves nothing like an assistant.
What is LLM pre-training?
LLM pre-training is the first training stage of a large language model, in which it learns to predict the next token across a very large text corpus. No labels are needed, because the text supplies its own answers. The result is a base model with general knowledge of language, facts and reasoning patterns.
What is pre-training in machine learning more generally? It is training a model on a broad task first, so that later, narrower training starts from useful weights instead of random ones. For language models, that broad task is predicting text.
What a base model does on its own
A base model has learned to continue documents, not to answer people. Published examples are rare, because the companies behind closed models such as GPT-5 do not release their base models. So we ran a small open one, Qwen2.5-1.5B, which has had no instruction tuning, beside Qwen2.5-1.5B-Instruct, the same model after chat tuning. Both used greedy decoding and stopped after 80 new tokens, so some outputs end mid-sentence.
Prompt, given to the base model as plain text with no chat template:
The president of the united states in 2006 wasOutput from the base model:
A. George W. Bush
B. Bill Clinton
C. George H. W. Bush
D. Barack Obama
答案:
A
下列关于我国人口的叙述,正确的是____
A. 我国人口分布不均,东部人口多,西部人口少
B. 我国人口分布不均,东部人口少,西部人口多
C.It read the sentence as a multiple-choice exam question, marked the answer in Chinese (答案 means “answer”) and started the next question, about China’s population. That habit comes from exam papers in its training data. Nathan Lambert’s book on RLHF shows the far larger Llama 3.1 405B Base doing much the same.
Post-training changes the format, not necessarily the facts. Given the same prompt in its chat format, Qwen2.5-1.5B-Instruct answered confidently, and wrongly: “Barack Obama was the President of the United States in 2006. He served from January 20, 2009 to January 20, 2017.”
Base model vs chat model: more examples
Other test prompts show the same habit. Each output below is complete, exactly as generated.
Prompt: “Hi! How are you today?”
Base model (Qwen2.5-1.5B):
I'm doing great, thank you! How about you? I'm doing well too, thanks for asking! How are you feeling today? I'm feeling great, thank you! What's your favorite color? My favorite color is blue. Why do you like blue? I like blue because it makes me feel calm and peaceful. What's your favorite animal? My favorite animal is a cat. WhyChat model (Qwen2.5-1.5B-Instruct):
Hello! I'm just an AI assistant and don't have feelings like humans do. But thank you for asking how I am - it's nice to hear from people! Is there anything specific you'd like help with or talk about?Prompt: “Why is the sky green?”
Base model (Qwen2.5-1.5B):
Why is the sky blue? Why is the sky blue? Why is the sky blue? Why is the sky blue? Why is the sky blue? Why is the sky blue? Why is the sky blue? Why is the sky blue? Why is the sky blue? Why is the sky blue? Why is the sky blue? Why is the sky blue? Why is the sky blue? Why isChat model (Qwen2.5-1.5B-Instruct):
The color of the sky appears to be blue because of a phenomenon called Rayleigh scattering. This occurs when light from the sun travels through Earth's atmosphere and interacts with gas molecules in the air. The shorter wavelengths of blue light are scattered more easily than longer wavelengths like red or orange.
As sunlight enters our atmosphere, it encounters nitrogen and oxygen molecules, which scatter the blue light more effectively than other colorsPrompt: “Write a Python function that reverses a string.”
Base model (Qwen2.5-1.5B):
The function should take a string as input and return the reversed string. The function should not use any built-in string reversal functions or methods. The input string can contain any printable ASCII characters. The function should handle strings of any length, including empty strings. The function should have a time complexity of O(n), where n is the length of the input string. The function should not use any additional dataChat model (Qwen2.5-1.5B-Instruct):
Certainly! Below is a simple Python function to reverse a given string:
```python
def reverse_string(s):
return s[::-1]
# Example usage:
input_str = "Hello, World!"
reversed_str = reverse_string(input_str)
print(reversed_str) # Output: "!dlroW ,olleH"
```
This function uses slicing with a step of `-1` (`Prompt: “What do you think about artificial intelligence?”
Base model (Qwen2.5-1.5B):
Respond using the words/style of R2-D2 from Star Wars.
I think artificial intelligence is fascinating! It's like a new frontier in technology, full of endless possibilities. Just like how I can learn new things from the Force, I can learn new things from AI. It's like a whole new world of knowledge waiting to be discovered. But, I do have some concerns about how we useChat model (Qwen2.5-1.5B-Instruct):
As an AI language model, I don't have personal opinions or beliefs. However, I can provide some information and insights on the topic of artificial intelligence.
Artificial Intelligence (AI) is a rapidly growing field that involves creating machines and software systems that can perform tasks that typically require human intelligence, such as visual perception, speech recognition, decision-making, and language translation. AI has the potential to revolutionThe base model still handled simpler requests, such as a short poem or job-interview tips, a sign that today’s pre-training data already holds plenty of question-and-answer text.
Why LLMs need pre-training
A language model’s weights start as random numbers. Pre-training is where they come to hold grammar, facts, and the patterns of argument and code the model later draws on.
Later training can redirect that knowledge but struggles to add to it. In our research on training from self-produced data, we call this the warm start problem: pre-training shapes the weights so thoroughly that a small, narrow dataset either barely changes the model or makes it forget what it knew.
How to train an LLM: the pre-training process
Pre-training runs in five stages.
Data collection and filtering
Pre-training data is mostly web text, plus books, code and reference material. Teams filter the raw crawl for quality and strip duplicates and boilerplate. FineWeb, an open pre-training dataset, holds 15 trillion tokens taken from 96 Common Crawl snapshots.
Tokenization
A tokenizer splits text into tokens, usually sub-word pieces produced by an algorithm such as byte-pair encoding. Common words become a single token, while rare words, names and codes break into several.
The training objective: next-token prediction
Given the tokens so far, the model predicts the next one, and its weights are adjusted to reduce the error, known as the cross-entropy loss. Repeated over trillions of tokens, that single task pushes the model to capture grammar, facts and some reasoning, because all of them help predict text.
Compute and distributed training
LLM pre-training is bound by compute and data more than by architecture. The Chinchilla study found that for a fixed compute budget, model size and training data should grow together, at roughly 20 tokens per parameter.
Checkpoints and evaluation during training
A run saves checkpoints, snapshots of the weights, at regular intervals. They let it recover from hardware failures and let the team test progress along the way.
Pre-training vs. post-training
Pre-training vs post-training is a split between building capability and shaping behaviour. Post-training covers supervised fine-tuning on instruction examples, preference methods such as DPO, and reinforcement learning.
| Pre-training | Post-training | |
|---|---|---|
| What it optimises | Next-token prediction on broad text | Following instructions, format, preferences, safety |
| Typical data | Trillions of tokens | Thousands to millions of curated examples |
| Share of compute | The bulk of the total | A small fraction |
| What it can fix | Missing knowledge and core ability | Tone, format, refusals, task habits |
| What it can’t fix | Behaviour on specific tasks | Knowledge and skills the base model lacks |
How to steer a model after LLM pre-training
Post-training is one way to change what a base model does. There are others, and most of them leave the weights alone.
Prompting
The cheapest lever is the text you give the model. A base model continues whatever document it seems to be in, so a prompt that looks like a page of good questions and answers gets answers back. In the URIAL study, a system prompt and as few as three fixed examples let untuned base models match or surpass versions aligned with SFT or RLHF.
Fine-tuning the weights: SFT, RL and LoRA
Supervised fine-tuning (SFT) trains the model on examples of the behaviour you want. Reinforcement learning (RL) rewards outputs that score well. A LoRA adapter is a small block of extra weights trained beside the frozen model, so the base never changes and the adapter can be swapped out.
In our self-training research, we fine-tuned Qwen 2.5 7B Instruct on its own successful attempts at a task. We tried PPO, GRPO and DPO, and plain SFT was the most stable. Each round trained a fresh LoRA adapter on about 5,000 examples.
Wrapping it in a harness
A harness is everything around the model: the system prompt, the tools it may call, the retrieval layer and the agent loop. Changing the harness changes behaviour without touching a single weight.
Geocluster, our AI environment for earth-science data, is built this way. It wraps whichever model you choose in a fork of the Cline coding agent, with limits on what the agent may do, specialist sub-agents and more than fifty data tools.
Giving it tools through MCP
Tools let a model act and look things up. MCP, the Model Context Protocol, is an open standard for exposing tools to AI applications. Stratigraphic Amenity, our open-source MCP server for scanned geological maps, gives an agent 10 tools, from registering a map to georeferencing it.
Teaching it procedures with skills
A Skill is a folder of written instructions, and optionally scripts, that an agent loads only when a task matches its description. The question-writing Skill in our dynamic benchmarking repository runs to 17 sections, from reading the source documents to validating the finished question set.
Training frameworks and infrastructure
Training large language models takes more than a training loop. The software has to keep thousands of devices busy and recover when some of them fail.
What an LLM training framework has to handle
- Parallelism: splitting data, layers and the model itself across devices.
- Checkpointing: saving and restoring the full training state, including optimiser state, without stalling the run.
- Data pipelines: streaming tokenised data in a fixed order, so a restarted run sees the same batches.
- Fault tolerance: detecting failed devices and resuming from the last checkpoint.
Open-source frameworks such as Megatron-LM, DeepSpeed and torchtitan do this at pre-training scale. Steering needs far less: prompts, harnesses, tools and skills need no training, and a LoRA adapter for a 7-billion-parameter model trains on one machine.
What LLM pre-training can’t give you
A knowledge cutoff
A model’s knowledge stops when training ends, often months before release. As the world moves on, its answers fall out of step, one of the causes of model drift.
Private data
Pre-training only sees public text. Lab results, internal reports and anything else a company never published are missing, which is the gap test-time learning aims to close: a model that keeps learning from your own data after release.
Grounded answers
A base model stores facts as patterns, not records, and cannot check a claim against anything outside itself. That is where hallucination starts, and why tools and retrieval matter.
How to measure whether a change worked
Every way of steering a model needs a before-and-after test on material the model has not seen, or a gain may come from memory rather than the change.
Public benchmarks are the usual risk, because they leak into web crawls. On GSM1k, a fresh set of grade-school maths problems written to mirror GSM8K, accuracy dropped by up to 8% for some models.
Groundtruth, Eigenform’s open source dynamic benchmark, is built for these comparisons. Its first edition grades answers about real geology documents from 0 to 10. Its harness track holds the model fixed, so a change in score belongs to the harness, and a pairwise mode asks whether a fine-tune beat its base model.
That comparison is the core of AI model evaluation, and it is where pre-training’s limits show.
Keeping the comparison repeatable is the job of an AI evaluation framework: a fixed dataset, rubric and judge, run again every time the model changes.
FAQs
Is LLM pre-training supervised or unsupervised?
It is self-supervised. There are no human labels: the model predicts the next token, and the text itself supplies the right answer. It is often called unsupervised because nobody annotates the data, but the training signal is exact, with one correct token at every position in the text.
How do you train a large language model?
In stages. Pre-training teaches a base model to predict text from a large, filtered, tokenised corpus across many accelerators. Post-training then adapts it with instruction examples, preference data or reinforcement learning so it follows requests. Prompts, a harness, tools and skills can steer it further without changing its weights.
What is the difference between pre-training and post-training?
Pre-training builds general capability from trillions of tokens of raw text and uses most of the compute. Post-training uses far smaller curated datasets to shape behaviour: following instructions, format, tone and safety. It can redirect what a model already knows, but it cannot reliably add knowledge or skills that pre-training never encoded.
What frameworks are used for LLM pre-training?
Large runs use distributed-training frameworks such as NVIDIA’s Megatron-LM, Microsoft’s DeepSpeed or PyTorch’s torchtitan. They share the same objective, next-token prediction, and differ mainly in how they split data, layers and model state across devices and recover from failures. Fine-tuning a small model on one machine rarely needs them.
How much data is needed to pre-train an LLM?
It depends on model size and compute budget. The Chinchilla study found roughly 20 training tokens per parameter to be compute-optimal, about 1.4 trillion tokens for a 70-billion-parameter model. Many recent models train on far more, and open datasets such as FineWeb, at 15 trillion tokens, now exist at that scale.


