SFT vs RL for sparse-feedback agents: modified SFT trains a small adapter on a window of accepted episodes and archives the failures, GRPO groups that all failed give no gradient, and RLVR learns from a verifier that marks every rollout
Match the method to the signal you have

Key Takeaways

  • When an agent gets only a pass/fail signal, rarely succeeds, and every attempt is an expensive episode, RL spends most of its compute on rollouts that teach it nothing. That is where the SFT vs RL choice matters, and where we use neither RL nor classic SFT.
  • If you need continuous learning the default is reinforcement learning, but that requires an established policy and a reliable, dense reward system. What if you don’t have those? Maybe SFT isn’t entirely tapped out. It’s possible to use a combination of overnight LoRAs and modified SFT approaches to approximate policy-free online learning. Here we show you how.
  • Classic SFT fits a fixed dataset once; RL needs a reward that varies across attempts. Formally, our loop is close to REINFORCE with a 0/1 reward, and it departs through the window, the fresh adapter and the loss.
  • With a cheap, reliable verifier, RLVR is the stronger default. Modified SFT is for the agent tasks that have none.

TL;DR

Use modified SFT when an environment can recognise a successful episode but cannot grade it, and rollouts are too expensive for RL. Policy-free means the update has no reward pulling it in a fixed direction: the model tracks a moving average of what recently worked. RL learns from failures and generalises further, so when a cheap verifier exists, use RLVR.

Why Choose Between SFT and RL at All?

The SFT vs RL question usually gets a quick answer. The default advice for post-training is RL: it optimises the real objective and generalises beyond its training examples. That advice assumes a reward that varies across attempts and enough attempts to estimate it. Agent tasks often break both assumptions.

The question matters when:

  • feedback is pass/fail: the environment can say that an attempt worked, not how good it was;
  • success is rare: most attempts fail;
  • each rollout is expensive: a full agent episode, not a single completion;
  • the environment keeps changing: tools, APIs and data shift, so the agent has to keep learning after deployment;
  • compute is a GPU or two: no room for a reward model, a critic and rollout workers.

Without a deliberate choice, teams build the RL stack first and then find that most rollouts contribute no gradient. The usual fallback, a one-off fine-tune on curated examples, stops learning the day the agent ships.

What Is the Difference Between SFT and RL?

The difference between SFT and RL is the training signal. Supervised fine-tuning (SFT) raises the likelihood of given target outputs with a cross-entropy loss. Reinforcement learning (RL) samples outputs from the current model and weights them by a reward, so better-scored outputs become more likely. SFT needs targets; RL needs a reward that varies across attempts.

Supervised fine-tuning vs reinforcement learning therefore comes down to the signal you have. Classic SFT takes its targets from a fixed dataset, trains once and stops. Our modified SFT is a new approach to SFT for continuous learning: it keeps the loss but changes where the targets come from and how often it trains. The targets are the agent’s own episodes that passed a task-owned check, and training runs online between rounds of episodes.

That is what we mean by policy-free learning. No reward defines which direction the model should improve in. The environment only decides which episodes count as successes, and the model follows a moving average of what worked. The choice arises at all because AI agent memory isn’t learning: improving an agent means changing its weights, not its notes.

SFT vs RL in NSL: How Our Modified SFT Works

Negative Space Learning (NSL), our research loop, runs modified SFT in four steps; its simplest form is a nightly LoRA adapter trained on recent successes:

  1. A task-owned gate marks each episode a success or not; failed episodes are archived and contribute zero rows.
  2. Each LLM call in a successful episode becomes an SFT row, with the episode as its label.
  3. A fresh LoRA adapter is trained from the base model on a sliding window of recent generations, training_window_size, default 3.
  4. training.inner_loss can replace the plain SFT loss with DFT or anchored SFT.

We tested PPO, DPO and GRPO on our own datasets before settling on modified SFT, and our public research tree keeps them as approaches we may revisit locally. Those runs are not written up, so the reasons below are the general mechanism, not a post-mortem of our experiments.

SFT vs GRPO, PPO and RLVR in Sparse-Feedback Environments

In sparse-feedback environments, SFT vs RL compares four methods that differ in the signal they need and what they hold in memory:

Modified SFT (ours)PPOGRPORLVR
Training signalPass/fail from a task-owned gateScalar reward per rolloutScalar reward per rollout, compared within a groupReward from an automatic verifier
Extra networksNoneValue model (critic)NoneDepends on the optimiser
Samples per promptOne accepted episode is enoughManyA group per promptMany
Learns from failuresNo: failed episodes are archivedYesYes, when the group’s rewards differYes
Typical failureImitates whatever the gate admitsCritic and reward-model cost and instabilityZero gradient when a group’s rewards are all equalNo verifier, no method

PPO: a critic to train

PPO scores each action against a learned value model, then takes clipped policy steps. The critic is a second network of similar size to hold in memory, and it needs reward signal dense enough to fit.

SFT vs GRPO: no critic, but a group to fill

GRPO drops the critic: it samples a group of answers per prompt and scores each against the group, A_i = (r_i − mean(r)) / std(r). That saves memory, but a group with equal rewards carries no signal.

At a 5% success rate and eight rollouts per prompt, the chance that a whole group fails is 0.95⁸, about 66%, so roughly two thirds of the groups produce no gradient, while each rollout still costs a full agent episode. Modified SFT needs only one success to produce training rows.

SFT vs RLVR: no verifier, no RLVR

RL with verifiable rewards, named in Tülu 3, replaces a learned reward model with an automatic check, such as a maths answer or a unit test. Where that check exists and is cheap, RLVR works well and is hard to beat. Many agent tasks have no such check: a hypothesis about sparse data, a change to a running system, a report summarised for a specialist.

What Can SFT Do That RL Can’t?

The honest answer is “less than it looks”, because the two are formally related. In an online loop, though, four things separate our modified SFT from RL.

Learn from a single success

One accepted episode becomes training rows immediately. RL needs that success to recur, and to contrast with failures in the same group, before it moves the policy. For a task solved once in fifty attempts, modified SFT starts learning on the first success.

Keep learning without a reward to climb

A policy gradient pushes the model towards whatever the reward defines, so the reward has to say in advance what better means. Our loop never states that. The gate keeps judging episodes, and the window holds only recent successes, so behaviour that stops passing ages out of the training data. When tools or data change, nothing has to be re-specified.

Bound drift with a window and a fresh adapter

Each step starts a new adapter from the base model, so earlier rounds carry over only through the rows still in the window. Every adapter stays on disk as after_generation_<N>, so rollback means serving an older one. Drift is bounded without a KL penalty or a reference model in memory.

In our self-training experiment, the lineage trained on a sliding window kept improving, while cumulative retraining plateaued and best-runs-only collapsed.

A window also decides what the model forgets, the trade-off at the heart of catastrophic forgetting in continual fine-tuning.

Reweight tokens instead of trusting every one

Plain SFT treats every token of an accepted episode as equally right. DFT scales each token’s loss by the model’s own probability of it, so rare, high-surprisal tokens pull less. The DFT paper shows that the plain SFT gradient implicitly encodes a problematic reward.

Anchored SFT adds a KL term to a reference model; in NSL the reference is the base model. Both run through training.inner_loss.

With on-policy samples and a 0/1 reward, the REINFORCE gradient E[r · ∇log π(y)] is the SFT gradient on the successes, scaled by the success rate. Training on the attempts that passed is the rejection sampling family of expert iteration, STaR, ReST and RAFT.

NSL departs from that equivalence in three ways. Its window mixes rows from earlier adapters, so the data is partly off-policy and uncorrected. Each step restarts from the base model instead of applying a KL penalty. And the loss can reweight tokens.

SFT vs RL: What RL Does Better

RL wins on four counts:

  • It learns from failures. A baseline turns a failed attempt into a negative signal; modified SFT archives it.
  • It optimises the objective itself. SFT optimises the likelihood of accepted episodes, which only approximates the goal.
  • It generalises further. A comparison of SFT and RL on unseen variants found that RL generalised to new rule and visual variants while SFT tended to memorise its training data.
  • It can exceed its best trace. SFT imitates accepted episodes; it rarely does better than the best one it trained on.

So this is not a case of SFT over RL. Use RLVR when a cheap verifier exists, and modified SFT for policy-free learning when the environment can only say pass or fail.

How to Test SFT vs RL on Your Own Agent

For AI researchers, whether RL with verifiable rewards adds abilities a model did not already have is still an open question. RLVR-trained models beat their base models on the first attempt but not at large pass@k, and SFT tends to memorise where RL generalises. Pass/fail agents are where the two are hardest to compare.

A protocol that isolates the method:

  1. Fix everything except the method: the same base model, task variations and episode budget for each arm.
  2. Run two or three arms: modified SFT on accepted episodes with a sliding window, GRPO with the same pass/fail reward, and RLVR if a verifier exists.
  3. Measure the signal, not only the score: pass@1 and pass@k, the share of GRPO groups with zero reward variance, and compute per point of improvement.
  4. Hold out variants neither arm trained on, to separate memorisation from generalisation.
  5. Repeat over several rounds, since both methods can compound or drift once a model trains on its own output.

Report the outcome as an experiment from a stated setup, not a general verdict.

What to Watch For

  • SFT imitates the gate. A pass/fail check that admits weak episodes teaches weak habits. The gate is the most important design decision in a modified SFT loop.
  • Episode-level labels reward dead ends. Every call in a successful episode becomes a positive example, including the detours.
  • SFT memorises. Hold out task variants the model never trained on, and measure them every round.
  • Contradictory signals defeat both. When situations reward opposite behaviour, RL has no coherent direction, and SFT averages the contradiction into its data.

FAQs

Which is better, SFT or RL?

In SFT vs RL, neither is better in general. RL learns from failures, optimises the objective directly and generalises further, so it is the stronger choice when a cheap, reliable verifier exists. Modified SFT for policy-free learning is the practical choice when the environment can only say pass or fail, successes are rare and rollouts are expensive.

What is the difference between SFT and GRPO?

SFT trains on target outputs with a cross-entropy loss. GRPO samples a group of outputs per prompt, scores each against the group’s mean reward and pushes probability towards the better ones. GRPO needs a scalar reward and variation within each group; with rare binary success, most groups give it no gradient.

What is policy-free learning?

Policy-free learning improves a model from its own experience without a reward that sets the direction in advance. In our loop, the environment only decides which episodes count as successful. The model is fine-tuned on a sliding window of those, so it tracks a moving average of what recently worked instead of climbing a reward.

When should I use RLVR instead of SFT?

Use RLVR when an automatic check can verify each answer cheaply, such as maths answers, unit tests or formal proofs, and when you can afford several rollouts per prompt. The verifier then turns both successes and failures into signal. Without such a check, RLVR has nothing to optimise against.