Review date / author: 2026-09-15 / Zhongzhu Zhou
Paper reviewed: Soft Adaptive Policy Optimization (SAPO)
Paper authors: Chang Gao, Chujie Zheng, Xiong-Hui Chen, Kai Dang, Shixuan Liu, Bowen Yu, An Yang, Shuai Bai, Jingren Zhou, Junyang Lin
arXiv: 2511.20347v2, revised 1 December 2025
1. Reading map and scope
I read SAPO as a proposal about how much to trust each sampled token during a policy update. Its most useful question is operational: after generating answers with yesterday’s policy snapshot, which parts of those answers should still influence today’s parameters? This question appears whenever a rollout batch supports several optimization steps, even without a large asynchronous system.
The core construction uses a sigmoid surrogate whose derivative is a smooth weight. That distinction matters. Multiplying a conventional loss by a sigmoid is generally a different algorithm because automatic differentiation also differentiates the multiplier. In this review I derive the intended derivative, inspect sign-dependent clipping, and separate scalar gate similarity from similarity of full parameter gradients.
The paper is nine pages including references. The longer treatment here is an instructional companion: calculations, counterexamples, implementation reasoning, and suggested experiments are my analysis. The three cropped training figures retain the paper’s measurements; the other six figures are original explanatory diagrams or deterministic calculations. No model training was run for this review. Exact-looking coordinates are not extracted from plotted curves.
A reader familiar with policy gradients can start at Section 5. A reader implementing the objective should read Sections 3, 6, and 10 together. Sections 12 and 13 distinguish the paper’s displayed evidence from what would establish a reliable engineering improvement. The conclusion is deliberately conditional: a smoother local update is useful, but it is not a guarantee of a globally safe training trajectory.

2. Prerequisites: an LLM as a stochastic policy
An autoregressive model gives a probability to each next token. The state at position is the prompt followed by the already sampled prefix. If the response is , the chain rule gives
Taking a logarithm turns a product into a sum:
The score vector tells us which parameter changes locally increase the log probability of that sampled token. It is a vector with one component per parameter, not a scalar confidence score. Its magnitude and direction may differ dramatically across tokens.
For a reward that is fixed with respect to the differentiable policy computation, the likelihood-ratio identity follows by writing inside the expectation:
The identity assumes that the usual interchange of summation and differentiation is valid. If a reward depends differentiably on the policy parameters, its own derivative adds another term. In the sampling-based language-model setting, rewards are normally treated as fixed labels for an update.
Subtracting a baseline that depends only on the prompt does not change the ideal on-policy expectation: the score has mean zero because probabilities sum to one. A baseline reduces variance. A baseline estimated from the same finite response group, however, changes the finite-sample estimator, so practical group-normalized objectives should not automatically inherit every unbiasedness statement about an ideal baseline.
3. Group advantages, masks, and normalization
For one prompt, collect rewards. Write their population mean and standard deviation as and . A practical fixed advantage is
The small is an implementation safeguard, not a claim that the paper specifies this exact constant. If rewards are , then and ; ignoring , the advantages are . All tokens in one response share its advantage. Thus a positive response does not mean every intermediate token was individually useful. Credit assignment remains coarse.
If every reward is the same, the centered rewards are zero. A denominator safeguard should yield zero learning signal, not NaNs. A group containing one rare success can give that success a relatively large normalized advantage. Reward scaling and group composition therefore interact with the gate; the gate does not repair a poor reward definition.
For variable response lengths, SAPO’s displayed objective averages over tokens within each response and then over responses. With mask and valid length , implement this as
This differs from pooling every valid token into one large average. Suppose one response has 10 tokens and another has 100. Per-response averaging gives each answer half the total weight; token pooling gives the longer answer roughly ten times as much weight. Neither convention is automatically right for every application. Matching the convention is necessary before comparing algorithms.
Algorithm 1: prepare a response group.
- Freeze the behavior snapshot and sample complete responses for each prompt.
- Store sampled token IDs, valid response masks, and behavior log probabilities.
- Evaluate rewards using the declared verifier, including truncation and invalid-answer rules.
- Compute fixed group means, deviations, and normalized advantages.
- Exclude empty responses explicitly; retain zero-advantage responses without dividing by zero.
- Keep prompts and padding out of the response loss, and record all normalization choices.
The alternative is a learned value function, which can give richer credit assignment but adds training cost and another approximation. A group estimator is simple and critic-free; it can waste a batch when all sampled rewards coincide. These tradeoffs precede the choice between clipping and a smooth gate.
4. Off-policy ratios and the exact meaning of clipping
Let denote the behavior probability at the sampled prefix and the current-policy probability at the same prefix. Define
The prefix must be identical in the numerator and denominator. Re-generating a prefix under the new policy computes a different quantity. Stored behavior log probabilities stay fixed across optimization steps; replacing them by a freshly evaluated current policy can accidentally make every ratio one.
If and , then . This is relative probability change, not a probability: ratios can exceed one. A tiny behavior probability makes ratios sensitive to small absolute changes. Log-space subtraction improves arithmetic, but exponentiation can still overflow when a batch is very stale.
A PPO-style scalar surrogate is
Derive its pieces before describing its behavior. For , multiplying preserves order, so the objective is . For , multiplication reverses order, giving . Away from the boundary, its derivative with respect to is therefore
Clipping is directional. A positive-advantage token with is still active, and a negative-advantage token with is still active when . It is inaccurate to say that every token outside a two-sided band loses its gradient. The favorable movement saturates; the movement that worsens the surrogate retains a corrective gradient.
This distinction is central when evaluating SAPO. Its smooth derivative attenuates deviations on both sides, whereas hard clipping is one-sided after conditioning on the advantage sign. Therefore the experiment changes more than differentiability: it also changes which corrective gradients are suppressed. A clean ablation must separate those mechanisms.
5. Deriving SAPO’s soft surrogate
For a positive temperature parameter , let . The SAPO construction is
The advantage sign selects a fixed temperature during differentiation. The paper’s controlled configuration uses and . In the formula, larger makes the gate narrower; the parameter acts like an inverse smoothing width despite being called a temperature.
Take the derivative in three explicit steps. First set . Second use . Third cancel the chain-rule factor against :
Because and is fixed,
Consequently one token contributes . At , and , so the first-order gradient agrees with the unclipped objective for any positive temperature. The scale factor is what makes this local agreement hold.

Figure 2 subtracts from each curve to align their origins. That subtraction is a constant and does not change gradients. Without it, curves can look different merely because their vertical offsets depend on temperature. Matching objective values is less informative than matching derivatives.
6. Gate weight is not the complete update coefficient
The weight lies between zero and one, but the multiplier on the score vector is . Bounded alone is insufficient to explain the update. For example, at and , while . The token is down-weighted relative to an unclipped importance-weighted update, yet its multiplier exceeds the on-policy value one.
At , the corresponding values are approximately and . At , they are approximately and . Small ratios receive a weak coefficient partly because of , not merely because of gate attenuation. These are deterministic evaluations of the displayed formula, not training measurements.

A useful identity makes the shape transparent:
To derive it, rewrite as and use . Near , put and expand:
There is no linear attenuation at the center. If is small, changing temperature slightly makes only a second-order local difference. For versus , the leading difference is . At this is about . A large long-horizon difference between training runs can therefore reflect rare tails, accumulated feedback, or interacting hyperparameters rather than a large change for a typical near-on-policy token.
For , the coefficient behaves like and tends to zero. For , the gate approaches a finite positive value while . These limits do not bound the norm of the full parameter gradient unless score-vector norms and advantages are also controlled. They do not impose a KL constraint either.
The alternative of an explicit KL penalty constrains a different quantity. It can be combined with a gate but adds another coefficient and a reference distribution. Gradient norm clipping also acts at a different level: it rescales an aggregate vector after token contributions have interacted. Treating these mechanisms as interchangeable hides important failure cases.
7. Why positive and negative updates differ
For logits and sampled token , write . Differentiating gives
Multiplication by advantage changes the direction. A positive advantage raises the sampled logit and lowers the alternatives under gradient ascent. A negative advantage reverses every sign. With probabilities and the first token sampled, the positive gradient is ; the negative gradient is .

The vocabulary argument needs care. Summing the unsampled magnitudes gives , regardless of how many vocabulary entries exist. A large vocabulary does not by itself prove a larger total logit-gradient magnitude. The concern is where that pressure goes, how it couples through shared parameters, and how repeated updates affect useful alternatives.
My interpretation is that sign asymmetry is an empirical design hypothesis supported by an ablation, rather than a universal theorem that negative advantages are harmful. Negative updates can eliminate wrong answers, discourage repetitive text, and support exploration. Suppressing them too strongly can preserve undesirable behaviors. A useful test must measure both stability and the ability to unlearn errors.
8. From token ratios to a sequence statistic
For a response of length , let . The length-normalized sequence ratio is
The geometric mean reduces sensitivity to a single multiplicative outlier compared with the raw sequence product. It can also hide compensating changes: ratios have geometric mean one, although neither token is on-policy. Thus a sequence statistic is a summary, not a complete measure of distribution mismatch.
SAPO’s relationship to a sequence gate uses two distinct approximations: is close to one, and has small variance within a response. The first gives . Define
Taylor-expand each token gate around :
Averaging cancels the linear term because the centered deviations sum to zero. For , two differentiations give
Let . The polynomial ranges from to , so its largest absolute value is two. Therefore , yielding
This is a useful, explicit bound on an average of approximate scalar gates. Its strength depends on temperature as well as variance. Saying that variance is small without its scale and is incomplete. The earlier substitution also has error: when , Taylor’s theorem gives . A gate bound should account for that term separately when ratios are not extremely close to one.
9. Where sequence coherence needs another assumption
The parameter gradient involves a weighted sum of vectors. Define , , and . Even after approximating by one, the difference between a token-weighted gradient and a sequence-weighted gradient is
The scalar Taylor bound controls only the second term. The first is a covariance-like term between token weights and gradient directions. An additional bound on score-vector magnitudes and their variation is needed to control it.
By Cauchy-Schwarz, a valid conservative bound for the first term is
This makes the missing ingredient visible. A tiny average gate error can coexist with meaningful error in the gradient direction, particularly when token gradients cancel. Relative error becomes unstable near a zero aggregate gradient. It is better to report absolute vector error and cosine similarity with a defined low-norm convention.

In Figure 5, , , and illustrative score components are . The average log ratio is zero. The gap between mean approximate gates and their sequence gate is about , below the variance bound. Nevertheless, the exact SAPO coefficient averaged against those score components is about , while the sequence approximation is zero. These values are a diagnostic construction, not evidence about the frequency of such a pattern in trained models.
The example uses moderate rather than infinitesimal deviations, so it does not refute a carefully qualified local approximation. It shows why a scalar histogram cannot by itself establish the desired vector statement. A stronger empirical evaluation would measure gradient-vector agreement across actual minibatches, separate dense from MoE models, and stratify by ratio tails and response length.
10. A loss implementation that preserves the intended derivative
A direct implementation differentiates through . A numerically convenient equivalent subtracts the constant at :
The centered form yields a zero scalar objective at the on-policy point while retaining a nonzero derivative. Therefore zero loss does not imply zero learning. Advantages, behavior probabilities, and chosen temperatures must remain fixed with respect to the current-policy gradient.
Algorithm 2: one SAPO optimization step.
- Read a rollout minibatch with frozen behavior log probabilities and advantages.
- Evaluate current log probabilities by teacher forcing on the stored responses.
- Subtract behavior log probabilities to obtain and calculate .
- Choose or from the fixed advantage sign.
- Calculate the centered sigmoid surrogate and multiply by the fixed advantage.
- Apply the response mask, average by valid response length, then average over the group or batch as declared.
- Negate the objective when using a minimization optimizer; backpropagate once.
- Apply the declared optimizer and any separately specified gradient clipping, then record diagnostics.

An alternative implementation treats as a detached weight on . It gives the same first-order gradient at the current parameter value if the entire multiplier is detached, but it does not define the same higher-order derivative or scalar objective. Those differences matter for meta-gradients and Hessian-based analysis.
A tempting incorrect expression is with no stop-gradient. Its derivative contains , whereas the desired derivative is . A second mistake is differentiating through a behavior model that shares current parameters. A third is averaging across padding or prompt tokens. These mistakes can produce smooth-looking training curves while implementing another objective.
Algorithm 3: minimal numerical audit.
- Use a scalar current log probability and several fixed behavior log probabilities to generate ratios below, at, and above one.
- Approximate the loss derivative by symmetric finite differences in the current log probability.
- Compare it with for a minimization loss, for both signs of and both temperatures.
- Check that the derivative at is and that zero advantage gives zero gradient.
- Test equal-reward groups, empty masks, and unequal response lengths separately.
- Include extreme log ratios to detect numerical failure before attempting a long training run.
For enormous positive log ratios, compute the effective coefficient in log space if using a detached gradient-weight formulation. Since
one can reason about underflow and overflow explicitly. This identity follows by expanding the two logistic denominators. It is not a complete production implementation: itself can overflow, and hard clamping silently changes the objective. A robust system should detect extreme staleness, record it, and apply a declared batch rejection or stabilization policy rather than hiding it inside an undocumented clamp.
11. Design choices and their boundaries
| Choice | Why it is useful | Alternative | Boundary to inspect |
|---|---|---|---|
| Token-level gates | Preserve locally useful parts of heterogeneous responses | One sequence-level multiplier | Per-token control does not solve credit assignment |
| Smooth sigmoid surrogate | Gives continuous derivatives near a clipping threshold | Hard clipping or another smooth kernel | Both tails are attenuated, changing corrective behavior |
| Separate temperatures | Allows different treatment of signed updates | Shared or adaptive temperature | Too much negative suppression may impair unlearning |
| Group advantages | Avoid a learned critic | Value baseline or process reward | Equal rewards provide no relative signal |
| Per-response averaging | Balances responses of different lengths | Token pooling | Changes the length distribution’s effective weighting |
| Reusing rollout data | Amortizes generation cost | Fresh rollouts after every step | Stale prefixes are not corrected by a local token ratio alone |
The last row deserves emphasis. A token ratio compares action probabilities at the observed prefix. It does not fully correct the distribution of prefixes generated by the behavior policy. A complete trajectory importance ratio has different variance and weighting properties. SAPO should be understood as a useful training surrogate, not an automatically unbiased estimator of the new policy’s full trajectory return.
12. Reading the experimental evidence
The controlled experiment starts from a cold-start checkpoint based on Qwen3-30B-A3B-Base, uses mathematical queries, and divides each rollout batch into four update minibatches. The reported comparisons include SAPO, GSPO, and GRPO with routing replay. Figure 7 reproduces the original curves; it supports a qualitative comparison over the displayed run, not an exact numerical leaderboard reconstructed from pixels.

The useful observation is temporal: the curves diverge after an initially shared improvement phase. An endpoint-only comparison can conflate optimization quality with whether a baseline collapsed before the stopping time. I would compare best checkpoint, fixed-step checkpoint, fixed-token checkpoint, and wall-clock budget separately. A method that survives longer can be valuable even if its early sample efficiency is similar.

The three displayed temperature choices are informative about this setup. They do not establish a universal optimum. Because near-center weights change only slightly, I would request tail-conditioned gradient statistics and repeated runs to understand the mechanism. The absence of error bars in these figures prevents estimating the variability of collapse time from the visual evidence alone.

The multimodal comparison uses Qwen3-VL-30B-A3B and two update minibatches per rollout batch. The paper reports evaluations involving AIME25, LiveCodeBench v6, ZebraLogic, and MathVision. Different numbers of sampled responses are used to estimate Pass@1 in different evaluations; this is not the same metric as Pass@k. The average of many independent single-sample correctness indicators estimates the chance that one sample succeeds. It does not estimate the chance that at least one of all samples succeeds.
An aggregate score also requires task weights and uncertainty. For task scores and weights , an aggregate may improve even when one task regresses. I would retain the per-task breakdown, evaluation prompts, decoding parameters, and any benchmark contamination checks. None of these can be recovered reliably from a single aggregate curve.
13. Limitations and reproducibility
The loss is compact; a matching experiment is not. A reproducible comparison needs the initial checkpoint, prompt and reward distributions, optimizer, learning-rate schedule, rollout group size, response length limits, update batch sizes, sampling temperature, routing policy, and evaluation protocol. Referring to another algorithm’s configuration helps locate information but does not substitute for one complete, versioned configuration.
A verifier can also be wrong. Reward hacking, ambiguous answer extraction, and truncation can create misleading advantages. Smoothing a gradient derived from an incorrect reward does not make the reward correct. Dataset quality and solver evaluation remain separate responsibilities.
The displayed experiments do not give a theorem ruling out collapse, a uniform claim over model sizes, or a causal decomposition of every benefit. Their value is to motivate controlled follow-up work. This review did not reproduce model training; the local numerical checks concern the stated objective and explanatory examples only.
14. Independent critical analysis
14.1 A scalar-to-vector step needs stronger justification
The main theoretical concern is the transition from average scalar gates to a full sequence-style gradient. Section 9 isolates the covariance term. An actionable improvement is to state a bound containing score-vector norms and gate-gradient dependence, then measure both terms on training batches. A scalar versus variance plot alone cannot certify alignment of parameter updates.
14.2 Smoothness is entangled with bidirectional suppression
The hard-clipped baseline and the smooth method differ in the treatment of corrective updates outside the opposite boundary. To identify which mechanism helps, compare four variants under equal tuning budgets: directional hard clipping, directional smooth clipping, bidirectional smooth attenuation, and a different smooth kernel with matched local slope and width. Report both validation quality and the fraction of signed learning signal removed.
14.3 The negative-token explanation is incomplete as a norm argument
The softmax derivative redistributes pressure over alternatives, but total unsampled probability mass is bounded by one. More vocabulary entries do not alone produce a larger logit-gradient norm. A sharper explanation would examine probability concentration, score-vector norms, MoE routing transitions, and reward sign together. Measure whether instability correlates with these quantities after controlling for advantage magnitude and policy lag.
14.4 Fairness requires a budget for retuning baselines
Identical hyperparameters do not guarantee equally strong baselines. A new optimizer may simply tolerate settings that are poor for another. Report both a shared-configuration comparison and an equal-search-budget comparison. Keep rollout tokens, optimizer steps, verifier calls, and wall-clock cost explicit. Include several seeds and a prespecified collapse criterion rather than choosing a favorable endpoint after observing the trajectories.
14.5 Stability should include the cost of conservatism
A stable run that learns too slowly is not necessarily a better run. Track reward gains per generated token, time to reach a validation threshold, response length, entropy, and failure to suppress known bad behaviors. A stronger temperature can remove exactly the negative feedback required to correct systematic errors. Targeted unlearning probes would reveal this tradeoff more clearly than aggregate reward alone.
15. A concrete follow-up experiment
I would start with a small model and a fixed, auditable mathematical task set. Freeze the data split before examining curves. Use one loss implementation and vary only the gating function, with routing behavior and normalization explicitly held constant.
Algorithm 4: compare optimization behavior without hiding costs.
- Reserve development prompts for tuning and untouched evaluation prompts for reporting.
- Allocate the same number of configurations and random seeds to every method.
- Record rollout tokens, update tokens, optimizer steps, verifier calls, and elapsed time.
- Log advantage sign, ratio quantiles, quantiles, zero-advantage groups, response lengths, and gradient norms.
- Sample batches for token-versus-sequence gradient cosine similarity and absolute error.
- Report taskwise Pass@1 with uncertainty, collapse timing, and best versus final checkpoint results.
- Repeat with increased rollout lag and longer outputs to test the intended failure boundary.
This proposal is a reproducibility plan, not a completed experiment or a claim about likely numerical gains. Its purpose is to distinguish smoother optimization, stronger conservatism, and better use of heterogeneous samples.
16. Conclusion
SAPO offers a small and interpretable change to a group-based policy objective. Its derivative can be derived exactly, audited cheaply, and understood at both token and sequence scales. I find the implementation idea compelling, while treating broad stability claims and sequence equivalence as questions that need stronger qualification. The most useful next step is a controlled comparison that measures the full update vector and the cost of discarded learning signal, alongside downstream scores.
References and evidence notes
- Gao et al. Soft Adaptive Policy Optimization, v2, 2025. Primary paper; all reproduced training curves are cropped from its Figures 4–6. The note adds original derivations and criticism rather than reproducing its text.
- Shao et al. DeepSeekMath, 2024. Background for the group-relative policy objective; the comparison here uses the precise objective written in SAPO.
- Zheng et al. Group Sequence Policy Optimization, 2025. Background for geometric-mean sequence ratios.
- Qwen SAPO research announcement. Author-team account; distinguish its qualitative interpretation from independent validation.
- ModelScope ms-swift SAPO documentation. An independently maintained implementation reference; not evidence that this review ran that code or reproduced the paper’s training setup.
Source and discussion checks were performed on 2026-09-15. The paper is an existing work from 2025, selected to fill a gap in the current reading collection, not presented as a newly released paper.