Unlearnable, or Unmeasured? On the Reliability of Difficulty Labels in RLVR

Chandak Chakma, Syed Nazmus Sakib, Nafiul Haque, Shifat E. Arman*

Department of Robotics and Mechatronics Engineering, University of Dhaka

NeurIPS 2026 Workshop: Transitioning from Pre-Training to Post-Training

*Correspondence to Shifat E. Arman (shifatearman@du.ac.bd).

Bar chart of the size of the unlearnable set after the no_reward exclusion under four constructions of the same 292 sub-threshold candidates: union of no_reward over 5 seeds (published) gives 74, single-seed no_reward gives 147.8 on average (145 to 150), zero correct in 128 gives 181, and intersection of no_reward over 5 seeds gives 228.
The same 292 candidate hard prompts, four ways of building the “unlearnable” set. Changing only how the five training seeds are combined moves the set from 74 to 228 prompts, while any single seed on its own gives 145–150. The label depends more on the aggregation rule than on which model was trained.

Overview

Reinforcement learning with verifiable rewards (RLVR) is a central tool for improving reasoning in post-training. Recent work reports that some hard prompts stay “unlearnable”: they occasionally produce a correct solution but barely improve during training, which has been linked to low gradient similarity and a possible representation failure. We revisit this phenomenon and ask a prior question: is the set of prompts used to support the claim measured reliably?

Difficulty labels are estimated from a limited number of sampled responses, and combining them across seeds can change which prompts are selected rather than simply reduce noise.

Key findings:

The slow-learning phenomenon survives our reanalysis, but both the prompts used to define it and the evidence used to explain it require more careful measurement.

Motivation

Prompt difficulty has become an operational quantity in RLVR. Methods use a model's empirical success rate to decide which prompts to replay, filter, route, reweight, or give extra training effort, and difficulty is also used to probe the limits of RLVR itself. Most strikingly, Chen et al. (2026) identified hard prompts that occasionally succeed yet barely improve, called them unlearnable, and connected their behavior to lower within-group gradient similarity.

All of these uses depend on identifying hard prompts reliably. Difficulty is estimated from sampled responses, so the same prompt can land on either side of a threshold across repeated evaluations even when the model is unchanged. We treat the aggregation rule and the rollout budget as part of the measurement, and ask when a prompt-level trainability label is stable enough to support a scientific claim.

Setup

We reproduce the unlearnability study of Chen et al. (2026) and keep its published cohorts, threshold, and verifier.

ComponentConfiguration
ModelQwen2.5-0.5B trained with GRPO, G = 8 rollouts per prompt, five independent seeds.
Data1,023 prompts from MATH; the original answer verifier, unmodified.
CohortsThe published easy, learnable, and unlearnable groups; difficulty threshold τ = 0.1. Du denotes the set produced by the published unlearnability construction.
Set membershipN = 128 evaluation rollouts per prompt; independent sets compared by Jaccard similarity.
GradientsN = 200 initial-policy rollouts; each prompt-level gradient averages its correct rollouts.
Uncertainty95% bootstrap intervals, resampling prompts at the level appropriate to each statistic.

Our run is not a byte-for-byte rerun. The key difference is a 1,024-token generation cap instead of the original 5,120, which we test directly on the evaluation side; the full constant-by-constant audit is below.

“Unlearnable” Prompts Do Learn, Slowly

Training reward over 120 optimizer steps for the published cohorts. Easy prompts rise from about 0.15 to 0.62, learnable prompts from about 0.05 to 0.15, and unlearnable prompts from about 0.01 to 0.10: low, but not flat.
Training reward for the published easy, learnable, and unlearnable cohorts. The separation is present; the flatness is not.

The three published groups stay clearly separated throughout training, but the Du group is not flat. Reward slopes per 100 optimizer steps:

+0.380easy
+0.112learnable
+0.042Du [+0.026, +0.059]

The Du interval excludes zero: these prompts improve at roughly one third of the learnable rate while staying at a much lower reward level.

We use the slope, not the terminal reward, because GRPO drops a prompt from the update once its sampled rewards have zero variance, so the prompts still active late in training are not the same prompts as early on. Throughout, “unlearnable” refers to the published label, not to a claim that these prompts cannot improve.

The Aggregation Rule Changes the Set

To separate training variation from the aggregation rule, we fix a common candidate pool: evaluating all 1,023 prompts 128 times on one reference checkpoint leaves 292 prompts at or below τ = 0.1. From these same candidates:

Three panels. Left: set size falls as more seeds are intersected, from about 170 to 41 for unlearnable and 190 to 122 for learnable. Middle: many prompts are flagged by only one to four of five seeds. Right: prompts flagged by fewer seeds cluster just above the threshold tau = 0.1.
Left: set size keeps falling as more seeds are combined instead of settling. Middle: many prompts are flagged by only some seeds. Right: that disagreement is concentrated near τ = 0.1, where one extra success moves a prompt across the threshold.

What fails is the estimator, not the target. For a fixed policy with true success probability p(x), the set Dτ = {x : p(x) ≤ τ} is well defined. But requiring a prompt to be labeled difficult in every independent run means that, as runs are added, only prompts almost guaranteed that label survive, pushing the retained set toward p(x) = 0. The no_reward exclusion then removes exactly those prompts. The procedure does not converge to Dτ; it converges to a degenerate outcome of the aggregation rule. Pooling the same rollouts into one deeper estimate avoids this.

“No reward was observed in training” is not the same as “unsolvable”: of the 218 prompts removed by the no_reward criterion, 116 are solved at least once in a later N = 128 evaluation.

Most of the Disagreement Is Evaluation Noise

To remove training variation entirely, we evaluate the same checkpoint twice, changing only the sampled rollouts. At N = 32 the two evaluations agree at Jaccard 0.798, versus 0.751 across independently trained seeds: resampling one fixed model reproduces about 81% of the apparent cross-seed disagreement. A binomial model of the threshold decision predicts the same-model agreement to within 0.004.

A longer generation cap does not explain the movement either. Re-evaluating all prompts with 5,120 instead of 1,024 tokens cuts truncation from 9.6% to 3.2%, yet the two sets agree at 0.898, no less than two passes at the same cap (0.877).

Table: Key robustness controls.

ControlComparisonResultInterpretation
Same-model resamplingfixed checkpoint vs. cross-seed, N = 320.798 vs. 0.751Most apparent cross-seed disagreement can arise from evaluation sampling alone.
Token-cap sensitivity1,024 vs. 5,120 tokens0.898 JaccardSet movement is no larger than the measured repeatability scale.
Equal-cost allocationone N = 128 pass vs. 4×320.877 vs. 0.791A deeper estimate is more reproducible at the same rollout cost.

How Many Rollouts Does a Difficulty Label Need?

Spend the budget deeply, not in shallow passes. With the same 128 rollouts per prompt, one deep estimate thresholded once agrees across repeats at Jaccard 0.877, while four blocks of 32 intersected agree at only 0.791 (difference +0.086 [+0.038, +0.135]). Both rules use the same effective threshold, ⌊0.1 × 128⌋/128 = ⌊0.1 × 32⌋/32 = 0.09375, so the gain comes from allocation, not from extra compute.

Predicting reproducibility. With N rollouts, a prompt with success probability p is flagged with probability qN(p) = Pr(K ≤ ⌊τN⌋), where K ~ Binomial(N, p). We estimate the distribution F of latent success probabilities from the pooled N = 128 evaluation with the Kiefer–Wolfowitz nonparametric MLE, which accounts for the noise in observed pass rates, and predict the agreement of two independent evaluations:

JN = ∫ qN(p)2 dF(p)  /  ∫ [2qN(p) − qN(p)2] dF(p)
Top: predicted Jaccard between two independent passes rises with per-prompt budget N, in a sawtooth, from about 0.63 at N=8 to about 0.88 at N=160; the shallow four-block rule sits below it. Bottom: the effective threshold floor(tau N)/N saws between about 0.05 and 0.10.
Predicted agreement versus per-prompt budget N. The sawtooth comes from the integer threshold (bottom), so we report the first budget from which a target stays satisfied.

For Qwen2.5-0.5B on MATH at τ = 0.1:

≈ 40rollouts for 0.80 agreement
242projected for 0.90
2,111projected for 0.95

The 0.90 and 0.95 budgets lie beyond the N = 128 fitting range. These numbers depend on how many prompts sit near the threshold, so they do not transfer to other models, datasets, or thresholds. What transfers is the procedure for computing them.

Table: Predicted Jaccard agreement at representative budgets (deep, single-pass rule).

N rollouts per prompt8163264128256512
Predicted Jaccard0.63220.72650.79980.84130.87590.90220.9171

The Budget Model Predicts Unmeasured Budgets

Predicted versus measured Jaccard. Points lie close to the y = x line within a ±0.03 band: three retrospective Qwen 0.5B points, three cross-model consistency points, and two red diamonds for the prospectively frozen predictions, which sit slightly above the line.
Predicted versus measured agreement. Filled diamonds are the two prospectively frozen predictions; the Qwen2.5-1.5B point is excluded from validation.

We froze and timestamped predictions for two budgets that had not yet been measured, then generated 212,784 new responses:

BudgetFrozen predictionMeasured
N = 400.8131 [0.7699, 0.8533]0.8450 [0.8063, 0.8833]
N = 640.8397 [0.7980, 0.8792]0.8742 [0.8369, 0.9091]

Both measurements fall inside their frozen intervals. The model underpredicts by about 0.03 in both cases, which is conservative: it recommends slightly more sampling, not less.

Set size alone is not enough. A null model that gives every prompt the same success probability matches the number of flagged prompts by construction but misses the measured Jaccard by 0.19–0.53, while the distributional model misses by only 0.008–0.032. What matters is how prompt probabilities are spread around the threshold.

Table: Retrospective and cross-model consistency checks. Cross-model rows reuse the pass their distribution was estimated from, so they are consistency checks rather than independent validation; the Qwen2.5-1.5B row (greyed) is an excluded diagnostic.

Model / ruleMeasuredPredictedError
Qwen2.5-0.5B, one pass N = 320.79770.8000+0.0023
Qwen2.5-0.5B, one pass N = 1280.87700.8759−0.0011
Qwen2.5-0.5B, four blocks of 320.79060.8228+0.0322
Qwen2.5-0.5B, two blocks of 160.70280.7248+0.0220
Llama-3.2-3B, two blocks of 160.68440.6763−0.0081
Llama-3.2-3B v2, two blocks of 160.70700.6926−0.0144
Qwen2.5-1.5B, two blocks of 160.68490.6718−0.0131

The Gradient Gap Shrinks When Sample Count Is Matched

Left: within-group mean gradient cosine rises with the number K of correct rollouts averaged, for D_u, learnable and easy prompts. Right: the easy-to-D_u similarity ratio is about 1.4 to 1.6 at matched K = 4, 6, 8 and 11, compared with 2.33 unmatched; every matched value stays above 1.

The gradient explanation compares how similar per-prompt gradients are within each group, but each prompt-level gradient averages its correct rollouts, and hard prompts have far fewer: in the N = 200 cohort, a median of 8 for Du versus 65 for easy prompts.

2.327×easy / Du ratio, unmatched
1.539×matched at K = 11 [1.26, 1.88]
48.9%of the log-gap removed [29.3, 71.2]

Every matched level (K = 4, 6, 8, 11) lies below the unmatched ratio and above one, so matching weakens the separation substantially without eliminating it. Matching also selects: at K = 11 only 28 of the 71 Du prompts remain, and they are the easier ones, so this shows sensitivity to sample count, not a corrected full-cohort estimate.

Width of the 95% interval that set-membership uncertainty adds to the gradient-similarity ratio, falling with per-prompt budget N from about 0.41 at N = 8 to about 0.12 at N = 200, crossing a 0.25 materiality threshold around N = 32.
Set uncertainty propagates downstream: at N = 8 it alone adds a 95% interval width of 0.41 to the gradient ratio, falling to 0.14 by N = 128.

Two Simple Optimization Explanations Do Not Hold

Exposure starvation. With G = 8, a prompt only updates the policy when its sampled rewards differ. Du prompts contribute on 2.14 steps per run, versus 3.13 for learnable and 5.26 for easy prompts. The gap is real, but none of the 156 prompts in the exposure cohort is fully starved across all five runs, and 22.4% receive at least the learnable group's mean exposure while still improving slowly.

GRPO normalization. If reward-std normalization suppressed hard prompts, Du updates would be smaller. The opposite holds: when a prompt does contribute, the mean 1/σ scale is 2.836 for Du and 2.426 for easy prompts, about 17% larger. The asymmetry is in how often an update forms, not in its scale.

Content starvation, where rewarded trajectories carry weak learning signal, remains possible; we cannot test it because the training-time responses were not stored. The cause of the slow-learning effect is still unresolved.

An Unstable Label Does Not Necessarily Mean Unstable Training

Does a noisy difficulty rule make difficulty-aware RLVR ineffective? We train five seeds under two neighboring constructions that remove very different amounts of usable training signal, 5.0% versus 27.6% after GRPO's zero-variance filter, alongside size-matched random-removal controls. The resulting models differ by less than one held-out accuracy point. The experiment cannot resolve effects that small, so this is not evidence of equivalence, but it bounds our claim: a large change in the reported set need not produce an equally large change in training. Our conclusion is about measurement and inference, not a claim that every method using noisy difficulty signals must fail.

Replication Audit

We audited the effective configuration from the actual training and evaluation paths: of 47 effective constants, 30 match the original and the 17 below differ. The generation cap matters most, since it affects hard prompts more than easy ones; its evaluation-side effect is tested above, but our training-dependent magnitudes should not be read as an exact reproduction of the original configuration. All runs used a single NVIDIA RTX 3090 Ti (24 GB).

Configuration constantOursOriginal
Training
Dataset namesimplelr_qwen_level1to4_sub1ksimplelr_qwen_level1to4
Maximum response length1,0245,120
Train batch size32256
PPO mini-batch size1664
PPO micro-batch size232
Log-probability micro-batch size4128
Micro rollout batch size1281,024
Tensor parallelism12
GPUs per node14
Rollout GPU memory utilization0.450.75
Total epochs1250
Save frequency4020
Test frequency100,0005
Validation batch size2001,000
Evaluation
Test dataTraining subset used hereOriginal test split
Maximum generation length1,0245,120
GPU memory utilization0.600.75

Conclusion

We revisited the unlearnability phenomenon in RLVR and found a more nuanced result. The affected prompts improve at roughly one third of the learnable rate rather than not at all, but the prompts used to define them are far less stable than the effect itself. The published aggregation rule does not consistently recover a fixed-threshold set, and finite-sample evaluation explains much of the apparent disagreement across training seeds.

We provide a way to determine how much evaluation reproducible difficulty assignment requires, and validate its predictions at previously unmeasured budgets. The gradient-similarity difference proposed to explain unlearnability shrinks substantially when estimator sample count is matched, although a residual gap remains. The lesson is not that unlearnability disappears, but that low-budget difficulty labels are unreliable for classifying individual prompts as “unlearnable,” or for supporting mechanistic claims built on that classification.

BibTeX

@inproceedings{chakma2026unlearnable,
    title={Unlearnable, or Unmeasured? On the Reliability of Difficulty Labels in RLVR},
    author={Chakma, Chandak and Sakib, Syed Nazmus and Haque, Nafiul and
            Arman, Shifat E.},
    booktitle={NeurIPS 2026 Workshop: Transitioning from Pre-Training to Post-Training},
    year={2026}
}