Do Larger Reasoning Models Actually Use Tokens More Efficiently?
Starting from an Interview Claim: A Systematic Test with Open-Weight Reasoning Models
The project began with a Yann Dubois interview
This project did not begin with a benchmark table. It began with Yann Dubois's interview on The MAD Podcast with Matt Turck. Dubois, who co-leads OpenAI's post-training frontiers team, made a direct connection between pretrained model scale and reasoning-token efficiency:
“If you have larger models, the amount of thinking time—so the amount of tokens they will think for—will usually decrease. […] The model already thinks through its weights when it generates a certain token. So you can decrease the number of tokens that you need to generate for thinking by increasing the size of the model […] if you basically pre-train larger models, you will get better efficiency.”
—Yann Dubois, excerpt from 24:34–25:23 in the original interview.
This passage directly motivated the study. It connects three claims: a larger pretrained model, more “thinking” already performed through its weights, and fewer explicit thinking tokens. The mechanism feels intuitive—if more capability is internalized in the parameters, inference may not need to reconstruct the same search through an equally long visible trace.
But intuition is not evidence. This is an interview statement, not an official OpenAI technical report; and “usually decrease” is not a universal claim across every family, task, or budget. A test also cannot simply compare leaderboard scores or treat any shorter output as more efficient: the model may have truncated early or failed to emit a final answer. The controlled question is:
Within the same model family, under matched tasks, output budgets, decoding rules, and repeated-sampling protocols, can the larger model reach equal or better quality with fewer visible thinking tokens?
I was also exploring Auto Research: whether frontier models could help formulate hypotheses, design experiments, orchestrate remote runs, audit artifacts, diagnose failures, and converge the evidence into a reproducible report. I therefore turned Yann's claim into a concrete research question.
Public discussions by Noam Brown and Josh McGrath then supplied two missing pieces. Brown stressed that additional test-time compute becomes useful only when the model is capable enough to exploit it. McGrath proposed judging progress on a two-dimensional quality–token plane, where the valuable movement is up and to the left. Together, the three perspectives define the study: Yann supplies the scale–weights efficiency mechanism, Brown the capability precondition, and McGrath the measurement coordinates.
The experimental design follows from that framing: compare two sizes within each family, keep the task and generation protocol fixed, sweep multiple output budgets, repeat each problem, and admit token-efficiency evidence only when final-answer validity, token partitioning, and truncation pass predefined gates.
GPT-5.5 and GPT-5.6 were research collaborators throughout the process. They helped trace and challenge the hypothesis, design the experimental matrix, audit remote runs, diagnose measurement failures, analyze accepted artifacts, and prepare the report, paper draft, and this article. They are not evidence for the conclusion: the empirical claims in this article come from logged jobs and validated result files.
Figure 1 summarizes eight decisions that materially changed the study's direction. Orange side branches are research paths diagnosed and downgraded to supplementary or boundary evidence, not simply failed executions.
From a broad model screen to two controlled families
The project did not begin with GPT-OSS and Qwen3.5 as two hand-picked “winners.” The initial candidate pool also included Ministral-3, Qwen3, and GLM-4.5. We ran pilots first, then asked whether each family could support an interpretable within-family scale comparison.
GPT-OSS had one additional study-design advantage. OpenAI's official model card already supplied a 20B/120B within-family scale pair, the official Harmony format, and multi-effort evaluations on AIME and GPQA-Diamond. Figure 3 in particular places low, medium, and high reasoning effort on an accuracy-versus-average-CoT+Answer plane, where 120B often reaches a comparable or higher-quality region with shorter output. Those results are not part of our experiment and are not counted below, but they made GPT-OSS unusually well suited to a controlled follow-up: comparable scale points, reliable token channels, and a pre-existing quality–token prior.
GPT-OSS and Qwen3.5 Dense survived that screen and formed the cleanest controlled comparisons. Ministral-3, Qwen3, and GLM-4.5 were excluded from the main matrix because of insufficient measurement validity, architectural or training-recipe confounding, real quality reversals, or strong inference-budget interactions. Their results are retained as failure-mode and boundary evidence.
The final evidence scope was frozen by comparability and measurement validity, not by whether a family confirmed the hypothesis. Families outside the main matrix remain part of the result.
Experimental setup
| Family | Smaller | Larger |
|---|---|---|
| GPT-OSS | 20B | 120B |
| Qwen3.5 Dense | 9B | 27B |
The experiment uses a unified K=8 protocol. It compares four models across four tasks—MATH500, GPQA-Diamond, AIME 2024, and AIME 2025—and four maximum output budgets of 4K, 8K, 16K, and 32K. Every model–problem–budget cell contains eight independent rollouts.
| Task | Independent problems n | Rollouts per problem |
|---|---|---|
| MATH500 | 500 | 8 |
| GPQA-Diamond | 198 | 8 |
| AIME 2024 | 30 | 8 |
| AIME 2025 | 30 | 8 |
Inference treats the problem as the independent sampling unit. The paired 95% intervals use 2,000 problem-bootstrap resamples; whenever a problem is sampled, all K=8 rollouts for both model sizes are kept together. K=8 reduces within-problem decoding noise, but AIME still has an effective benchmark sample size of 30 rather than 240.
GPT-OSS high-effort is treated only as an auxiliary analysis and does not gate the primary conclusion; the other candidate families are retained in the failure diagnostics below.
Quality is measured primarily by average accuracy over eight rollouts (avg@8). pass@8 is auxiliary. We separately count thinking tokens, answer tokens, and total generated tokens.
What the final aggregate says
The result is family-conditional, not universal.
| Family | Eligible cells | Joint dominance | Quality-matched savings | Verdict |
|---|---|---|---|---|
| GPT-OSS | 5 | 4 | 4/4 | Strong support |
| Qwen3.5 Dense | 1 | 1 | 1/1 | Strong support |
The table below reports the paired 32K avg@8 effect, defined as larger minus smaller. An interval excluding zero supports the accuracy direction over the observed benchmark-problem distribution. “Ineligible” limits token-efficiency inference only; strict accuracy and its uncertainty are retained.
| Family | Task | n | avg@8 difference | Paired problem-bootstrap 95% CI | Token-efficiency gate |
|---|---|---|---|---|---|
| GPT-OSS | MATH500 | 500 | +1.35 pp | [+0.60, +2.18] pp | eligible |
| GPT-OSS | GPQA-Diamond | 198 | +6.25 pp | [+3.41, +9.34] pp | eligible |
| GPT-OSS | AIME 2024 | 30 | −2.92 pp | [−8.33, +2.50] pp | eligible |
| GPT-OSS | AIME 2025 | 30 | +9.17 pp | [+3.75, +15.00] pp | ineligible |
| Qwen3.5 Dense | MATH500 | 500 | +4.68 pp | [+3.65, +5.73] pp | eligible |
| Qwen3.5 Dense | GPQA-Diamond | 198 | +8.14 pp | [+5.49, +10.98] pp | ineligible |
| Qwen3.5 Dense | AIME 2024 | 30 | +8.33 pp | [+3.33, +14.17] pp | ineligible |
| Qwen3.5 Dense | AIME 2025 | 30 | +15.83 pp | [+8.75, +23.75] pp | ineligible |
Relation to the official GPT-OSS model card. Its Figure 3 already plots 20B/120B quality–CoT+Answer curves on AIME 2025 and GPQA-Diamond. The plotted points align with Table 3's with tools rows and primarily demonstrate test-time scaling as reasoning effort increases. The closest comparison to our no-tools, fixed-medium, hard-cap protocol is instead Table 3's medium / no tools column: GPQA-Diamond is 66.0%→73.1%, nearly identical to our 66.9%→73.1%; AIME 2025 is 72.1%→80.0%, directionally consistent with our 74.2%→83.3%. MATH500 does not appear in the model card.
We therefore do not claim the two-dimensional presentation or the basic GPT-OSS direction as a first discovery. The more accurate contribution is a controlled replication under fixed effort, K=8, and explicit measurement gates, extended with thinking-only tokens, MATH500, and a second family, Qwen3.5. The official numbers serve as external consistency evidence; they are not pooled with our rows, confidence intervals, or acceptance statistics.
GPT-OSS: the clearest positive result
In all five eligible fixed-budget comparisons, GPT-OSS 120B uses fewer thinking tokens than 20B. In four, it is also more accurate. All four quality-matched frontier points show token savings.
- MATH500 improves from 95.8% to 97.2%. The paired effect is +1.35 pp (95% CI [+0.60, +2.18]), while mean thinking falls by 849 tokens (95% CI [−1,028, −685]).
- GPQA-Diamond improves from 66.9% to 73.1%. The paired effect is +6.25 pp (95% CI [+3.41, +9.34]), while mean thinking falls by 2,804 tokens (95% CI [−3,289, −2,323]).
- AIME 2025 improves from 74.2% to 83.3%, a +9.17 pp effect (95% CI [+3.75, +15.00]). The cell is nevertheless ineligible for token-efficiency inference, so only strict accuracy is retained.
AIME 2024 has a directional reversal in the point estimate: accuracy falls from 84.6% to 81.7%. But the paired −2.92 pp effect has a 95% CI of [−8.33, +2.50], and an exact paired sign-flip test is also non-significant (two-sided p=0.383). We therefore treat it as a statistically inconclusive deviation, not an established counterexample. The thinking-token reduction is more stable: a mean difference of −4,252, with a 95% CI of [−6,071, −2,663].
Qwen3.5: positive, but narrower
On MATH500 at 32K, Qwen3.5 27B improves from 93.3% to 98.0%. The paired effect is +4.68 pp (95% CI [+3.65, +5.73]), while mean thinking falls by 1,949 tokens (95% CI [−2,155, −1,752]), corresponding to an 18.7% point estimate.
GPQA-Diamond strict accuracy also rises from 74.1% to 82.2%, a paired +8.14 pp effect (95% CI [+5.49, +10.98]). AIME 2024 and AIME 2025 have effects of +8.33 pp (95% CI [+3.33, +14.17]) and +15.83 pp (95% CI [+8.75, +23.75]). These GPQA/AIME cells fail a final-validity or truncation gate, so they remain strict-accuracy observations and are excluded from token-efficiency inference.
The parser question
Qwen3.5's AIME gate failures could initially be mistaken for a parser problem. Artifact audits show that the thinking/answer partition is valid; the dominant failure occurs during the reasoning-to-final transition, when long reasoning consumes the budget before a terminal answer is emitted.
Increasing the sample count under the same configuration would not remove this systematic finalization behavior. We therefore kept strict accuracy, excluded token-efficiency claims where required, and avoided “repairing” the dataset by selecting only well-finalized traces.
What the three ineligible families taught us
The earlier screening funnel summarized the outcome. Here we separate four diagnostic layers: whether the parser worked, whether the model completed the reasoning-to-final transition, whether scale produced a real quality advantage, and whether architecture or inference settings made the comparison uninterpretable.
| Family | Parser and token partition | Model/scale behavior | Configuration interaction | Classification |
|---|---|---|---|---|
| Ministral-3 | Official parser handles completed streams; parser is not the primary cause. | AIME often fails to finish; 14B is materially worse than 8B on MATH500. | Long caps amplify non-finalization and truncation. | Measurement failure plus a real quality reversal. |
| Qwen3 | No evidence of a primary parser failure. | AIME leans larger; MATH500 favors the smaller model in both quality and tokens. | MoE activation scale, post-training, and incomplete matrix are confounded. | Valid pilot, but scale is not identifiable. |
| GLM-4.5 | Token partition is valid, ruling out a global parser failure. | Air beats full GLM on AIME; GPQA reverses direction. | High truncation and low final validity interact strongly with the 32K cap. | A substantive AIME counterexample, not a stable cross-task law. |
Ministral-3: not merely a parser bug
The official parser correctly handles completed streams; the main AIME failure occurs earlier, when the model spends its reasoning budget without emitting a final answer. MATH500 at 32K also shows a real reversal—7.0% for 14B versus 31.6% for 8B—so parser repair cannot recover the hypothesis. The report describes a 14B→8B→3B pruning-and-distillation cascade and an RL length increase from 32K to 80K after observed truncation. Those details make distillation path and long-reasoning sensitivity plausible clues, not proven causes. The appropriate classification is measurement failure plus a genuine smaller-model win.
Qwen3: successful execution, but scale is not a clean variable
The Qwen3 pilot showed no primary parser failure, but its direction depends on the task: AIME leans toward 235B-A22B, while on MATH500 the 30B-A3B model is both more accurate (94.2% versus 89.8%) and shorter (3.34K versus 4.04K total tokens). The report shows that the pair also differs in active parameters (3B versus 22B), depth, distillation, and post-training; official AIME inference uses a stop-thinking transition and a 38,912-token cap rather than our hard 32K cutoff. This is therefore a valid but confounded pilot, not an identifiable parameter-scale treatment. Future MoE comparisons should control active scale, training recipe, and finalization protocol together.
GLM-4.5: a real counterexample plus inference-configuration sensitivity
At 32K on AIME, GLM-4.5-Air reaches 64.4% avg@8 versus 55.0% for the full model; valid token partitions rule out a parser explanation. GPQA reverses direction and final validity is low, so the counterexample interacts strongly with output cap and exit-reasoning behavior. The report uses hybrid expert distillation and trains reasoning RL directly at 64K; our 32K cap may amplify the issue, but a seed-0 64K diagnostic still did not restore larger-model dominance. The best classification is a genuine AIME counterexample plus configuration sensitivity, best followed up with targeted budget, template, stop-condition, and forced-finalization ablations.
Three recommendations for future community studies
- Verify measurement eligibility before interpreting coverage. Complete coverage does not imply valid finals, valid token partitions, or acceptable truncation. A missing final should count as an error, but its token length should not enter the efficiency comparison.
- Control comparability before attributing an effect to scale. Test multiple output caps and report active MoE parameters, routing, and post-training differences so configuration effects are not mistaken for scale effects.
- Preserve genuine counterexamples. If the smaller model is better under valid measurement, it should constrain the conclusion rather than be relabeled as dirty data.
Exclusion from the main result does not mean experimental failure: it may indicate invalid measurement, an unidentifiable comparison, or a genuine counterexample.
Takeaways: what the experiment supports—and what it does not
Within some controlled model families, a larger model can achieve equal or better quality with fewer visible reasoning tokens. The confidence intervals support GPT-OSS on MATH500 and GPQA-Diamond and Qwen3.5 on MATH500; the admissible Qwen3.5 scope is narrower.
- This is a conditional within-family pattern, not a law of parameter scale. The supported 32K accuracy effects are +1.35 pp for GPT-OSS MATH500, +6.25 pp for GPT-OSS GPQA, and +4.68 pp for Qwen3.5 MATH500, while visible thinking also falls. These results do not establish a universal relationship across architectures or training recipes.
- Reasoning efficiency must combine quality, thinking tokens, and measurement validity. A shorter response that never reaches a valid final answer is not efficient. Final validity, token partitioning, and truncation are part of the metric, not post-hoc cleanup.
- Families outside the main result reveal measurement failure, comparison confounding, and genuine counterexamples. “Scale” is meaningful only when architecture, post-training, active MoE scale, and tokenizer are sufficiently comparable; counterexamples should not be filtered away to manufacture a clean law.
- Uncertainty matters most on small benchmarks. AIME has only 30 independent problems; K=8 stabilizes each problem estimate but does not turn it into 240 independent items. GPT-OSS AIME 2024 is therefore an inconclusive deviation, not an established counterexample. Visible tokens also remain different from total compute.
Research scale
Across the two Codex tasks, the validation required approximately 126.2 hours of effective Codex work—about five days and six hours. This counts active execution, tool use, and remote monitoring, while excluding long idle gaps between working sessions.
The final auditable experiment artifacts processed approximately 874 million tokens, including prompts and model generations. That remains a conservative lower bound: some early explorations and failed runs were not consolidated into the ledger, and the figure excludes tokens consumed by the GPT-5.5/GPT-5.6 research agents themselves.
Sources
- Noam Brown, Scaling Test Time Compute to Multi-Agent Civilizations, Latent Space, June 20, 2025.
- Yann Dubois, OpenAI's Yann Dubois: Why AI Progress Suddenly Feels Real, The MAD Podcast with Matt Turck, May 21, 2026; scale–weights–thinking-token passage at 24:34–25:23.
- Josh McGrath, [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency, December 31, 2025; token-efficiency chapter at 13:12.
- Liu et al., Ministral 3, arXiv:2601.08584, 2026.
- Yang et al., Qwen3 Technical Report, arXiv:2505.09388, 2025.
- Zeng et al., GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models, arXiv:2508.06471, 2025.
这个项目从 Yann Dubois 的一段访谈开始
这个项目的起点不是一张 benchmark 表,而是我读到的一段访谈。共同领导 OpenAI post-training frontiers 团队的 Yann Dubois 在 The MAD Podcast with Matt Turck 中谈到预训练规模与推理效率时,英文原话是:
“If you have larger models, the amount of thinking time—so the amount of tokens they will think for—will usually decrease. […] The model already thinks through its weights when it generates a certain token. So you can decrease the number of tokens that you need to generate for thinking by increasing the size of the model […] if you basically pre-train larger models, you will get better efficiency.”
——Yann Dubois,原始英文访谈节选,24:34–25:23
这段话的核心是:更大的模型通常会用更少的 thinking tokens;模型生成 token 时,已经通过自身权重完成了一部分“思考”,因此扩大模型并预训练更大的模型,往往能提高推理效率。
这段判断直接启发了本研究。它把三个判断连在了一起:更大的预训练模型、已经发生在权重中的“思考”,以及更少的显式 thinking tokens。这个解释直觉上很对——能力如果已经更多地内化进参数,推理时或许就不必用同样长的可见轨迹重新搜索。
但直觉不是证据。这段话来自访谈而不是 OpenAI 的正式技术报告;Dubois 说的也是“通常会减少”,并不等于对所有模型家族、任务和预算都成立。要验证它,不能只比较 leaderboard 分数,也不能把输出更短直接当作更高效:一个模型可能只是更早截断,甚至没有给出最终答案。真正需要问的是:
在同一模型家族内,当任务、输出预算、解码规则和重复采样协议保持一致时,较大模型能否用更少的可见 thinking tokens,达到相同或更高的质量?
当时我正在探索 Auto Research:能否让前沿模型参与提出假设、设计实验、编排远端任务、审计产物、诊断失败,再把证据收敛成可复核的报告。因此,我把 Yann 的判断设为一个具体研究问题。
在把直觉变成实验时,Noam Brown 和 Josh McGrath 的公开讨论又补齐了两块:Brown 强调额外的测试时计算只有在模型能力足够时才有价值;McGrath 建议把模型进步画在“质量—token”二维平面上,看它是否真正向左上方移动。三者合在一起,形成了完整的实验逻辑:Yann 提供效率机制的直觉,Brown 给出能力前提,McGrath 给出测量坐标。
实验设计也由此推出:在每个家族内比较两个规模,固定任务和生成协议,扫描多档输出预算,对每道题进行重复采样;只有最终答案有效率、token 分区和截断率都通过预设门禁时,才允许进入 token 效率比较。
GPT-5.5 和 GPT-5.6 作为研究协作者参与了整个流程:追溯和质疑原始假设、设计实验矩阵、审计远端任务、诊断测量失败、分析验收后的产物,以及协助准备实验报告、论文初稿和这篇文章。它们本身不是结论的证据;文章中的实证数字均来自可复核的任务日志和通过验收的结果文件。
图 1 概括了研究过程中八个关键决策节点。橙色侧枝表示被诊断后降级为补充或边界证据的研究路线,而不是简单的实验执行失败。
从广泛选型到两个正式实验家族
项目并不是从 GPT-OSS 和 Qwen3.5 两个“答案正确”的家族开始。最初的候选池还包括 Ministral-3、Qwen3 和 GLM-4.5。我们先做小规模 pilot,再检查能否形成可解释的同家族尺度对照。
GPT-OSS 还有一个额外的先验优势:OpenAI 的官方 model card已经提供了 20B/120B 同家族尺度对、官方 Harmony 格式,以及 AIME 与 GPQA-Diamond 的多档 reasoning-effort 评测。尤其是其中的图 3,它把 low、medium、high 三档推理强度画在“准确率—平均 CoT+Answer tokens”平面上,已经显示 120B 往往能以更短输出进入相近或更高质量区。这不是我们实验的一部分,也不计入下文统计,但它说明 GPT-OSS 同时具备可比较尺度、可靠 token 通道和现成的质量—token 先验,是很适合进一步做受控验证的模型家族。
筛选之后,GPT-OSS 和 Qwen3.5 Dense 能形成最清楚的受控对照,因此进入面向公开结论的主矩阵。Ministral-3、Qwen3 和 GLM-4.5 则因测量有效性不足、尺度变量存在架构或训练配方混杂,或出现真实质量反转与明显的推理预算交互,没有纳入主矩阵;相关结果作为失败模式与边界证据单独分析。
最终证据范围是按可比性与测量有效性冻结的,而不是按某个家族是否支持原始假设来筛选。未进入主矩阵的家族同样构成实验结果。
实验设置
| 家族 | 较小模型 | 较大模型 |
|---|---|---|
| GPT-OSS | 20B | 120B |
| Qwen3.5 Dense | 9B | 27B |
实验采用统一的 K=8 协议。主实验比较 4 个模型,覆盖 MATH500、GPQA-Diamond、AIME 2024 和 AIME 2025 4 个任务,扫描 4K、8K、16K 和 32K 4 档最大输出预算;每个“模型—题目—预算”单元进行 8 次独立 rollout。
| 任务 | 独立题目数 n | 每题 rollout |
|---|---|---|
| MATH500 | 500 | 8 |
| GPQA-Diamond | 198 | 8 |
| AIME 2024 | 30 | 8 |
| AIME 2025 | 30 | 8 |
统计推断把题目作为独立抽样单位。95% 区间采用 2,000 次按题配对 bootstrap;每次抽中一道题时,会同时保留大小模型在该题上的全部 K=8 rollouts。K=8 能降低题内解码噪声,但 AIME 的有效基准样本量仍是 30,而不是 240。
GPT-OSS high-effort 仅作为辅助分析,不参与主结论门禁;其他候选家族则保留在后文的失败诊断中。
质量指标以 8 次 rollout 的平均准确率 avg@8 为主,pass@8 为辅助。token 成本被拆成 thinking、answer 和 total generated tokens。
最终聚合结果
结论不是普适规律,而是家族条件成立。
| 家族 | 合格单元 | 质量与 thinking 同时占优 | 等质量节省 | 结论 |
|---|---|---|---|---|
| GPT-OSS | 5 | 4 | 4/4 | 强支持 |
| Qwen3.5 Dense | 1 | 1 | 1/1 | 强支持 |
下表给出 32K 下“较大模型减较小模型”的配对 avg@8 效应。区间不跨 0,表示在当前题目分布上准确率方向获得统计支持。“不合格”只限制 token 效率推断;严格准确率及其不确定性仍然保留。
| 家族 | 任务 | n | avg@8 差值 | 按题配对 95% CI | token 效率门禁 |
|---|---|---|---|---|---|
| GPT-OSS | MATH500 | 500 | +1.35 pp | [+0.60, +2.18] pp | 合格 |
| GPT-OSS | GPQA-Diamond | 198 | +6.25 pp | [+3.41, +9.34] pp | 合格 |
| GPT-OSS | AIME 2024 | 30 | −2.92 pp | [−8.33, +2.50] pp | 合格 |
| GPT-OSS | AIME 2025 | 30 | +9.17 pp | [+3.75, +15.00] pp | 不合格 |
| Qwen3.5 Dense | MATH500 | 500 | +4.68 pp | [+3.65, +5.73] pp | 合格 |
| Qwen3.5 Dense | GPQA-Diamond | 198 | +8.14 pp | [+5.49, +10.98] pp | 不合格 |
| Qwen3.5 Dense | AIME 2024 | 30 | +8.33 pp | [+3.33, +14.17] pp | 不合格 |
| Qwen3.5 Dense | AIME 2025 | 30 | +15.83 pp | [+8.75, +23.75] pp | 不合格 |
与官方 GPT-OSS model card 的关系。官方图 3 已经在 AIME 2025 和 GPQA-Diamond 上画过 20B/120B 的质量—CoT+Answer 曲线;图中点位与表 3 的 with tools 结果一致,主要用于说明提高 reasoning effort 带来的 test-time scaling。与我们无工具、固定 medium effort、四档硬预算的设置最接近的是表 3 的 medium / no tools:GPQA-Diamond 为 66.0%→73.1%,与我们的 66.9%→73.1% 几乎一致;AIME 2025 为 72.1%→80.0%,也与我们的 74.2%→83.3% 同向。MATH500 没有出现在 model card 中。
因此,我们不把二维图或 GPT-OSS 上的基本方向主张为首次发现。更准确的定位是:对官方先验做固定策略、K=8 和测量门禁下的受控复验,并用 thinking-only token、MATH500 与 Qwen3.5 把证据扩展到官方 model card 没有覆盖的分析口径和模型家族。官方结果只作为外部一致性证据,不与我们的实验行数、置信区间或验收统计合并。
GPT-OSS:最清晰的正向结果
在 5 个通过门禁的固定预算比较中,GPT-OSS 120B 的 thinking tokens 全部少于 20B;其中 4 个单元同时具有更高准确率。4 个等质量前沿点也全部显示 token 节省。
- MATH500 从 95.8% 提升到 97.2%,配对效应为 +1.35 pp(95% CI [+0.60, +2.18]);平均 thinking 减少 849 tokens(95% CI [−1,028, −685])。
- GPQA-Diamond 从 66.9% 提升到 73.1%,配对效应为 +6.25 pp(95% CI [+3.41, +9.34]);平均 thinking 减少 2,804 tokens(95% CI [−3,289, −2,323])。
- AIME 2025 从 74.2% 提升到 83.3%,配对效应为 +9.17 pp(95% CI [+3.75, +15.00])。但该单元未通过 token 效率门禁,因此只保留严格准确率结论。
AIME 2024 的点估计出现方向性反转:准确率从 84.6% 降到 81.7%。但配对效应 −2.92 pp 的 95% CI 为 [−8.33, +2.50],精确配对符号翻转检验也不显著(双侧 p=0.383)。因此我们把它视为统计上尚不确定的偏离,而不是已确立的反例。thinking-token 减少更稳定:平均差值 −4,252,95% CI 为 [−6,071, −2,663]。
Qwen3.5:方向一致,但证据更窄
在 MATH500 的 32K 设置下,Qwen3.5 27B 的准确率从 93.3% 提升到 98.0%。配对效应为 +4.68 pp(95% CI [+3.65, +5.73]);平均 thinking 减少 1,949 tokens(95% CI [−2,155, −1,752]),点估计降幅为 18.7%。
GPQA-Diamond 的严格准确率也从 74.1% 提升到 82.2%,配对效应为 +8.14 pp(95% CI [+5.49, +10.98])。AIME 2024 和 AIME 2025 的效应分别为 +8.33 pp(95% CI [+3.33, +14.17])和 +15.83 pp(95% CI [+8.75, +23.75])。但这些 GPQA/AIME 单元未通过 final validity 或截断率门禁,因此只作为严格准确率结果,不进入 token 效率推断。
到底是不是解析器的问题?
Qwen3.5 的 AIME 门禁失败最初可能被误判为解析器问题。产物审计表明,thinking/answer token 分区有效;主要失败发生在 reasoning-to-final 转换阶段:长推理耗尽预算后,模型未能输出终止性的最终答案。
因此,按同一配置增加样本量不会消除这一系统性收尾问题。我们保留严格准确率,并按门禁将相关单元排除在 token 效率分析之外,也不通过筛选“收尾良好”的样本来人为修复数据。
三个未通过家族告诉了我们什么
前面的选型漏斗只概括了筛选结果。这里进一步把失败拆成四层:解析器是否正确、模型是否完成 reasoning-to-final、大小模型是否真的形成质量优势,以及推理预算/架构配置是否让比较失去可解释性。
| 家族 | 解析器与 token 分区 | 模型/尺度现象 | 配置交互 | 最终判断 |
|---|---|---|---|---|
| Ministral-3 | 完整生成流可被官方解析器处理;解析器不是主因。 | AIME 经常没有完成 reasoning-to-final;MATH500 上 14B 明显弱于 8B。 | 长预算放大未收尾与截断。 | 测量不合格,同时存在真实质量反转。 |
| Qwen3 | 没有证据表明解析器是主要故障。 | AIME 偏向大模型,MATH500 却由小模型同时赢得质量与 token。 | MoE 激活规模、后训练和不完整矩阵混杂。 | pilot 有效,但无法归因于“规模”。 |
| GLM-4.5 | token 分区有效,排除全局解析器故障。 | AIME 上 Air 胜过完整模型;GPQA 方向相反。 | 高截断、低 final validity,与 32K cap 强交互。 | AIME 是实质反例,但不能形成稳定跨任务规律。 |
Ministral-3:不是单纯的解析器故障
官方解析器能够正确处理完整生成流;AIME 的主要问题是模型耗尽 reasoning 预算后仍没有输出 final,因此不是单纯的 parser 故障。与此同时,MATH500 32K 上 14B 只有 7.0%,明显低于 8B 的 31.6%,说明还存在真实尺度反转。技术报告显示该家族沿 14B→8B→3B 级联剪枝与蒸馏,并曾因截断把 RL 生成上限从 32K 提高到 80K:这让蒸馏路径和长推理敏感性成为合理线索,但不足以完成因果归因。最终分类是“测量不合格 + 小模型真实胜出”。
Qwen3:运行成功,但“规模”不是干净变量
Qwen3 pilot 没有主要解析器故障,但结果依赖任务:AIME 偏向 235B-A22B,MATH500 却由 30B-A3B 同时赢得准确率(94.2% 对 89.8%)和 total tokens(3.34K 对 4.04K)。技术报告表明两者不仅总参数不同,激活参数(3B 对 22B)、层数、蒸馏和后训练路线也不同;官方 AIME 还使用“停止思考再作答”机制与 38,912 token 上限,而我们的设置是硬 32K。因而这是一组有效但混杂的 pilot,不能把差异归因于单一的“规模”;后续 MoE 对照必须同时控制激活规模、训练配方与收尾协议。
GLM-4.5:真实反例叠加推理配置敏感性
32K AIME 上,GLM-4.5-Air 的 avg@8 为 64.4%,高于完整模型的 55.0%;两者 token 分区有效,因此不能归咎于 parser。GPQA 的方向却相反,而且 final validity 偏低,说明这个反例与预算和退出推理行为存在强交互。技术报告采用 hybrid expert distillation,并直接在 64K 长度上做 reasoning RL;我们的 32K cap 可能放大问题,但 seed0 的 64K 诊断仍未恢复大模型优势。最稳妥的判断是“真实 AIME 反例 + 配置敏感性”,下一步应做预算、模板、终止条件和强制收尾的小规模消融。
给后续社区实验的三点建议
- 先确认结果“能测量”。覆盖完整不代表 final validity、token partition 和截断率合格;缺失 final 应计为错误,但其 token 长度不应进入效率比较。
- 先控制可比性,再讨论规模效应。应比较多档输出预算,并报告 MoE 激活参数、路由和后训练差异,避免把配置差异误认为规模效应。
- 保留真实反例。如果较小模型在有效测量下确实更好,这一结果应当限制结论,而不是被归类为“脏数据”。
没有进入主结果不等于实验失败:它可能代表测量不合格、比较不可识别,或一个真实反例。
Takeaways:实验支持什么、不支持什么
在部分受控模型家族内,较大模型可以用更少的可见推理 token 达到相同或更高质量。置信区间支持 GPT-OSS 在 MATH500、GPQA-Diamond 以及 Qwen3.5 在 MATH500 上的方向;Qwen3.5 可进入 token 效率推断的范围更窄。
- 这是有条件的家族内规律,不是参数规模定律。32K 下获得统计支持的准确率效应为:GPT-OSS MATH500 +1.35 pp、GPT-OSS GPQA +6.25 pp、Qwen3.5 MATH500 +4.68 pp,同时可见 thinking 也减少;这些结果不能外推为跨架构、跨训练配方的普适关系。
- 推理效率必须联合衡量质量、thinking tokens 与测量有效性。更短但没有完成 final 的响应不是高效;final validity、token partition 和截断率是指标的一部分,而不是事后清洗。
- 未进入主结果的家族分别揭示了测量失效、比较混杂和真实反例。“规模”只有在架构、后训练、MoE 激活规模和 tokenizer 足够可比时才是有意义的变量,不能通过筛掉反例制造整齐规律。
- 小数据集尤其需要报告不确定性。AIME 只有 30 道独立题目;K=8 能稳定每道题的估计,却不会把它变成 240 个独立样本。因此 GPT-OSS AIME 2024 只能视为尚不确定的偏离,而不是已确立的反例。可见 token 也不等于总计算成本。
这次验证的投入规模
跨前后两个 Codex 任务累计,这次验证约投入 126.2 小时的有效 Codex 工作时间,也就是 5 天 6 小时。这个口径统计实际执行、工具调用和远端监控时间,不把两次工作之间的长空档算作投入。
最终可审计的实验产物共处理约 8.74 亿 tokens,包含 prompt 与评测模型生成内容。它仍是保守下界:部分早期探索、失败任务没有进入统一账本,也不包含 GPT-5.5/GPT-5.6 研究代理自身消耗的 tokens。
来源
- Noam Brown,Scaling Test Time Compute to Multi-Agent Civilizations,Latent Space,2025 年 6 月 20 日。
- Yann Dubois,OpenAI's Yann Dubois: Why AI Progress Suddenly Feels Real,The MAD Podcast with Matt Turck,2026 年 5 月 21 日;规模—权重—thinking token 原话位于 24:34–25:23。
- Josh McGrath,[State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency,2025 年 12 月 31 日;token efficiency 章节从 13:12 开始。
- Liu 等,Ministral 3,arXiv:2601.08584,2026。
- Yang 等,Qwen3 Technical Report,arXiv:2505.09388,2025。
- Zeng 等,GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models,arXiv:2508.06471,2025。