Reasoning models

Reasoning models, honestly

Reasoning models are how labs buy accuracy: RL on verifiable rewards over long chain-of-thought, then distillation. The recipe works — and enterprise adoption is a procurement question. The evidence, the serving bill, and what to measure first.

STEAV13 min read

What a reasoning language model is

A reasoning language model generates an extended internal chain-of-thought — thousands to tens of thousands of tokens — before producing a final answer, and is trained so that this deliberation improves correctness on math, code and multi-step logic. You have read its output: a worked solution with dead ends crossed out, a program with its test reasoning attached, a plan that checks itself before committing.

OpenAI's o1, in September 2024, introduced this as a distinct inference paradigm: accuracy bought with test-time compute rather than with a larger training run. The step that made the paradigm reproducible for everyone else came four months later. DeepSeek-R1 (arXiv 2501.12948, January 2025) showed that reinforcement learning with rule-based verifiable rewards on top of a base model can elicit long chain-of-thought — including self-verification and reflection — without a learned reward model. DeepSeek's own reported results for R1: 79.8% on AIME 2024, 97.3% on MATH-500, 71.5% on GPQA Diamond. Those are vendor-reported numbers, but the recipe was released with the weights, and that is what changed the field.

The naming matters for reading the literature. The training signal is called RLVR — reinforcement learning with verifiable rewards — a term coined by Ai2's Tülu 3 (arXiv 2411.15124), which formalized verifier-function rewards for math and instruction following, and made mainstream by R1. A reward is verifiable when a rule can score it: exact answer match, unit tests, a checker. No learned reward model sits in the loop, which means the reward is auditable line by line.

Reasoning is a training recipe, not a model type. GRPO, distillation, rejection sampling are configuration choices inside a pipeline, and a recipe can be measured, compared and owned. That is the lens for this post: what the recipe is, what the evidence says it buys, what the open-weight landscape looks like, and what it costs to run.

The recipe, step by step

To train a reasoning model in 2026 is to run a loop that R1 made public. Start from a base model. Optionally cold-start with a small set of long-chain-of-thought examples so the model's outputs are parseable before RL begins. Then run the RL stage that the whole recipe is named for.

THE RECIPE, STEP BY STEPBase model, cold-start SFT, GRPO-family RL against a verifiable reward, rejection sampling, and an optional distillation into smaller models. The dashed arrow is R1's own loop.
After DeepSeek-R1's open recipe (arXiv 2501.12948): RL with verifiable rewards elicited long chain-of-thought without a learned reward model; roughly 600k rejection-sampled traces fed a further SFT plus RL pass; roughly 800k R1-generated samples trained the 1.5B–70B distills. DAPO (arXiv 2503.14476), Dr. GRPO (arXiv 2503.20783) and GSPO (arXiv 2507.18071) are the stability fixes that came after.

The optimizer at the center is GRPO — group relative policy optimization, from DeepSeekMath (arXiv 2402.03300). GRPO removes the value critic that PPO carries: for each problem it samples a group of responses, scores each with the verifier, and mean-centers those rewards within the group to estimate the advantage. When the reward is a cheap rule, a learned critic adds cost without adding signal, which is why the critic-free variant won for RLVR. By 2026 GRPO is the central reference point of reasoning RL, and the tooling is ordinary: Hugging Face TRL's GRPOTrainer, veRL and OpenRLHF all implement it.

Vanilla GRPO has failure modes, and the field named them. DAPO (arXiv 2503.14476) fixed unstable training with dynamic sampling and decoupled clipping, reaching an AIME 2024 score of 50 with Qwen2.5-32B in a fully open RL system. Dr. GRPO (arXiv 2503.20783) showed that GRPO's advantage normalization penalized long traces unfairly, and corrected the length bias. GSPO (arXiv 2507.18071) moved the importance ratio from tokens to whole sequences and made RL on mixture-of-experts models markedly more stable. The 2026 crop — DHPO, EP-GRPO, TR-GRPO, DPPO, VPO and others mapped by Turing Post in September 2026 — is optimizing stability, credit assignment and cost. Most of those are preprints; treat them as a map of where the field is pushing, not settled consensus.

Data curation beats data volume, twice over. R1's own pipeline used rejection sampling to generate roughly 600k correct reasoning samples and then trained on them again — the loop in the figure. And the s1 result (arXiv 2501.19393) showed that 1,000 carefully filtered questions plus budget forcing — appending “wait” to extend deliberation — matches strong reasoning models at small scale. The lesson a platform should take is not that RL is magic; it is that every step is a measurable data decision.

What RL adds — and what it doesn't

The strongest version of the RLVR story comes from R1 itself. R1-Zero — pure RL from the base model, no SFT warm-start — produced self-verification, reflection and length generalization that the authors described as emergent. If taken at face value, RL is not just sharpening a distribution, it is inventing behaviors the base did not have. That claim deserves a hard look, and the field gave it one.

The best-controlled counter-result is Yue et al. In Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? (arXiv 2504.13837, NeurIPS 2025 oral), they compared RLVR-trained models against their bases at pass@1 and at pass@k — the probability that at least one of k sampled solutions is correct. The RLVR models won clearly at pass@1. But as k grew, the base models caught up and at large k matched or exceeded them. Read literally: RLVR reweights the base model's existing sampling distribution toward its solvable problems. It sharpens; the evidence that it expands the frontier of what the model can solve at all is, in that study, not there. Distillation, by contrast, could add capability the base lacked.

PASS@1 UP, PASS@K FLATThe shape of Yue et al.'s result: the RLVR-tuned model sits far above its base at pass@1, and the base catches up or passes at large k.
Schematic of the result's shape in Yue et al. (arXiv 2504.13837), not measured data. The curves are illustrative; the paper's numbers, not this figure, are the claim.

RLVR is best understood as a reliable way to harvest a base model's existing reasoning into higher pass@1 — distillation and better bases move the frontier.

The debate is live, and honest coverage says so. A counter-paper, RLVR Implicitly Incentivizes Correct Reasoning in Base LLMs (arXiv 2506.14245), argues the other side, and a 2026 reward-design survey (arXiv 2602.09305) devotes a section to “genuine reasoning improvement or optimized selection?” and reports mixed evidence on both. What you should not say, on the current evidence, is that RLVR creates genuinely novel reasoning capability. What you can say is cheaper and still useful: RLVR reliably converts a base model's existing reasoning into higher pass@1, at known cost, and it is the distillation step and the base-model choice that move the frontier. Both things are true at once, and both are operationally important: if sharpening is what RL buys, then base selection and recipe measurement are the buyer's real problem, because the base is doing more of the work than the marketing suggests.

The control is cheap, so run it before you believe any tuning result. Sample k solutions per problem from the base model and from the tuned candidate, and compare pass@1 against pass@k. If the tuned model wins pass@1 while pass@k does not move against the base, you bought sharpening, not new capability — genuinely useful, but price it as what it is. A tuning claim without that control, from a vendor or from your own pipeline, is a claim about sharpening dressed up as a claim about capability.

The open-weight landscape

For an enterprise evaluator the landscape splits in two: frontier systems sold as APIs, and open-weight families that can be run and audited inside your own boundary. The API tier's numbers are vendor-reported from top to bottom and change with every release; the open-weight tier is the one an evaluation can hold to a fixed standard, rerun next quarter, and compare against a claim. Sorted by the role a family plays in an enterprise evaluation, rather than by benchmark rank — ranks mislead across suites, as the caveats below make clear — the open-weight landscape as of October 2026 is this:

FamilyWhat to shortlist it for
DeepSeek R1, R1-0528, R1-Distill 1.5B–70BThe reference family: the open recipe everything else derives from, and the baseline to benchmark vendor claims against. R1-0528-Qwen3-8B remains a default small-reasoner baseline in 2026 papers. DeepSeek-R2 is unreleased as of 2026-08-25; the flagship 2026 DeepSeek release is the V4 family.
Qwen3 thinking, 235B-A22B down to 4BSmall reasoners and hybrid thinking/non-thinking modes for high-volume internal workloads; Qwen3-4B-Thinking-2507 is a standard fast baseline in 2026 work.
GLM-4.6/4.7 and GLM-5.x (Z.ai)The flash tier (GLM-4.7-Flash 30B-A3B) for latency-sensitive assistants, and the open agentic/coding line at GLM-5.x, MIT-licensed at 5.2.
Kimi K2 Thinking, then K3 (Moonshot)The open-weight agentic reasoner for tool-use evaluation: vendor-stated 1T-parameter MoE, 32B active, 200–300 sequential tool calls held coherently (Moonshot's own post). K3 is reporting-grade and thinking-by-default.
gpt-oss-20b / 120b (OpenAI-open)Open reasoners in two MXFP4 sizes for private deployment; the 120B is the strongest open option in this table.
MiniMax M1, then M2.5/M2.7The 1M-context reasoning line for long-document and long-session evaluation. M1 shipped hybrid linear attention; the M2.x releases reportedly dropped it.

Read the table with the caveats attached. Several rows rest on reporting-grade sources, and we say so rather than launder them: Kimi K3's specifications (reported 2.8T dense, 1M context) and the MiniMax attention reversal come from aggregator and survey reporting, not primary cards; K2 Thinking's own numbers above are vendor-stated in Moonshot's launch post; the Qwen3.6–3.8 release details are community-reported, with Qwen3.8's AIME 2026 FP8 claims explicitly unconfirmed by the vendor in our pass. R2 status gets the plainest sentence in this section: DeepSeek-R2 is not released as of 2026-08-25 (decodethefuture status roundup, 2026-08-25, reporting; corroborated by Layer3 Labs, 2026-07-21). Plan around what is released.

Benchmark numbers in this tier need discipline. AIME 2024 and 2025 are effectively saturated at the frontier, which is why AIME 2026 (February 2026) is the current suite; SWE-bench Verified is clustered near its ceiling and labs have moved to SWE-bench Pro, whose integrity is itself contested — OpenAI reportedly pulled its Pro claims in 2026 after finding roughly 30% of tasks broken (reporting). And across ten agent benchmarks, an automated auditor found 219 flaws and reached near-perfect scores “without solving a single task” (BenchJack, arXiv 2605.12673). The house rule, which our own benchmark posts follow: quote a number with its suite and date, label vendor-reported numbers as vendor-reported, and never compare across suites.

What reasoning costs to serve

Here is the part that changes your infrastructure. A standard LLM workload is heavy prefill and light decode: a long prompt in, a short answer out. A reasoning model inverts it: a short problem in, ten thousand tokens of decode out, with wild variance in length between two problems that look identical. The one measured study of this workload as of October 2026 is the ICLR 2026 empirical study, Reasoning Language Model Inference Serving Unveiled (OpenReview 6CGjZYp6ft), run across eight paired LLM/RLLM configurations on vLLM and SGLang. Its findings are the spine of this section.

WHAT LONG CHAIN-OF-THOUGHT COSTSKV cache occupancy swings of 3–70% for reasoning models against under 3% for standard LLMs, the measured direction of each popular optimization, and the overthinking result.
Measured on eight paired LLM/RLLM serving configurations (vLLM/SGLang): ICLR 2026 empirical study, OpenReview 6CGjZYp6ft, summarized by papernotes 2026-05-08. DeepSeek's published pricing encodes the prefix-cache effect: $0.14 per million input tokens on a cache hit versus $0.55 on a miss (open-infra-index inference overview, 2025-02-27).

The KV cache becomes the binding resource. In the study's measurements, KV cache occupancy swings 3–70% for reasoning models — nearing 100% under bursty load — against under 3% for standard LLMs. A few hard problems become long-tail stragglers that dominate batch latency, and runtime correlates with problem difficulty. Capacity planning on parameter count alone will mislead you; what matters is p95 trace length times batch size times KV per token.

TechniqueWhat the measurements say
Prefix cachingA clear win for reasoning models at ≥14B, where shared system prompts and repeated problem prefixes dominate; a net latency loss below 8B, where hash overhead dominates. DeepSeek's pricing encodes the same economics: $0.14/M input tokens on a cache hit versus $0.55/M on a miss.
Speculative decodingCuts end-to-end latency for reasoning models at all scales with no accuracy loss, but reduces throughput and worsens first-token latency. A trade-off, not a free lunch — under high concurrency the throughput loss is the number you feel.
KV cache compressionA correctness caveat, not just a speedup: under compression, final-answer accuracy can hold while the evidence chain silently breaks, an answer–evidence gap (arXiv 2608.01631). For reasoning models, report trace faithfulness, not just final accuracy, when evaluating memory optimizations.
Longer token budgets4096–8192 token budgets suffice on most of the study's suites, and longer budgets hurt accuracy on GPQA and AIME — the overthinking result. Length is not a free dial.

The economics are now first-order. OpenAI's own September 2026 account of its research organization reports the median researcher using more than $600 per day of inference, with agents contributing 3.1 agent-workdays per human workday (OpenAI, 2026-09-06, vendor-reported). You will also see a claim circulating that inference now accounts for 70–80% of total AI compute. We could not verify it from a primary source, so this post does not repeat it. The structural point survives either way: the bill scales with tokens spent thinking, not with answers delivered — which is why the capacity math above belongs in the procurement review as well as the architecture review.

What this means for an enterprise platform team

Evaluation moves ahead of procurement. The measurement discipline in this post is a pilot design. Put the candidate against your own problem distribution — the tasks you are actually buying for, not the vendor's demo set — run the pass@k control from section three, and write the result into the pilot's acceptance criteria before anyone signs. A benchmark number, from a vendor or from your own tuning run, without that control is a claim about sharpening sold as a claim about capability.

Serving economics become a capacity problem you own either way. Whether the model is served inside your own boundary or bought as an API, the quantities to plan around are not parameter counts. They are the trace-length distribution of your workload, the KV footprint at your concurrency, long-tail stragglers against your latency targets, and cost per solved task rather than cost per token. API pricing encodes the same dials — DeepSeek's published cache-hit and cache-miss input prices are the clearest published example — so even a pure buy decision deserves a measurement loop on top of it, and a renewal clause that survives a price change.

Decide what to own. Owning the training loop — verifiers, rollout infrastructure, the stability fixes of section two — is a platform program, and it pays only if the recipe itself is a differentiator. Owning the evaluation harness and the outcome record pays in every scenario: it is what makes a vendor's claim checkable next quarter and your own claims defensible in an audit. CID is STEAV's platform for owning that loop, and the next section shows what it looks like in practice — one platform, every move above made concrete. The discipline itself is stack-independent, and it starts with a held-out problem set, not a purchase.

Where CID comes in

CID is STEAV's platform for owning the model layer, and it is the worked example of the three moves above. Every recipe this post has described — SFT, rejection sampling, GRPO-family RL with verifiable rewards, distillation — runs as an auditable experiment through the same seven stages: data_prep → train → aggregate → eval → benchmark → validate → deploy. Nothing is skipped because a search launched it. The validate stage plus the platform's outcome ledger is pre-registration by construction: the criteria a recipe must meet are recorded before its results are, and every recipe change is an experiment with a recorded outcome — the audit trail a procurement review, or an internal one, asks for.

Model Studio runs the search and carries the advisor. Measured paired search queues trials for idle compute and selects by validation metrics, not by a leaderboard glance, and a recipe advisor — trained from the platform's own outcome ledger — predicts, given two configurations, which will reach the better outcome. The prediction is logged and timestamped before evaluation runs. On our own deployment that loop has promoted an advisor over its incumbent on holdout evidence, audited and reversible, with no human touching the decision (the record, internal and verifiable). What the advisor's accuracy is on reasoning recipes specifically is a number we do not have: the paired bake-off that would have scored its predictions against measured pass@1 was cancelled before it ran. When we run it, the protocol is the one this post has argued for throughout — predictions fixed before results, scored against measurement, reported either way.

The tooling and the serving layer are one system. Training is TRL-based — SFT, DPO and GRPO-family trainers — with rollouts served through the same engine the benchmark stage times, so the recipe loop and the serving layer measure one artifact rather than two versions of it. That engine is our own llama.cpp fork: one engine across NVIDIA and AMD unified-memory hardware, serving EXL3-quantized long-chain-of-thought models, after a ground-up correctness effort that we consider the price of entry for heterogeneous serving. System One, CID's calibrated-decision-model layer — a shipped recipe, not a roadmap item — is the piece that bears on the trace-trust question: which reasoning traces to trust, and with what confidence.

None of this is required for the measurements of the next section; those are stack-independent. CID is simply what the loop looks like when it is owned end to end on one platform — every configuration, every outcome, and every promotion earned or rejected on recorded evidence.

What to measure before you adopt a recipe

None of the above is an argument against reasoning models; it is an argument for measuring them on your own base before you spend. The recipe is public, the verifiers are cheap, and four measurements separate a recipe that pays for itself from one that only demos:

  1. pass@1 against pass@k, against a fixed base. This is the Yue et al. control, and it is cheap: sample k solutions per problem from the base and from the tuned model. pass@1 tells you what you bought; pass@k tells you what the base could already do. The gap between those two sentences is the honest price of the recipe.
  2. Trace length, p50 and p95, on your workload. Not the context-window rating in the datasheet — your distribution, on your problems. It decides the memory bill, the straggler tail, and whether a batch scheduler survives contact with production.
  3. Serving optimizations ablated in both directions. Prefix caching and speculative decoding, on and off, at the scale you plan to deploy, on the serving stack you plan to run. The measured evidence says the advice inverts across scales; measure which side your deployment lands on before committing capacity to it.
  4. Trace faithfulness, not just final accuracy. Especially under KV cache compression or any other memory optimization: final answers can hold while the evidence chain silently breaks, and a reasoning model you cannot interrogate is a liability wearing a benchmark score.

Reasoning is a recipe and a bill. Both can be measured before you pay either — and a platform that records every run, with its configuration and its measured outcome, is what makes those measurements compound instead of evaporate.

FAQ

  • Is RLVR the same as RLHF? No. RLHF trains against a learned reward model fitted to human preference labels. RLVR replaces that learned reward with a rule verifier — exact answer match, unit tests — that anyone can audit. The distinction is the point: you can read a rule, and you cannot read a reward model's mind. PPO is the ancestor of both; GRPO is the critic-free variant that won for verifiable rewards.
  • Do I need a reward model to train a reasoning model? Not for math and code with exact-match or unit-test verifiers — that is precisely what RLVR removes, and why the recipe is cheap to replicate. You need a verifier, which is a different thing: a function that scores correctness, not a model that predicts preference. Open-ended or agentic tasks without checkable answers are the harder case, and an active research direction.
  • Does RL make a model reason novelly? Contested, and we decline to settle it. Yue et al. (arXiv 2504.13837) found pass@1 up and pass@k flat against the base; counter-papers argue genuine improvement; a 2026 survey reports mixed evidence. The defensible claim is narrower: RL reliably harvests a base model's existing reasoning into higher pass@1. The control is cheap enough to run in the pilot, before procurement — sample k solutions per problem from the base and from the tuned model, and compare pass@k — because a tuning result without a control is a number about the prior, not about the recipe.
  • Do we need RL infrastructure, or is distillation enough? Probably start with distillation. Distilling from a strong reasoning teacher buys much of the pass@1 gain at a fraction of the cost — that is the pattern behind R1-Distill — while owning GRPO means standing up rollout serving, verifier engineering and the stability fixes of section two, which is a platform program, not a notebook. Whether RL beats matched-data distillation on your workload is an empirical question the field has not settled, so run the pass@k control on whatever you deploy. And if a vendor sells you a tuning stage, ask for the control.
  • Can we bring our own base model and reward verifiers? Yes. A base model registers from your own storage, and verifiers are ordinary functions — exact answer match, unit tests — that you can read and version like code, because that is what they are. Your data, checkpoints and outcome records stay inside your boundary; the platform's job is to make every run through it auditable, not to make your models hostage to it.
  • Why did my speculative-decoding speedup disappear? Because the speedup was always a trade-off. The ICLR 2026 measurements found speculative decoding cuts end-to-end latency for reasoning models at all scales with no accuracy loss, but it reduces throughput and worsens first-token latency. On a lightly loaded deployment you see the latency win; under high concurrency, or when first-token time is what your users feel, the loss side dominates and the optimization inverts. Measure both directions before keeping it.
All news

Bring your own base. The first experiment is the control.

Book a live demoExplore CID