Reasoning Models and Test-Time Compute

Parent: LLM Models and APIs · researched 2026-05-31T21:02:22.465Z· 17 sources · 14 concepts · skill reasoning-models

The frontier (2024–2026) shift from "scale the model and prompt it well" to "train the model to reason, then spend extra compute at inference to reason harder." Two coupled ideas drive it:

Reasoning Models & Test-Time Compute

Scope boundary (read first)

1. Chain-of-thought (CoT) and long-CoT

2. Reinforcement learning for reasoning — GRPO and the DeepSeek-R1 recipe

GRPO (Group Relative Policy Optimization)

The R1-Zero and R1 recipe (arXiv:2501.12948)

3. RLVR — Reinforcement Learning with Verifiable Rewards

The RLVR effectiveness debate (important, unresolved as of mid-2026)

4. Process reward models (PRM) vs outcome reward models (ORM)

5. Parallel test-time compute — best-of-N, self-consistency, verifiers

6. Search-based test-time compute — beam, lookahead, MCTS, reward-guided decoding

7. Inference-time scaling laws and compute-optimal test-time scaling

8. The reasoning-model landscape — the o-series, R1-class, and mid-2026 SOTA

9. Budget forcing and thinking-token control

10. Distilling reasoning into smaller models

11. Reasoning benchmarks — and why they keep breaking

12. Cost, latency, and accuracy trade-offs (the operating decision)

Practical patterns

Anti-patterns

Troubleshooting

References (primary sources)

Children

Frontier under this node: Budget forcing and thinking-token control, Chain-of-thought and long-CoT, Cost / latency / accuracy trade-offs and overthinking, Inference-time scaling laws and compute-optimal test-time scaling, Parallel test-time compute (self-consistency / majority vote, best-of-N, generative verifiers), Process reward models (PRM) vs outcome reward models (ORM) and process supervision, RLVR — reinforcement learning with verifiable rewards, Reasoning benchmarks (AIME, GPQA-Diamond, MATH, LiveCodeBench), Reasoning distillation (R1-Distill), Reinforcement learning for reasoning (GRPO + DeepSeek-R1/R1-Zero recipe), Search-based test-time compute (beam, lookahead, MCTS, reward-guided decoding), The o-series / R1-class reasoning-model landscape

← the whole tree · 3D view· how to read this page