Reasoning Models: Papers to Read

A bottom-up reading list for understanding reasoning in LLMs: from chain-of-thought prompting through process reward models, RL-trained reasoners (o1 / R1 style), test-time compute scaling, and self-improvement. I'll keep adding to this over time rather than starting a new post for every batch.

Type key: paper = peer-reviewed / arxiv preprint · blog = blog/post · report = technical report · book = textbook/chapter · survey = survey paper · website = tool/org site · dataset = data/benchmark release


Part 1 - Chain-of-Thought & Prompting Foundations


Part 2 - Structured & Search-Based Reasoning


Part 3 - Tool Use & Program-Aided Reasoning


Part 4 - Self-Improvement & Bootstrapping


Part 5 - Verification & Process Reward Models


Part 6 - RL for Reasoning (o1 / R1 Era)


Part 7 - Test-Time Compute Scaling


Part 8 - Distillation & Small Reasoners


Part 9 - Faithfulness & Interpretability of Reasoning


Part 10 - Math & Theorem Proving


Part 11 - Benchmarks for Reasoning


Part 12 - Surveys & Big-Picture Reading

← Back to all posts