Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought

Violet Xiang, C. Snell, K. Gandhi, A. Albalak, A. Singh, C. Blagden, D. Phung, R. Rafailov, N. Lile, D. Mahan, L. Castricato, J.-P. Fränken, N. Haber, C. Finn.

Preprint, 2025

A training roadmap combining linearized search traces, process supervision, synthetic data, and reinforcement learning, supported by evidence of in-context search behavior in frontier reasoning models.

← All publications