Test-Time Self Improvement by Scaling Unsupervised World Internalization

Violet Xiang, M. Khalifa, I. Wu, Y. Shen, J. Leskovec, N. Haber, A. Kumar.

Under review at ICLR 2027

U-WIN enables agents to explore new environments and train on self-written notes and self-generated trajectories without access to downstream tasks, rewards, or correctness feedback. Across legal, codebase, and math environments, downstream performance improved near log-linearly as training scaled from 10M to 100M tokens.

Experiments used Qwen 9B and 27B models across Harvey, StudyBench, and Equational Theories. U-WIN increased the share of agentic trajectories that recall environment knowledge from 20.6% to 99.8%. Self-written notes also outperformed token-matched raw documents or observations for knowledge retention.

← All publications