Test-Time Self Improvement by Scaling Unsupervised World Internalization
Violet Xiang, M. Khalifa, I. Wu, Y. Shen, J. Leskovec, N. Haber, A. Kumar.
Under review at ICLR 2027
U-WIN enables agents to explore new environments and train on self-written notes and self-generated trajectories without access to downstream tasks, rewards, or correctness feedback. Across legal, codebase, and math environments, downstream performance improved near log-linearly as training scaled from 10M to 100M tokens.
Experiments used Qwen 9B and 27B models across Harvey, StudyBench, and Equational Theories. U-WIN increased the share of agentic trajectories that recall environment knowledge from 20.6% to 99.8%. Self-written notes also outperformed token-matched raw documents or observations for knowledge retention.
← All publications