About Me
I am a PhD candidate in Psychology at Stanford University and a member of the Stanford Autonomous Agents Lab. I work with Nick Haber and Aviral Kumar.
My research interests span exploration and scientific discovery through reinforcement learning, test-time learning, and learning in open-ended domains such as creative writing. I bring a computational cognitive science perspective to questions about how agents learn, reason, and explore.
Background
Previously, I earned an M.S. in Computer Science and undergraduate degrees in Informatics and Mathematics at Indiana University.
CV (PDF) · Google Scholar · GitHub · Email
Selected research
Test-Time Self Improvement by Scaling Unsupervised World Internalization
Under review at ICLR 2027
U-WIN enables agents to explore new environments and train on self-written notes and self-generated trajectories without access to downstream tasks, rewards, or correctness feedback. Across legal, codebase, and math environments, downstream performance improved near log-linearly as training scaled from 10M to 100M tokens.
ExpRL: Exploratory RL for LLM Mid-Training
COLM 2026
RL-based mid-training with dense outcome- and process-level rewards from a reference-guided LLM judge. On held-out AIME 2026 with Qwen3-4B, ExpRL improved pass@1 over the strongest baseline at each stage by 3.6 points after priming and 4.7 points after identical sparse-GRPO post-training.
One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models
ICML 2026
Identified persistent length, uncertainty, position, sycophancy, and model-style biases in language reward models. Null-space probe projections reduced three of these biases without retraining or loss on RewardBench-2.
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing
EACL 2026 · Long Paper
A creative-writing evaluation benchmark with 43k training pairs and a 2.5k-pair debiased, human-labeled test set. Trained reward models improved human-preference agreement by 5 points over the strongest zero-shot judge.
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
NeurIPS 2025 Efficient Reasoning Workshop · Spotlight
A GRPO-compatible adaptive length penalty reduced DeepScaleR-1.5B generation length by more than 50% at matched accuracy across MATH-500, AIME, and OlympiadBench, while allocating 5.35× more tokens to hard prompts than easy ones.
Industry research
Radical Numerics
Member of Technical Staff Intern
October 2025 - January 2026
Built an end-to-end post-training and evaluation stack for DPO-tuning genomic foundation models, including preference-data processing, training orchestration, and evaluation for biomedical DNA-sequence design.
Toyota Research Institute
Research Scientist Intern, Human-AI Interactive Learning
June - September 2025
Developed a teacher-training objective for corrective feedback, scoring teacher outputs by how much they improved a frozen student's likelihood of the reference solution. Established a final-answer reward signal and studied why it did not transfer to long reasoning traces.
