About Me

I am a PhD candidate in Psychology at Stanford University, working in the Stanford Autonomous Agents Lab with Nick Haber. My research focuses on LLM post-training, reinforcement learning, and reward modeling, with roots in computational cognitive science.

I study how to train language models to reason more effectively, how to provide useful feedback during learning, and how to evaluate what models learn. My recent work spans dense rewards for exploratory RL, adaptive reasoning budgets, reward-model biases, and human-preference evaluation. I also work on theory of mind in multi-agent systems and benchmarks connecting human and machine learning.

Previously, I earned an M.S. in Computer Science and undergraduate degrees in Informatics and Mathematics at Indiana University.

CV (PDF) · Google Scholar · GitHub · Email

Selected research

One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models

ICML 2026

Identified persistent length, uncertainty, position, sycophancy, and model-style biases in language reward models. Null-space probe projections reduced three of these biases without retraining or loss on RewardBench-2.

View all publications →

Industry research

Radical Numerics

Member of Technical Staff Intern

October 2025 - January 2026

Built an end-to-end post-training and evaluation stack for DPO-tuning genomic foundation models, including preference-data processing, training orchestration, and evaluation for biomedical DNA-sequence design.

Toyota Research Institute

Research Scientist Intern, Human-AI Interactive Learning

June - September 2025

Developed a teacher-training objective for corrective feedback, scoring teacher outputs by how much they improved a frozen student's likelihood of the reference solution. Established a final-answer reward signal and studied why it did not transfer to long reasoning traces.