I am a Ph.D. student in Machine Learning (CSE) at Georgia Tech, advised by Prof. Bo Dai. My research focuses on post-training for language models and self-evolving agentic systems.
I am interested in how language models learn to reason, take actions, and improve from feedback. My current interests include reinforcement learning for reasoning, on-policy distillation (OPD), and agentic RL, as well as systems that improve their own strategies through interaction and evaluation. My work also covers AI safety, including watermarking for language models and synthetic data.
I am interested in the learning methods and feedback mechanisms that help language models become more capable reasoners and agents.
Post-training & reasoning
Reinforcement learning for reasoning, on-policy distillation, and learning from verifiable feedback. I am particularly interested in how training objectives, credit assignment, and the distribution of training problems shape what a model learns and how well it generalizes.
Agentic reinforcement learning
Learning policies for multi-step tasks that involve decisions, tools, and interaction with an environment. I am interested in how agents use feedback across a trajectory, learn from their own experience, and develop strategies that transfer to new tasks.
Self-evolving agentic systems
Agents that improve their own problem-solving processes over time. I am interested in the loop between generating experience, evaluating outcomes, and revising strategies, and in how to make these improvements reliable rather than specific to a narrow evaluation.
A framework for prompt reweighting in reinforcement learning with verifiable rewards. CurveRL uses the rank and density of prompt pass rates to account for their distribution during training.
NeurIPS 2026 (Poster) · Language models · Watermarking
An on-policy fine-tuning framework for watermarking open-weight language models. It uses a watermark signal as a reward while regularizing text quality to improve the quality–detectability trade-off.
Yizhou Zhao, Xiang Li, Peter X. K. Song, Qi Long, Weijie Su
Statistical Learning and Data Science · Synthetic data · Watermarking
A post-editing watermark for synthetic tabular data that embeds signals in the frequency domain. The method handles heterogeneous features and uses rank-based pseudorandom bits for robustness to common data transformations.
EngDesign evaluates language models on practical design tasks across nine engineering domains. Simulation-based evaluation tests whether generated designs meet functional goals and constraints, beyond factual recall or textbook questions.