Ph.D. student in Machine Learning (CSE)

Yizhou Zhao

Georgia Institute of Technology

I am a Ph.D. student in Machine Learning (CSE) at Georgia Tech, advised by Prof. Bo Dai. My research focuses on post-training for language models and self-evolving agentic systems.

I am interested in how language models learn to reason, take actions, and improve from feedback. My current interests include reinforcement learning for reasoning, on-policy distillation (OPD), and agentic RL, as well as systems that improve their own strategies through interaction and evaluation. My work also covers AI safety, including watermarking for language models and synthetic data.

Previously, I received my master's degree in Applied Mathematics and Computational Science (AMCS) from the University of Pennsylvania, advised by Prof. Weijie Su and Prof. Qi Long. I completed my bachelor's degree in Information and Computing Science at Zhejiang University.

Yizhou Zhao

News

Two papers have been accepted to NeurIPS 2026 (CurveRL and MarkTune)! See you in Atlanta!

I started my Ph.D. in Machine Learning (CSE) at Georgia Tech, advised by Prof. Bo Dai.

I completed my master's degree in AMCS at the University of Pennsylvania.

Our work on robust spectral watermarking for synthetic tabular data appeared in Statistical Learning and Data Science.

Research interests

I am interested in the learning methods and feedback mechanisms that help language models become more capable reasoners and agents.

Post-training & reasoning

Reinforcement learning for reasoning, on-policy distillation, and learning from verifiable feedback. I am particularly interested in how training objectives, credit assignment, and the distribution of training problems shape what a model learns and how well it generalizes.

Agentic reinforcement learning

Learning policies for multi-step tasks that involve decisions, tools, and interaction with an environment. I am interested in how agents use feedback across a trajectory, learn from their own experience, and develop strategies that transfer to new tasks.

Self-evolving agentic systems

Agents that improve their own problem-solving processes over time. I am interested in the loop between generating experience, evaluating outcomes, and revising strategies, and in how to make these improvements reliable rather than specific to a narrow evaluation.

Selected publications

Full list on Google Scholar

* Equal contribution

2025

Robust Spectral Watermark for Synthetic Tabular Data

Yizhou Zhao, Xiang Li, Peter X. K. Song, Qi Long, Weijie Su

Statistical Learning and Data Science · Synthetic data · Watermarking

A post-editing watermark for synthetic tabular data that embeds signals in the frequency domain. The method handles heterogeneous features and uses rank-based pseudorandom bits for robustness to common data transformations.

2025

Toward Engineering AGI: Benchmarking the Engineering Design Capabilities of LLMs

Xingang Guo, Yaxin Li, Xiangyi Kong, …, Yizhou Zhao, …, Bin Hu

NeurIPS 2025 · Datasets & Benchmarks (Poster) · LLM evaluation · Engineering design

EngDesign evaluates language models on practical design tasks across nine engineering domains. Simulation-based evaluation tests whether generated designs meet functional goals and constraints, beyond factual recall or textbook questions.

Education

— Present
Georgia Institute of Technology seal

Ph.D. in Machine Learning (CSE)

Georgia Institute of Technology

Advised by Prof. Bo Dai

—
University of Pennsylvania shield

Master's in Applied Mathematics and Computational Science

University of Pennsylvania

Advised by Prof. Weijie Su and Prof. Qi Long

—
Zhejiang University emblem

Bachelor's in Information and Computing Science

Zhejiang University

Academic Service

Conference reviewer

  • NeurIPS 2026
  • ICLR 2027

Journal reviewer

  • Journal of the American Statistical Association (JASA)
  • The Annals of Applied Statistics (AOAS)
  • Harvard Data Science Review