I am a Ph.D. student at The University of Hong Kong, advised by Prof. Jiayu Chen. My research focuses on reinforcement learning, LLM alignment, preference optimization, and agentic AI systems.

I study the theory and systems behind self-improving language models: online alignment and RLHF, robust preference learning, self-play post-training, and long-horizon agents that reason with tools, memory, and feedback. I am particularly interested in turning principled learning algorithms into reliable, scalable AI systems.

Before HKU, I was selected through a competitive process for a dual-degree programme at the University of Edinburgh and Dalian University of Technology. I graduated from Edinburgh with First-Class Honours and ranked in the top 5% at Dalian University of Technology.

News

Joining Microsoft in Beijing as a Research Intern, working on Agent RSI and Auto Research.

Started leading EnergyBridge, a long-horizon LLM-agent system for virtual power plant coordination.

On the Convergence of Self-Improving Online LLM Alignment accepted to UAI 2026.

Occupancy Reward Shaping accepted to ICLR 2026.

Started my Ph.D. at The University of Hong Kong.

Experience

Research Intern

Microsoft · Beijing, China

Research on Agent RSI and Auto Research.

Project Lead, Agentic Intelligence Lab

The University of Hong Kong · Hong Kong SAR

Independently authored multiple research papers and led the development of several research projects.

Research Assistant

University of California, Irvine · Irvine, CA

Worked on LLM-assisted workflow analysis and reporting for the Texera data analytics platform.

Selected Publications

All research →
  1. On the Convergence of Self-Improving Online LLM Alignment

    Xudong Wu, Pangpang Liu, Vaneet Aggarwal, and Jiayu Chen.

    UAI 2026
  2. Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning

    Aravind Venugopal, Jiayu Chen, Xudong Wu, Chongyi Zheng, Benjamin Eysenbach, and Jeff Schneider.

    ICLR 2026
  3. Distributionally Robust Listwise Preference Optimization

    Xudong Wu, Jian Qian, Pangpang Liu, Vaneet Aggarwal, and Jiayu Chen.

    Under review

Research Interests

  • LLM post-training and alignment: RLHF/RLAIF, preference optimization, online alignment, and verifier-guided learning.
  • Agentic AI: self-improving agents, tool use, persistent memory, multi-agent systems, and long-horizon planning.
  • Reinforcement learning: offline and goal-conditioned RL, self-play, game-theoretic learning, and trust-region methods.

Education

The University of Hong Kong

Ph.D. · Reinforcement Learning and LLM Alignment

University of Edinburgh

BSc (Hons) Mathematics and Statistics · Competitively selected dual-degree programme · First-Class Honours

Dalian University of Technology

BSc Information and Computing Science · Top 5%