I am a Ph.D. student at The University of Hong Kong, advised by Prof. Jiayu Chen. My research focuses on reinforcement learning, LLM alignment, preference optimization, and agentic AI systems.

I study the theory and systems behind self-improving language models: online alignment and RLHF, robust preference learning, self-play post-training, and long-horizon agents that reason with tools, memory, and feedback. I am particularly interested in turning principled learning algorithms into reliable, scalable AI systems.

Before HKU, I earned a BSc (Hons) in Mathematics and Statistics from the University of Edinburgh with First-Class Honours, and studied Information and Computing Science at Dalian University of Technology.

News

Joining Microsoft in Beijing as a Research Intern, working on Agent RSI and Auto Research.

Started leading EnergyBridge, a long-horizon LLM-agent system for virtual power plant coordination.

On the Convergence of Self-Improving Online LLM Alignment accepted to UAI 2026.

Occupancy Reward Shaping accepted to ICLR 2026.

Started my Ph.D. at The University of Hong Kong.

Experience

Research Intern

Microsoft · Beijing, China

Research on Agent RSI and Auto Research.

Project Lead, Agentic Intelligence Lab

The University of Hong Kong · Hong Kong SAR

Leading EnergyBridge, an end-to-end long-horizon LLM-agent system for home-grid and virtual power plant coordination.

Research Assistant

University of California, Irvine · Irvine, CA

Worked on LLM-assisted workflow analysis and reporting for the Texera data analytics platform.

Selected Publications

All research →
  1. On the Convergence of Self-Improving Online LLM Alignment

    Xudong Wu, Pangpang Liu, Vaneet Aggarwal, and Jiayu Chen.

    UAI 2026
  2. Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning

    Aravind Venugopal, Jiayu Chen, Xudong Wu, Chongyi Zheng, Benjamin Eysenbach, and Jeff Schneider.

    ICLR 2026
  3. Distributionally Robust Listwise Preference Optimization

    Xudong Wu, Jian Qian, Pangpang Liu, Vaneet Aggarwal, and Jiayu Chen.

    Under review

Research Interests

  • LLM post-training and alignment: RLHF/RLAIF, preference optimization, online alignment, and verifier-guided learning.
  • Agentic AI: self-improving agents, tool use, persistent memory, multi-agent systems, and long-horizon planning.
  • Reinforcement learning: offline and goal-conditioned RL, self-play, game-theoretic learning, and trust-region methods.

Education

The University of Hong Kong

Ph.D. · Reinforcement Learning and LLM Alignment

University of Edinburgh

BSc (Hons) Mathematics and Statistics · First-Class Honours

Dalian University of Technology

BSc Information and Computing Science