I am a Ph.D. student at The University of Hong Kong, advised by Prof. Jiayu Chen. My research focuses on reinforcement learning, LLM alignment, preference optimization, and agentic AI systems.
I study the theory and systems behind self-improving language models: online alignment and RLHF, robust preference learning, self-play post-training, and long-horizon agents that reason with tools, memory, and feedback. I am particularly interested in turning principled learning algorithms into reliable, scalable AI systems.
Before HKU, I earned a BSc (Hons) in Mathematics and Statistics from the University of Edinburgh with First-Class Honours, and studied Information and Computing Science at Dalian University of Technology.
News
Joining Microsoft in Beijing as a Research Intern, working on Agent RSI and Auto Research.
Started leading EnergyBridge, a long-horizon LLM-agent system for virtual power plant coordination.
On the Convergence of Self-Improving Online LLM Alignment accepted to UAI 2026.
Occupancy Reward Shaping accepted to ICLR 2026.
Started my Ph.D. at The University of Hong Kong.
Experience
Research Intern
Microsoft · Beijing, China
Research on Agent RSI and Auto Research.
Project Lead, Agentic Intelligence Lab
The University of Hong Kong · Hong Kong SAR
Leading EnergyBridge, an end-to-end long-horizon LLM-agent system for home-grid and virtual power plant coordination.
Research Assistant
University of California, Irvine · Irvine, CA
Worked on LLM-assisted workflow analysis and reporting for the Texera data analytics platform.
Selected Publications
All research →-
On the Convergence of Self-Improving Online LLM Alignment
Xudong Wu, Pangpang Liu, Vaneet Aggarwal, and Jiayu Chen.
UAI 2026 -
Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning
Aravind Venugopal, Jiayu Chen, Xudong Wu, Chongyi Zheng, Benjamin Eysenbach, and Jeff Schneider.
ICLR 2026 -
Distributionally Robust Listwise Preference Optimization
Xudong Wu, Jian Qian, Pangpang Liu, Vaneet Aggarwal, and Jiayu Chen.
Under review
Research Interests
- LLM post-training and alignment: RLHF/RLAIF, preference optimization, online alignment, and verifier-guided learning.
- Agentic AI: self-improving agents, tool use, persistent memory, multi-agent systems, and long-horizon planning.
- Reinforcement learning: offline and goal-conditioned RL, self-play, game-theoretic learning, and trust-region methods.
Education
The University of Hong Kong
Ph.D. · Reinforcement Learning and LLM Alignment
University of Edinburgh
BSc (Hons) Mathematics and Statistics · First-Class Honours
Dalian University of Technology
BSc Information and Computing Science