Yihang Yao yihangya[at]andrew.cmu.edu
Hi, welcome to my website! I am a final-year Ph.D. candidate in the Safe AI Lab at Carnegie Mellon University
, advised by Prof. Ding Zhao. I am also a research intern at NVIDIA Nemotron
. I received my Bachelor's degree from Shanghai Jiao Tong University in 2022 and spent a wonderful time as a visiting student in the Intelligent Control Lab at CMU, working with Prof. Changliu Liu. I previously interned at Salesforce AI Research, MIT-IBM Watson AI Lab, IBM Research and have also been collaborating with Google DeepMind Robotics since 2023.
My research focuses on reinforcement learning and large language models, with a current emphasis on self-evolving agents and RL environment scaling for real-world applications such as terminal/coding agents [Technical Report 2026
]. My prior research experience also includes:
Yihang Yao, Bo Pang, Xuan Phi Nguyen, Ding Zhao, Shafiq Joty, Semih Yavuz
Technical Report, 2026 · Paper / Website / Salesforce Blog
Collaboration with Salesforce AI Research
TL;DR: We propose the agentic compositional generalization hypothesis, which posits that terminal-agent RL generalizes by shaping reusable behaviors that compose existing atomic skills. Using RL environments with high-quality verifiers to shape these behaviors, our training recipe achieves 106% and 30% larger average RL performance gains on Terminal-Bench-Lite and Terminal-Bench-v2.1, respectively, across models ranging from 2B to 27B parameters, while using fewer than 30% of the environments in the TMax collection.
Yihang Yao, Zhepeng Cen, Haohong Lin, Shiqi Liu, Zuxin Liu, Jiacheng Zhu, Zhang-Wei Hong, Laixi Shi, Ding Zhao
ICML 2026 · Paper / Website / Code
Collaboration with Salesforce AI Research
TL;DR: How can we better balance user burden and task performance in multi-turn interactive tasks? We introduce Behavioral Agentic Optimization (BAO), an agentic RL pipeline that combines simulated-user training with turn-wise reward regularization. On UserRL tasks, BAO achieves state-of-the-art performance while requiring less user intervention than the evaluated RL baselines.
Yihang Yao, Guangtao Zeng, Raina Wu, Yang Zhang, Ding Zhao, Zhang-Wei Hong, Chuang Gan
ACL 2026 Main · Paper
Collaboration with MIT-IBM Watson AI Lab, IBM Research
TL;DR: How can we improve RL efficiency and final performance? We introduce Tailor, an automated SFT data-curation pipeline that warm-starts RL with diverse reasoning primitives, including skills and reasoning patterns. This initialization broadens exploration during RL, improves sample efficiency, and raises the final performance ceiling.
Zhepeng Cen*, Yihang Yao*, William Han, Zuxin Liu, Ding Zhao
NeurIPS 2025 · Paper / Website / Code
Collaboration with Salesforce AI Research
TL;DR: Why do LLMs respond differently to RL and exhibit varying data efficiency? We explain these differences through data co-influence effects and identify reasoning patterns as an important factor in RL efficiency. We also introduce a structured data pipeline that prepares LLMs for RL by injecting exploratory and exploitative behaviors into their reasoning traces.
Yihang Yao, Zhepeng Cen, Miao Li, William Han, Yuyou Zhang, Emerson Liu, Zuxin Liu, Chuang Gan, Ding Zhao
ACL 2025 Findings · Paper
Collaboration with Salesforce AI Research
TL;DR: MEND enhances LLM robustness by improving query symmetry awareness, boosting reasoning performance through structured dataset curation.
Yihang Yao*, Zhepeng Cen*, Wenhao Ding, Haohong Lin, Shiqi Liu, Tingnan Zhang, Wenhao Yu, Ding Zhao
NeurIPS 2024 · Paper / Website / Code
Collaboration with Google DeepMind
TL;DR: We investigate offline RL from a data-centric perspective and propose a diffusion model-based data generator to curate training datasets aligned with user preferences.
Zhepeng Cen, Yihang Yao, Zuxin Liu, Ding Zhao
ICML 2024
TL;DR: We introduce FCSRL, a framework that improves safety constraint estimation in RL through representation learning and self-supervised techniques.
Yihang Yao, Zuxin Liu, Zhepeng Cen, Peide Huang, Tingnan Zhang, Wenhao Yu, Ding Zhao
L4DC 2024 · Paper / Website / Code
Collaboration with Google DeepMind
TL;DR: We introduce GradS, a gradient-based method for improving training efficiency in multi-constraint RL by manipulating gradients, optimizing both reward and constraint satisfaction.
Yihang Yao*, Zuxin Liu*, Zhepeng Cen, Jiacheng Zhu, Wenhao Yu, Tingnan Zhang, Ding Zhao
NeurIPS 2023 · Paper
Collaboration with Google DeepMind
TL;DR: We introduce CCPO, a framework for versatile/adaptive safe RL that enables efficient training and zero-shot adaptation to varying safety constraints.
Zuxin Liu*, Zijian Guo*, Yihang Yao, Zhepeng Cen, Wenhao Yu, Tingnan Zhang, Ding Zhao
ICML 2023
TL;DR: We propose CDT for offline safe RL, which leverages a multi-objective optimization approach to balance safety and task performance, achieving superior adaptability, robustness, and high-reward policies with zero-shot adaptation capabilities.