Yihang Yao yihangya[at]andrew.cmu.edu

Hi, welcome to my website! I am a final-year Ph.D. candidate in the Safe AI Lab at Carnegie Mellon University , advised by Prof. Ding Zhao. I am also a research intern at NVIDIA Nemotron . I received my Bachelor's degree from Shanghai Jiao Tong University in 2022 and spent a wonderful time as a visiting student in the Intelligent Control Lab at CMU, working with Prof. Changliu Liu. I previously interned at Salesforce AI Research, MIT-IBM Watson AI Lab, IBM Research and have also been collaborating with Google DeepMind Robotics since 2023.

My research focuses on reinforcement learning and large language models, with a current emphasis on self-evolving agents and RL environment scaling for real-world applications such as terminal/coding agents [Technical Report 2026 ]. My prior research experience also includes:

Selected Works (* indicates equal contribution)
Learning Generalizable Behaviors for Terminal Agents

Yihang Yao, Bo Pang, Xuan Phi Nguyen, Ding Zhao, Shafiq Joty, Semih Yavuz

Technical Report, 2026  ·  Paper / Website / Salesforce Blog

Collaboration with Salesforce AI Research

Coding Agents Agentic RL RL Environments

TL;DR: We propose the agentic compositional generalization hypothesis, which posits that terminal-agent RL generalizes by shaping reusable behaviors that compose existing atomic skills. Using RL environments with high-quality verifiers to shape these behaviors, our training recipe achieves 106% and 30% larger average RL performance gains on Terminal-Bench-Lite and Terminal-Bench-v2.1, respectively, across models ranging from 2B to 27B parameters, while using fewer than 30% of the environments in the TMax collection.

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization

Yihang Yao, Zhepeng Cen, Haohong Lin, Shiqi Liu, Zuxin Liu, Jiacheng Zhu, Zhang-Wei Hong, Laixi Shi, Ding Zhao

ICML 2026  ·  Paper / Website / Code

Collaboration with Salesforce AI Research

Agentic RL User Simulation Interactive Agents

TL;DR: How can we better balance user burden and task performance in multi-turn interactive tasks? We introduce Behavioral Agentic Optimization (BAO), an agentic RL pipeline that combines simulated-user training with turn-wise reward regularization. On UserRL tasks, BAO achieves state-of-the-art performance while requiring less user intervention than the evaluated RL baselines.

Tailored Primitive Initialization is the Secret Key to Reinforcement Learning

Yihang Yao, Guangtao Zeng, Raina Wu, Yang Zhang, Ding Zhao, Zhang-Wei Hong, Chuang Gan

ACL 2026 Main  ·  Paper

Collaboration with MIT-IBM Watson AI Lab, IBM Research

LLM Reasoning RL Foundations Learning Dynamics Automatic Data Pipeline

TL;DR: How can we improve RL efficiency and final performance? We introduce Tailor, an automated SFT data-curation pipeline that warm-starts RL with diverse reasoning primitives, including skills and reasoning patterns. This initialization broadens exploration during RL, improves sample efficiency, and raises the final performance ceiling.

Behavior Injection: Preparing Language Models for Reinforcement Learning

Zhepeng Cen*, Yihang Yao*, William Han, Zuxin Liu, Ding Zhao

NeurIPS 2025  ·  Paper / Website / Code

Collaboration with Salesforce AI Research

LLM Reasoning RL Foundations Learning Dynamics

TL;DR: Why do LLMs respond differently to RL and exhibit varying data efficiency? We explain these differences through data co-influence effects and identify reasoning patterns as an important factor in RL efficiency. We also introduce a structured data pipeline that prepares LLMs for RL by injecting exploratory and exploitative behaviors into their reasoning traces.

Your Language Model May Think Too Rigidly: Achieving Reasoning Consistency with Symmetry-Enhanced Training

Yihang Yao, Zhepeng Cen, Miao Li, William Han, Yuyou Zhang, Emerson Liu, Zuxin Liu, Chuang Gan, Ding Zhao

ACL 2025 Findings  ·  Paper

Collaboration with Salesforce AI Research

LLM Reasoning Robustness

TL;DR: MEND enhances LLM robustness by improving query symmetry awareness, boosting reasoning performance through structured dataset curation.

OASIS: Conditional Distribution Shaping for Offline Safe Reinforcement Learning

Yihang Yao*, Zhepeng Cen*, Wenhao Ding, Haohong Lin, Shiqi Liu, Tingnan Zhang, Wenhao Yu, Ding Zhao

NeurIPS 2024  ·  Paper / Website / Code

Collaboration with Google DeepMind

Reinforcement Learning Data-centric ML AI Safety

TL;DR: We investigate offline RL from a data-centric perspective and propose a diffusion model-based data generator to curate training datasets aligned with user preferences.

Gradient Shaping for Multi-Constraint Safe Reinforcement Learning

Yihang Yao, Zuxin Liu, Zhepeng Cen, Peide Huang, Tingnan Zhang, Wenhao Yu, Ding Zhao

L4DC 2024  ·  Paper / Website / Code

Collaboration with Google DeepMind

Reinforcement Learning Multi-Objective Optimization AI Safety

TL;DR: We introduce GradS, a gradient-based method for improving training efficiency in multi-constraint RL by manipulating gradients, optimizing both reward and constraint satisfaction.

Constraint-Conditioned Policy Optimization for Versatile Safe Reinforcement Learning

Yihang Yao*, Zuxin Liu*, Zhepeng Cen, Jiacheng Zhu, Wenhao Yu, Tingnan Zhang, Ding Zhao

NeurIPS 2023  ·  Paper

Collaboration with Google DeepMind

Reinforcement Learning Multi-Objective Optimization AI Safety

TL;DR: We introduce CCPO, a framework for versatile/adaptive safe RL that enables efficient training and zero-shot adaptation to varying safety constraints.



Services