Reinforcement Learning
Efficient decision-making
Offline RL, meta-RL, diffusion planning, policy transfer, multi-task learning, and compact symbolic policies.
Researching reinforcement learning, embodied intelligence, and large language models — with an emphasis on efficient decision-making, generalizable agents, and systems that connect reasoning with action.
About
I am a researcher at the Institute of Computing Technology, Chinese Academy of Sciences. My work studies how learning-based agents can make decisions more efficiently, transfer across tasks and environments, and connect high-level reasoning with low-level interaction.
My recent research spans reinforcement learning and meta-reinforcement learning, diffusion-based planning, embodied agents, vision-language-action systems, and large-language-model agents.
Research directions
Build agents that can learn reusable structure, reason over long horizons, and remain practical to deploy.
Reinforcement Learning
Offline RL, meta-RL, diffusion planning, policy transfer, multi-task learning, and compact symbolic policies.
Embodied Intelligence
Reusable motion primitives, VLA models, vision-and-language navigation, open-ended agents, and autonomous verification.
Large Language Models
LLM agents, grounded skill learning, code-driven planning, long-context RL, and evaluation of model capabilities.
Selected publications
Temporal Diffusion Planner reuses and progressively updates plans to substantially improve decision frequency while preserving planning performance.
Distribution-aware sparse attention and algorithm-hardware co-design for efficient Diffusion Transformer inference.
A structured CRUX intermediate space and two-stage training framework for more precise natural-language-to-Verilog generation.
Signal-aware verification, code extraction, and DPO turn partially correct Verilog generations into fine-grained functional learning signals.
Only Support Constraint (OSC) restricts the learned policy to behavior-policy support while avoiding unnecessary constraints within the support.
ACL 2024 Long Paper. † Equal contribution; * Corresponding author.
End-to-end differentiable symbolic policies for efficient, compact, and interpretable reinforcement learning.
A hindsight-conditioned value baseline that leverages future information to reduce policy-gradient variance in stochastic environments.