About Me
I previously worked as a Senior Researcher at Microsoft Research Asia (MSRA), Shanghai, focusing on resource management, cloud computing, and systems for machine learning. My research covered GPU scheduling, distributed training, and efficient LLM serving and inference.
That systems background now shapes my startup work on LLM agents. I am a Scientist at Nex-AGI, where we build end-to-end agent systems spanning models, data, and frameworks. My current interests include agent reinforcement learning, context and harness engineering, and Software 3.0. We aim to make LLM agents useful in real work by helping them learn from experience and apply the problem-solving strategies of domain experts.
I obtained my Ph.D. degree from The University of Hong Kong (HKU) in 2020, advised by Prof. Francis C.M. Lau. Before that, I received my B.Eng in Electronic and Information Engineering from the University of Electronic Science and Technology of China (UESTC) in 2014.
Selected Publications
Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses
AHE automatically evolves coding-agent harnesses through component, experience, and decision observability, with improvements that transfer across tasks and base models.
NexAU-AHE + GPT-5.5: 84.7% ± 2.1 on Terminal-Bench 2.0 — ranked #1 as of .
Featured in Lilian Weng's Harness Engineering for Self-Improvement ().
Nex-N1: Agentic Models Trained via a Unified Ecosystem for Large-Scale Environment Construction
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
Best PaperMinference 1.0: Accelerating pre-filling for long-context LLMs via dynamic sparse attention
Parrot: Efficient Serving of LLM-based Applications with Semantic Variable
PIT: Optimization of Dynamic Sparse Deep Learning Models via Permutation Invariant Transformation
Optimizing Dynamic Neural Networks with Brainstorm
Dynamic Resource Allocation for Deep Learning Clusters with Separated Compute and Storage
ElasticFlow: An Elastic Serverless Training Platform for Distributed Deep Learning
SiloD: A Co-design of Caching and Scheduling for Deep Learning Clusters
PilotFish: Harvesting Free Cycles of Cloud Gaming with Deep Learning Training
HiveD: Sharing a GPU Cluster for Deep Learning with Guarantees
Retiarii: A Deep Learning Exploratory-Training Framework
Automating Cloud Deployment for Deep Learning Inference of Real-time Online Services
Gandiva: Introspective Cluster Scheduling for Deep Learning
Online File Caching in Latency-Sensitive Systems with Delayed Hits and Bypassing
Regularization-Based Coflow Scheduling in Optical Circuit Switches
Online Dispatching and Scheduling of Jobs with Heterogeneous Utilities in Edge Computing
Scheduling Placement-Sensitive BSP Jobs with Inaccurate Execution Time Estimation
OnDisc: Online Latency-Sensitive Job Dispatching and Scheduling in Heterogeneous Edge-Clouds
Joint Online Coflow Routing and Scheduling in Data Center Networks
Camul: Online Caching on Multiple Caches with Relaying and Bypassing
Energy Efficient Dynamic Virtual Machine Management in Data Centers
Efficient Online Learning Based Cross-Tier Uplink Scheduling in HetNets
Congestion Game with Agent and Resource Failures
Professional Services
Program Committee
- • IEEE INFOCOM 2021, 2022
- • MSN 2020
Journal Reviewer
- • IEEE/ACM Transactions on Networking