I am a second-year PhD student at the State Key Laboratory of CAD&CG, Zhejiang University logoZhejiang University. I received my B.Eng. in Artificial Intelligence from Dalian University of Technology logoDalian University of Technology in 2025.

I work on agentic foundation models, especially computer-use agent models, with a focus on Agentic RL algorithms and training infrastructure. My goal is to build agents that learn and act reliably in unfamiliar real-world environments. I am also interested in multimodal generation and alignment.

I am currently a research intern with the Qwen logoQwen Team at Alibaba, working on computer-use agents. I previously interned at Microsoft Research Asia logoMicrosoft Research Asia (MSRA). I'm also a community collaborator at WebAgentLab, an open-source community for computer-use agents.

Open to research collaborations — feel free to get in touch.

News

Earlier milestones
  • 2026.03 Started a research internship at Microsoft Research Asia.
  • 2025.09 Started my PhD at Zhejiang University.

Research

My work spans three threads.

Agentic Reinforcement Learning

Training agents to reason and act over long horizons with scalable, stable RL.

System architecture: corpus synthesis and reinforcement curriculum learning.

LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent

Wanli Li, Bince Qu, Bo Pan, Jianyu Zhang, Zheng Liu, Pan Zhang, Wei Chen, Bo Zhang

COLM 2026 First author

A scalable Agentic RL framework whose 4B model scores 71.3% on GAIA-Text and 78.0% on Xbench-DS, leading the evaluated open-source models on Xbench-DS.

The outer self-improvement loop and inner research loop of AREX.

AREX: Towards a Recursively Self-Improving Agent for Deep Research

Shuqi Lu, Chaofan Li, Kun Luo, …, Wanli Li, et al.

Technical report Contributor

A deep research agent that checks its answers and launches targeted follow-up research, outperforming the evaluated baselines of comparable scale on BrowseComp, WideSearch and HLE.

Generalist Computer-Use Agents

Agents that operate real desktops across GUI, CLI and code — and how we evaluate them.

WeaveBench task taxonomy across eight real-world domains.

WeaveBench: Benchmarking Hybrid-Interface Computer-Use Agents

Wanli Li, Bowen Zhou, Yunyao Yu, Zhou Xu, Yifan Yang, Dongsheng Li, Caihua Shan

EMNLP 2026 First author

A long-horizon benchmark of 114 hybrid-interface tasks requiring agents to combine GUI interaction with CLI/code execution; the best evaluated agent achieves a 41.2% success rate.

ST-Lite Figure 1: limitations of existing methods compared with component-centric spatial saliency and trajectory-aware semantic gating.

ST-Lite: Training-Free KV Cache Compression with Spatio-Trajectory Guidance for Long-Horizon GUI Agents

Bowen Zhou*, Zhou Xu*, Wanli Li*, Jingyu Xiao, Pingan Gan, Haoqian Wang

EMNLP 2026 Co-first author

A training-free KV cache compression method for GUI agents, achieving up to 2.35× decoding speedup at fivefold compression in the reported experiments.

RecreationWorld overview, panels A and B: applications on five platforms and transfer gains from training.

RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents

Shuai Bai, Jiayong Deng, Yikun Fu, …, Wanli Li, et al.

Technical report Contributor

Software recreation provides scalable training environments with verifiable rewards and a 250-task benchmark across Ubuntu, macOS, Windows, Android and Web.

Multimodal Generation

Unified multimodal models and RL alignment for generation.

OmniGen2 image editing examples showing changes in style, color, objects, and accessories.

OmniGen2: Towards Instruction-Aligned Multimodal Generation

Chenyuan Wu, Jiahao Wang, Pengfei Zheng, …, Wanli Li, et al.

CVPR 2026 Core contributor 332 Citations · 4.1k GitHub Stars

A unified multimodal generation model; I built the in-context data pipeline and led the progressive Flow-GRPO RL alignment.

Images generated by TempFlow-GRPO with FLUX.1-dev.

TempFlow-GRPO: When Timing Matters for GRPO in Flow Models

Xiaoxuan He, Siming Fu, Yuke Zhao, Wanli Li, Jian Yang, Dacheng Yin, Fengyun Rao, Bo Zhang

ICLR 2026 Core contributor

A temporally aware RL framework for flow-matching models, using trajectory branching and noise-aware weighting for more stable alignment.

Progressive refinement from SDXL through successive SAIL iterations.

SAIL: Self-Amplified Iterative Learning for Diffusion Model Alignment with Minimal Human Feedback

Xiaoxuan He, Siming Fu, Wanli Li, Zhiyuan Li, Dacheng Yin, Kang Rong, Fengyun Rao, Bo Zhang

ICLR 2026 Core contributor

A self-amplified iterative framework that aligns diffusion models from minimal seed data, letting the model act as its own teacher.

Talks

  • Computer-Use Agents: Recent Advances and Future Directions Heart in the Bay · Computer Use Seminar
  • 2026.07 WeaveBench: Benchmarking Hybrid-Interface Computer-Use Agents Hunyuan Frontier Lab
  • 2026.07 How to Achieve Scaling in Agentic RL for Deep Research Agents AgenticAICon 2026

Open Source

LoopX

5.7k stars · 511 forks

Core Contributor · Long-Horizon Agent Control Plane

GitHub#1 on GitHub Trending · Aug 5, 2026

An open-source control layer for long-horizon agents, preserving task state and coordinating work across sessions and harnesses such as Codex and Claude Code. It supports recovery, human oversight, and empirical studies of agent behavior and harness design.

Experience

Qwen Team, Alibaba Group Aug 2026 – Present

Research Intern · Qwen Foundation Model Team

Hybrid Computer-Use Agents and human–agent collaboration: agents that move fluidly across GUI, CLI and code while staying steerable by a human collaborator.

Advisors: Chang Gao, Que Shen

Microsoft Research Asia Mar 2026 – Aug 2026

Research Intern · Shanghai AI/ML Group

Full-stack hybrid-interface Computer-Use Agents: VM sandbox infrastructure, fully asynchronous Agentic RL training, and the WeaveBench benchmark.

Advisors: Caihua Shan, Yifan Yang

SimplexAI May 2025 – Mar 2026

Research Intern · Agentic RL Team — First Author / Project Lead

Led LiteResearcher, an open-source 4B deep research model trained with scalable Agentic RL in a simulated search environment.

BAAI Nov 2024 – May 2025

Research Intern · VectorSpaceLab — Core Contributor

Core contributor to OmniGen2: the in-context data pipeline and progressive Flow-GRPO RL alignment.

Advisors: Zheng Liu, Shitao Xiao

Skills

Training Large-scale model training and Agentic RL · PyTorch, veRL, slime
Systems Agent frameworks, VM sandboxes and large-scale asynchronous rollouts · Docker, Kubernetes
Inference Efficient inference and deployment · vLLM, SGLang, TensorRT-LLM

Education

  • Zhejiang University — PhD in Computer Science, State Key Lab of CAD&CG 2025 – Present
    Advised by Prof. Wei Chen
  • Dalian University of Technology — B.Eng. in Artificial Intelligence, IIAU Lab 2021 – 2025
    Advised by Prof. Huchuan Lu and Prof. Lijun Wang

Honors

  • National Silver Award, China International College Students' Innovation Competition (“Internet+”), 2023
  • National Scholarship (Top 1%) — Ministry of Education, 2022 & 2024
  • Yulan Scholarship (Top 10 university-wide, highest honor) — DUT, 2024
  • Outstanding Graduate of Liaoning Province, 2025

View paper ↗