I am a first-year PhD candidate at the State Key Laboratory of CAD&CG, Zhejiang University. I earned my B.Eng. in Artificial Intelligence at the IIAU Lab, Dalian University of Technology, in 2025.

My research interest lies in the intersection of Reinforcement Learning, Reasoning and Machine Learning Systems, with their applications in complex, real-world environments. My long-term goal is to create an autonomous decision-making system that can act intelligently in any unknown environment. I am also interested in multimodal generation and its alignment.

Currently a research intern with the Qwen Team at Alibaba, previously at MSRA and BAAI. I'm also a community collaborator at WebAgentLab, an open-source community around GUI and Computer-Use Agents.

I'm always open to research collaboration. Feel free to reach out if you'd like to work together.

News

  • 2026.08 Started a research internship with the Qwen Team at Alibaba.
  • 2026.07 LiteResearcher accepted by COLM 2026.
  • 2026.03 Started a research internship at Microsoft Research Asia.
  • 2026.02 OmniGen2 accepted by CVPR 2026.
  • 2026.01 SAIL and TempFlow-GRPO accepted by ICLR 2026.
  • 2025.09 Started my PhD at Zhejiang University.

Research

My work spans three threads. * denotes equal contribution.

Agentic Reinforcement Learning

Training agents to reason and act over long horizons with scalable, stable RL.

LiteResearcher

LiteResearcher: Scalable Agentic RL Training for Deep Research Agents

Wanli Li, et al.

First Author COLM 2026

A scalable Agentic RL framework whose 4B model reaches open-source SOTA on GAIA (71.3%) and Xbench-DS (78.0%), rivaling commercial systems.

AREX

AREX: Towards a Recursively Self-Improving Agent for Deep Research

Shuqi Lu, Chaofan Li, Kun Luo, Zhang Zhang, Hui Wang, …, Wanli Li, et al.

Tech Report arXiv 2026

A recursively self-improving deep research agent that audits its own answer constraint-wise and relaunches targeted research; the 4B and 122B-A10B models lead comparable-scale baselines on BrowseComp, WideSearch and HLE.

Generalist Computer-Use Agents

Agents that operate real desktops across GUI, CLI and code — and how we evaluate them.

WeaveBench

WeaveBench: Benchmarking Hybrid-Interface Computer-Use Agents

Wanli Li, et al.

First Author Preprint 2026

A long-horizon benchmark of 114 hybrid-interface tasks forcing GUI and CLI/code to cooperate — the best agent reaches only 41.2% success.

ST-Lite

ST-Lite: Training-Free KV Cache Compression for Efficient GUI Agents

Wanli Li*, et al.

Co-first Preprint 2026

A training-free KV cache compression for GUI agents, giving 2.45× decoding speedup and +7.3% success on long-horizon tasks.

Multimodal Generation

Unified multimodal models and RL alignment for generation.

OmniGen2

OmniGen2: Towards Instruction-Aligned Multimodal Generation

Wanli Li (Core Contributor), et al.

CVPR 2026 332 Citations · 4.1k GitHub Stars

A unified multimodal generation model; I built the in-context data pipeline and led the progressive Flow-GRPO RL alignment.

TempFlow-GRPO

TempFlow-GRPO: When Timing Matters for GRPO in Flow Models

Wanli Li (Core Contributor), et al.

ICLR 2026

A temporally-aware RL framework for flow-matching models, using trajectory branching and noise-aware weighting for stabler alignment.

SAIL

SAIL: Self-Amplified Iterative Learning for Diffusion Model Alignment

Wanli Li*, et al.

Co-first ICLR 2026

A self-amplified iterative framework that aligns diffusion models from minimal seed data, letting the model act as its own teacher.

Talks

  • 2026.07 WeaveBench: Benchmarking Hybrid-Interface Computer-Use Agents Hunyuan Frontier Lab
  • 2026.07 LiteResearcher: Scalable Agentic RL for Deep Research Agents AgenticAICon 2026

Experience

Qwen Team, Alibaba Group Aug 2026 – Present

Research Intern · Base Model Team

Hybrid Computer-Use Agents and human–agent co-work: agents that move fluidly across GUI, CLI and code while staying steerable by a human collaborator.

Advisors: Chang Gao, Que Shen

Microsoft Research Asia Mar 2026 – Aug 2026

Research Intern · Shanghai AI/ML Group

Full-stack hybrid-interface Computer-Use Agents: VM sandbox infra, fully-async Agentic RL training, and the WeaveBench benchmark.

Advisors: Caihua Shan, Yifan Yang

SimplexAI May 2025 – Mar 2026

Research Intern · Agentic RL Team — First Author / Project Lead

Led LiteResearcher, an open-source SOTA 4B deep-research model trained with scalable Agentic RL over a local search world.

Advisors: Bo Zhang, Pan Zhang

BAAI Nov 2024 – May 2025

Research Intern · VectorSpaceLab — Core Contributor

Core contributor to OmniGen2 (4k★): the in-context data pipeline and progressive Flow-GRPO RL alignment.

Advisors: Zheng Liu, Shitao Xiao

Skills

Research Large-scale Model Training · Agentic RL · Agent Framework Development · Inference Optimization
Languages Python · C/C++ · LaTeX
Frameworks PyTorch · veRL · slime · vLLM · SGLang · DeepSpeed · TensorRT-LLM · Docker
Infra Kubernetes · Milvus · PostgreSQL · BGE-M3 · large-scale async rollout
AI Tooling Claude Code · Codex · Cursor · GitHub Copilot — rebuilding research & engineering workflows around LLM agents

Education

  • Zhejiang University — PhD in Computer Science, State Key Lab of CAD&CG 2025 – Present
    Advised by Prof. Bo Zhang
  • Dalian University of Technology — BEng in Artificial Intelligence, IIAU Lab 2021 – 2025
    Advised by Prof. Huchuan Lu and Prof. Lijun Wang

Honors

  • National Silver Award, China International College Students' Innovation Competition (“Internet+”), 2023
  • National Scholarship (Top 1%) — Ministry of Education, 2022 & 2024
  • Yulan Scholarship (Top 10 university-wide, highest honor) — DUT, 2024
  • Outstanding Graduate of Liaoning Province, 2024