Yuming Li

profile photo

I am a PhD student at the University of Hong Kong, affiliated with the Musketeers Foundation Institute of Data Science, where I am advised by Andrew Luo.

I received my Master's degree in Software Engineering from Peking University, where I worked with Shanghang Zhang. Before that, I earned my Bachelor's degree in Computer Science and Technology from Northwestern Polytechnical University.

My research focuses on real-time long video generation, reinforcement learning alignment for diffusion and video generation models, controllable multimodal generation, and embodied world modeling. I am interested in building generative video systems that are fast, controllable, and temporally consistent.

Email  /  Google Scholar  /  Github

profile photo

Research

I work on generative video systems for long-horizon, controllable, and real-time generation. My recent work studies structured RL post-training for diffusion models, streaming audio-visual generation, fast video initialization, and multimodal world models.

Selected Research

JoyAI-Echo-1.5 demo Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
Nan Duan, Haoyang Huang, Weiyang Jin, Haoran Li, Yaowei Li, Yuming Li, et al.
Core Contributor
JoyAI-Echo-1.5  ·  Technical Report, August 2026
project page / arXiv / code

A unified audio-visual system for persistent multi-shot stories and interactive 6-DoF worlds, powered by composable cross-shot memory and few-step causal generation.

JoyAI-Echo demo JoyAI-Echo: Pushing the Frontier of Long Audio-Visual Generation
Echo Team @ Joy Future Academy, JD; Yuming Li* (co-first author, listed second)
* Equal contribution
Technical Report, 2026  ·  1.9k GitHub stars
project page / technical report / code / Hugging Face

A memory-driven audio-visual generation framework for minute-level coherent video and interactive generation.

OmniForcing method OmniForcing: Unleashing Real-time Joint Audio-Visual Generation
Yaofeng Su*, Yuming Li*, Zeyue Xue, Jie Huang, Siming Fu, Haoran Li, Haoyang Huang, Nan Duan
* Equal contribution
ECCV 2026  ·  Oral Presentation
project page / arXiv / code / model

A real-time streaming framework for joint audio-visual generation with long-form multimodal consistency.

BranchGRPO method BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models
Yuming Li*, Yikai Wang*, Yuying Zhu, Zhongyu Zhao, Ming Lu, Qi She, Shanghang Zhang
* Equal contribution
ICLR 2026
OpenReview / arXiv / project page / code

A structured branching rollout strategy for stable and efficient GRPO training in diffusion models.

Experience

JD.com Exploratory Research Institute
Research Intern, 2025-present.

Kuaishou - Kling AI
Research Intern, 2025.

Hedra AI
Research Intern, 2025.

ByteDance
Research Intern, 2024-2025.

Tencent Robotics X
Research Intern, 2024.


Template inspired by Andrew Luo. Last updated October 2026.