Portrait of Junhao Zhuang

Researcher · Generative Intelligence

Junhao Zhuang 庄俊豪

JD Future Academy · Tech Genius Team (TGT)

Curiosity fuels discovery.
Persistence unlocks the unknown.

My research focuses on large-scale audio-visual generative models and interactive world models for games and real-world environments.

ChinaTsinghua University
Profile

👋 About Me

I am currently a Researcher at JD Future Academy, JD.com, as a member of the Tech Genius Team (TGT).

I received my Master’s degree in Computer Technology from Tsinghua University in 2026, under the supervision of Prof. Chun Yuan. I obtained my Bachelor’s degree in Computer Science and Technology from the Yingcai Honors College at the University of Electronic Science and Technology of China in 2023, where I was fortunate to be advised by Prof. Xile Zhao.

Previously, I worked as a Research Assistant at MMLab, The Chinese University of Hong Kong (CUHK), under the supervision of Prof. Tianfan Xue.

My research focuses on large-scale audio-visual generative models and interactive world models for games and real-world environments.

Research focus Audio-visual generation · Interactive world models · Autoregressive video diffusion models · Long video generation · Efficient visual synthesis
Updates

News

Recent releases, publications, and research milestones.

Earlier news
Selected work

Research

Research spanning persistent audio-visual worlds, long-form video generation, visual editing, and restoration.

* indicates equal contribution.

JoyAI-Echo-1.5 project preview

JoyAI-Echo-1.5: Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

Tech Report · 2026

Nan Duan, Haoyang Huang, Weiyang Jin, Haoran Li, Yaowei Li, Yuming Li, Yijun Liu, Xin Lu, Xiaoxiao Ma, Yanwen Ma, Yaofeng Su, Yilang Sun, Haoyu Wang, Zeyue Xue, Songchun Zhang, Junhao Zhuang

JoyAI‑Echo‑1.5 is a unified audio‑visual generation system featuring two purpose‑built variants: a long‑video variant with composable cross‑shot memory for sustained identity consistency, and a world‑model variant with geometry‑aware 6‑DoF camera control for interactive viewpoint navigation.

GitHub stars for JoyAI-EchoGitHub forks for JoyAI-Echo
EchoWM project preview

EchoWM: Open and Enterable Omnimodal World Models

Tech Report · 2026

Songchun Zhang*, Yaowei Li*, Junhao Zhuang*, Weiyang Jin*, Haoyu Wang, Xin Lu, Shiyi Zhang, Haoran Li, Xiaoxiao Ma, Yumin Li, Yijun Liu, Yaofeng Su, Yanwen Ma, Haoyu Wu, Zihan Su, Yue Ma, Lvmin Zhang, Haoyang Huang, Zeyue Xue, Anyi Rao, Nan Duan

An omnimodal world model for generative media that responds to continuous navigation while video, environmental sound, music, and speech evolve together. I proposed Short-Horizon and Long-Horizon Audio-Visual Self-Gradient Forcing and was responsible for EchoWM’s causal training.

GitHub stars for JoyAI-EchoGitHub forks for JoyAI-Echo
Self Gradient Forcing project preview

Self Gradient Forcing: Native Long Video Extrapolation

Tech Report · 2026

Junhao Zhuang, Shiyi Zhang, Yuxuan Bian, Yaowei Li, Yawen Luo, Weiyang Jin, Songchun Zhang, et al.

Self Gradient Forcing (SGF) recovers the missing context-gradient path for self-generated causal memory through a bounded two-pass replay, enabling models trained with only a 5-second window to extrapolate to minute-scale videos with stronger identity, layout, and temporal stability. It also supports the causal training of EchoWM.

GitHub stars for Self Gradient ForcingGitHub forks for Self Gradient Forcing
JoyAI-Echo project preview

JoyAI-Echo: Pushing the Frontier of Long Audio-Visual Generation

Tech Report · 2026

Haoran Li, Fredreic Li, …, Junhao Zhuang, …, Zeyue Xue, Nan Duan

JoyAI-Echo is an interactive long video generation framework with boosted speed, stable audio-visual consistency and real-time editing, outperforming baseline models.

GitHub stars for JoyAI-EchoGitHub forks for JoyAI-Echo
ShotStream project preview

ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling

ECCV · 2026

Yawen Luo, Xiaoyu Shi, Junhao Zhuang, Yutian Chen, Quande Liu, Xintao Wang, Pengfei Wan, Tianfan Xue

ShotStream is a novel causal multi-shot architecture that enables interactive storytelling and efficient on-the-fly frame generation, achieving 16 FPS on a single NVIDIA GPU.

GitHub stars for ShotStreamGitHub forks for ShotStream
FlashVSR project preview

FlashVSR: Towards Real-Time Diffusion-Based Streaming Video Super-Resolution

CVPR · 2026

Junhao Zhuang, Shi Guo, Xin Cai, Xiaohui Li, Yihao Liu, Chun Yuan, Tianfan Xue

FlashVSR is a streaming, one-step diffusion-based video super-resolution framework with block-sparse attention and a Tiny Conditional Decoder. It reaches ~17 FPS at 768×1408 on a single A100 GPU. A Locality-Constrained Attention design further improves generalization and perceptual quality on ultra-high-resolution videos.

GitHub stars for FlashVSRGitHub forks for FlashVSR
Cobra project preview

Cobra: Efficient Line Art COlorization with BRoAder References

SIGGRAPH · 2025

Junhao Zhuang, Lingen Li, Xuan Ju, Zhaoyang Zhang, Chun Yuan, Ying Shan

Cobra is a novel efficient long-context fine-grained ID preservation framework for line art colorization, achieving high precision, efficiency, and flexible usability for comic colorization. By effectively integrating extensive contextual references, it transforms black-and-white line art into vibrant illustrations.

GitHub stars for CobraGitHub forks for Cobra
FlexiAct project preview

FlexiAct: Towards Flexible Action Control in Heterogeneous Scenarios

SIGGRAPH · 2025

Shiyi Zhang*, Junhao Zhuang*, Zhaoyang Zhang, Yansong Tang

We achieve action transfer in heterogeneous scenarios with varying spatial structures or cross-domain subjects.

GitHub stars for FlexiActGitHub forks for FlexiAct
PowerPaint project preview

A Task is Worth One Word: Learning with Task Prompts for High-Quality Versatile Image Inpainting

ECCV · 2024

Junhao Zhuang, Yanhong Zeng, Wenran Liu, Chun Yuan, Kai Chen

PowerPaint is the first versatile image inpainting model that simultaneously achieves state-of-the-art results in various inpainting tasks such as text-guided object inpainting, context-aware image inpainting, shape-guided object inpainting with controllable shape-fitting, and outpainting.

GitHub stars for PowerPaintGitHub forks for PowerPaint
BrushEdit project preview

BrushEdit: All-In-One Image Inpainting and Editing

TPAMI · 2025

Yaowei Li, Yuxuan Bian, Xuan Ju, Zhaoyang Zhang, Junhao Zhuang, Ying Shan, Yuexian Zou, Qiang Xu

BrushEdit is an all-in-one image inpainting and editing framework that combines multimodal large language models (MLLMs) with the enhanced dual-branch diffusion inpainting model BrushNetX. It supports free-form instruction-guided interactive editing, achieves superior performance in background preservation and text alignment, and provides a user-friendly multi-round editing experience.

GitHub stars for BrushEditGitHub forks for BrushEdit
Safe-Sora project preview

Safe-Sora: Safe Text-to-Video Generation via Graphical Watermarking

NeurIPS · 2025

Zihan Su, Xuerui Qiu, Hongbin Xu, Tangyu Jiang, Junhao Zhuang, Chun Yuan, Ming Li, Shengfeng He, Fei Richard Yu

Safe-Sora: a framework for embedding graphical watermarks into video generation, achieving state-of-the-art quality, fidelity, and robustness through hierarchical adaptive matching and a 3D wavelet-enhanced Mamba architecture.

ColorFlow project preview

ColorFlow: Retrieval-Augmented Image Sequence Colorization

Tech Report · 2024

Junhao Zhuang*, Xuan Ju*, Zhaoyang Zhang, Yong Liu, Shiyi Zhang, Chun Yuan, Ying Shan

ColorFlow is the first model designed for fine-grained ID preservation in image sequence colorization, utilizing contextual information. Given a reference image pool, ColorFlow accurately generates colors for various elements in black and white image sequences, including the hair color and attire of characters, ensuring color consistency with the reference images.

GitHub stars for ColorFlowGitHub forks for ColorFlow
TextureDiffusion project preview

TextureDiffusion: Target Prompt Disentangled Editing for Various Texture Transfer

ICASSP Oral · 2024

Zihan Su, Junhao Zhuang, Chun Yuan

We proposed TextureDiffusion, a tuning-free image editing method applied to various texture transfer.

UConNet project preview

UConNet: Unsupervised Controllable Network for Image and Video Deraining

ACM MM · 2022

Junhao Zhuang, Yisi Luo, Xile Zhao, Taixiang Jiang, Bichuan Guo

We propose the UConNet for image and video deraining. Our UConNet learns a relationship between trade-off parameters of the loss function and weightings of feature maps. At the inference stage, the weightings can be adaptively controlled to handle different rain scenarios, resulting in high generalization abilities. Extensive experimental results validate the effectiveness, generalization abilities, and efficiency of UConNet.

Background

Experience

Research across generative video, efficient diffusion, visual restoration, and editing.

Present

JD Future Academy, JD.com

Researcher · Tech Genius Team (TGT)

Large-scale audio-visual generative models and interactive world models.

Sep 2025 — Present

Kuaishou / KlingAI

Research Intern

Supervised by Yunyao Mao and Xintao Wang.
Topics: Video Generation.

May — Sep 2025

Shanghai AI Laboratory

Research Intern

Supervised by Shi Guo and Tianfan Xue.
Topics: Video Super-Resolution · Diffusion Acceleration · Sparse Attention.

May 2024 — Apr 2025

Tencent ARC Lab

Research Intern

Supervised by Zhaoyang Zhang and Ying Shan.
Topics: Comic Colorization · Video Generation · Diffusion.

Jul 2023 — Feb 2024

Shanghai AI Laboratory

Research Intern

Supervised by Yanhong Zeng and Kai Chen.
Topics: Image Inpainting · Diffusion.

Class of 2026

Tsinghua University

M.E. in Computer Technology

Advised by Prof. Chun Yuan. Previously a Research Assistant at MMLab, CUHK, advised by Prof. Tianfan Xue.

Class of 2023

University of Electronic Science and Technology of China

B.E. in Computer Science and Technology · Yingcai Honors College

Advised by Prof. Xile Zhao.

Recognition

Honors

Academic distinctions received during undergraduate and graduate study.

Excellent GraduateTsinghua University
Outstanding Master’s Degree ThesisTsinghua University
Comprehensive Excellence ScholarshipTsinghua University · 2024, 2025
Outstanding Bachelor’s Degree ThesisUniversity of Electronic Science and Technology of China
Outstanding Student ScholarshipUniversity of Electronic Science and Technology of China · 2020–2022
Audience

Visitor Map

Real-time visitor locations and traffic to this homepage.

Open live visitor statistics ↗