Geometry-guided control几何引导的控制
Pixel-aligned irradiance from target HDR lighting.由目标 HDR 光照构造与画面对齐的辐照度。
* Corresponding author* 通讯作者
Geometry guides the light.
Joint generation refines the result.用几何引导光照,
让联合生成修正细节。
Relighting asks a video model to change how a scene is illuminated while preserving its identity and motion. JointLight turns estimated geometry and target HDR lighting into an explicit irradiance cue, then jointly generates the relit RGB video and a refined irradiance video.重光照要求视频模型在保留场景内容与人物动作的同时,改变环境中的光影关系。JointLight 将估计几何与目标 HDR 光照转化为显式辐照度线索,再联合生成重光照 RGB 视频与精细化辐照度视频。
The key is to treat imperfect irradiance as a cue that can be corrected. Asynchronous denoising and domain matching reduce the influence of geometric artifacts while retaining large-scale lighting structure. Relighting fine-tuning uses synthetic data only, with generalization evaluated on real videos and HumanOLAT.关键是让不完美的辐照度参与生成并被逐步修正。异步去噪和域匹配减弱几何伪影的影响,同时保留大尺度光照结构。模型仅使用合成数据进行重光照任务微调,并在真实视频和 HumanOLAT 上验证泛化效果。
Pixel-aligned irradiance from target HDR lighting.由目标 HDR 光照构造与画面对齐的辐照度。
RGB and irradiance, with separate denoising schedules.RGB 与辐照度采用独立的去噪进度。
Diverse paired data and training-time domain matching.以多样化配对数据和域匹配支持迁移。
Explore static lighting, rotating illumination, joint irradiance outputs, and more real-world results.查看静态光照对比、旋转光照、联合辐照度输出与更多真实视频结果。
A coarse geometric proxy knows something about light transport, but its errors should not dictate every pixel.粗糙几何提供了光传输的依据,但其中的误差不应成为最终画面的约束。
We estimate monocular geometry, convert it into a mesh, and render a white diffuse scene under the target environment map with Mitsuba 3. This provides a pixel-aligned irradiance cue.从单目视频估计几何并转换为网格,再用 Mitsuba 3 在目标环境光下渲染白色漫反射场景,构造与画面对齐的辐照度线索。
Compared with ground-truth irradiance, coarse cues contain missing boundaries, incorrect shadows, and lost detail. Direct conditioning can transfer these artifacts into the output.相比真值辐照度,粗糙线索包含边界缺失、阴影偏差和细节损失,直接作为条件可能将这些伪影传入结果。
Based on Wan2.2-TI2V-5B, a shared diffusion Transformer models RGB and irradiance jointly. Independent noise timesteps let the model exchange information even when the two modalities are at different denoising stages. A second training stage introduces noisy coarse cues to match the inference distribution.基于 Wan2.2-TI2V-5B,在同一个扩散 Transformer 中联合建模 RGB 与辐照度。独立的噪声时间步让模型在两种模态清晰程度不同时仍能交换信息;第二阶段引入带噪粗糙线索,使训练与推理时的分布更接近。
At inference, we add noise to the rendered irradiance cue. The RGB branch denoises first while the cue remains fixed; at 0.6T, both modalities continue denoising together. Noise suppresses high-frequency artifacts, and joint generation refines the final lighting. The same switching point is used for all reported examples.推理时先对渲染得到的辐照度加噪,由 RGB 分支先行去噪,辐照度暂时固定;到 0.6T 时两种模态共同去噪。噪声减弱高频伪影,联合生成进一步修正光照细节。所有报告案例均使用相同切换时间步。
An Unreal Engine 5 pipeline renders paired RGB and irradiance videos while holding the scene and motion fixed and changing only the illumination. Photorealistic assets and diverse combinations of environments, characters, and motion support transfer to real footage.Unreal Engine 5 管线固定场景与人物动作,仅改变照明,同步渲染配对的 RGB 与辐照度视频。逼真的素材与多样化的环境、角色和动作组合,支持模型向真实影像迁移。
40 dynamic scenes with fixed target lighting. Tests relighting quality as the subject moves.40 个动态场景,目标环境光保持固定,评估人物运动过程中的重光照质量。
10 dynamic scenes with rotating target lighting. Tests simultaneous changes in motion and illumination.10 个动态场景,目标环境光随时间旋转,评估运动与照明同时变化时的控制能力。
Relighting fine-tuning uses synthetic data only. HumanOLAT is used for evaluation, not training. Dataset download: coming soon.重光照任务微调仅使用合成数据。HumanOLAT 用于评估,未参与训练。数据集下载入口待发布。
We measure image fidelity on synthetic videos and real captured images, and assess perceptual quality on in-the-wild videos.我们在合成视频与真实采集图像上评估重光照质量,并通过真实视频用户研究评估感知效果。
| Method方法 | PSNR ↑ | SSIM ↑ | LPIPS ↓ |
|---|---|---|---|
| DiffusionRenderer (Cosmos) | 13.48 | 0.589 | 0.345 |
| UniRelight | 14.90 | 0.602 | 0.303 |
| JointLight | 16.39 | 0.631 | 0.263 |
Selected leading baselines are shown here; the paper includes the full comparison. Bold values are best among the displayed methods. PSNR and SSIM: higher is better. LPIPS: lower is better.此处展示主要对比方法,完整结果见论文。粗体表示所列方法中的最优值。PSNR、SSIM 越高越好,LPIPS 越低越好。
13 participants rated 15 indoor and outdoor videos on a 1–5 scale. JointLight achieved the highest mean scores among evaluated methods for both visual quality and lighting fidelity.13 名参与者对 15 个室内外视频案例进行 1–5 分评分。JointLight 在参与比较的方法中获得最高的视觉质量与光照保真度平均分。
Ablations use the same 6K-step training budget. Direct irradiance conditioning reaches 14.10 dB PSNR, while the full method with asynchronous scheduling and domain matching reaches 15.69 dB. Joint modeling alone does not bridge the gap between coarse and ground-truth irradiance.消融模型使用相同的 6K 步训练预算。直接辐照度条件的 PSNR 为 14.10 dB,加入异步调度与域匹配的完整方法达到 15.69 dB。仅使用联合建模仍不足以弥合粗糙辐照度与真值之间的差距。
The full pipeline takes about 223 seconds for 81 frames at 832 × 480, using 50 sampling steps on one GPU. Processing is offline. Severe geometry errors and strongly specular materials remain challenging; see the paper for limitations.单 GPU、81 帧、832 × 480 分辨率与 50 步采样设置下,完整流程约需 223 秒,属于离线处理。严重几何误差与强镜面反射材质仍具有挑战,详见论文讨论。
@inproceedings{lv2026jointlight,
title = {JointLight: Human-Centric Video Relighting with
Asynchronous RGB-Irradiance Joint Modeling},
author = {Lv, Henglei and Deng, Bailin and Liu, Xiaoqiang and
Wang, Yichen and Zhang, Haoxian and Wan, Pengfei and Gao, Lin},
booktitle = {SIGGRAPH Asia 2026 Conference Papers},
year = {2026},
publisher = {Association for Computing Machinery},
doi = {10.1145/3829340.3842252}
}