# Haoming Cai (蔡昊明) > CS PhD student (expected 2027) at the University of Maryland, College Park, advised by Christopher Metzler. Works on generative media: teaching generative models to respect light, geometry, and the camera. On the job market for full-time roles from 2027. Previously an undergraduate research assistant at XPixel (SIAT, Shenzhen) with Chao Dong and Jinjin Gu; BEng from The Chinese University of Hong Kong, Shenzhen. ## Contact - Homepage: https://www.hm-cai.com/ - Email: hmcai@umd.edu - CV (PDF, always the latest version): https://www.hm-cai.com/cv.pdf - Google Scholar: https://scholar.google.com/citations?user=Jzs37TkAAAAJ ## Research topics Generative media; computational imaging; image and video diffusion models; light field synthesis and bokeh rendering; portrait lighting control; reflection removal; event cameras; atmospheric turbulence mitigation; Gaussian splatting; image quality assessment; image super-resolution and restoration. ## Experience - Summer 2026 — Adobe, San Jose, CA. Ph.D. Research Intern, NextCam team (Marc Levoy). - Summer 2025 — Adobe, San Jose, CA. Ph.D. Research Intern, NextCam team (Marc Levoy), supervised by Shumian Xin and Zhoutong Zhang. Bokeh editing project; outcome: ECCV 2026. Patent filing in process. - Summer 2024 — Dolby Vision Lab, Sunnyvale, CA. Ph.D. Research Intern, supervised by Guan-Ming Su. Portrait lighting control for text-to-image diffusion models without light stage data. Outcome: ICCV 2025, one patent. - 2020 – 2022 — XPixel Group, SIAT, Shenzhen, China. Undergraduate Research Assistant, supervised by Chao Dong. Efficient and controllable image restoration; image quality and aesthetic assessment. Outcome: two ECCV papers, two CVPR workshops organized, one patent. ## Publications Newest first. Each title links to a page for that paper. - [Real2SAM2Real: Generative 3D Caches as Complementary Context for Video Diffusion](https://www.hm-cai.com/papers/real2sam2real/) — Manuscript (2026). Authors: Jiayi Wu*, Haoming Cai*, Cornelia Fermuller, Christopher Metzler, Yiannis Aloimonos. Control the camera and the objects in video generation with a generative 3D cache. PDF: https://arxiv.org/pdf/2606.00299 Project: https://jiayi-wu-leo.github.io/real2sam2real/ - [Large-Scale Light Field Synthesis from Videos Enables Geometrically Consistent Bokeh Editing](https://www.hm-cai.com/papers/vid2bokeh/) — ECCV 2026. Authors: Haoming Cai, Zhoutong Zhang, Christopher Metzler, Shumian Xin. Render geometrically consistent bokeh on scenes that defeat other methods — transparent surfaces, thin structures, cluttered depth. Project: https://www.hm-cai.com/vid2bokeh/ - [Parametric Shadow Control for Portrait Generation in Text-to-Image Diffusion Models](https://www.hm-cai.com/papers/shadowdirector/) — ICCV 2025. Authors: Haoming Cai, Tsung-Wei Huang, Shiv Gehlot, Brandon Y. Feng, Sachin Shah, Guan-Ming Su, Christopher Metzler. Add a shadow knob to AI portraits, without heavy compute. PDF: https://arxiv.org/pdf/2503.21943 Project: https://www.hm-cai.com/tn/projects_html/ShadowDirector/ - [Flash-Split: 2D Reflection Removal with Flash Cues and Latent Separation](https://www.hm-cai.com/papers/flashsplit/) — CVPR 2025. Authors: Tianfu Wang*, Mingyang Xie*, Haoming Cai, Sachin Shah, Christopher Metzler. Separate reflection from transmission with flash cues and a diffusion prior. PDF: https://arxiv.org/pdf/2501.00637 - [Repurposing Pre-trained Video Diffusion Models for Event-based Video Interpolation](https://www.hm-cai.com/papers/timerewind/) — CVPR 2025. Authors: Jingxi Chen, Brandon Y. Feng, Haoming Cai, Tianfu Wang, Levi Burner, Dehao Yuan, Cornelia Fermuller, Christopher Metzler, Yiannis Aloimonos. Turn event streams into smooth video with a video diffusion model. PDF: https://arxiv.org/pdf/2412.07761 Project: https://vdm-evfi.github.io/ - [Temporally Consistent Atmospheric Turbulence Mitigation with Neural Representations](https://www.hm-cai.com/papers/convrt/) — NeurIPS 2024. Authors: Haoming Cai*, Jingxi Chen*, Brandon Y. Feng, Weiyun Jiang, Mingyang Xie, Kevin Zhang, Cornelia Fermuller, Yiannis Aloimonos, Ashok Veeraraghavan, Christopher Metzler. Low-pass the temporal dimension to strip atmospheric turbulence out of video. PDF: https://proceedings.neurips.cc/paper_files/paper/2024/file/4eb91efe090f72f7cf42c69aab03fe85-Paper-Conference.pdf - [Flash-Splat: 3D Reflection Removal with Flash Cues and Gaussian Splats](https://www.hm-cai.com/papers/flashsplat/) — ECCV 2024. Authors: Mingyang Xie*, Haoming Cai*, Sachin Shah, Yiran Xu, Brandon Y. Feng, Jia-bin Huang, Christopher Metzler. Separate reflections in 3D using flash-induced cues. PDF: https://arxiv.org/pdf/2410.02764 Project: https://mingyangx.github.io/Flash-Splat/ - [CodedEvents: Optimal Point-Spread-Function Engineering for 3D-Tracking with Event Cameras](https://www.hm-cai.com/papers/codedevent/) — CVPR 2024. Authors: Sachin Shah, Matthew Albert Chan, Haoming Cai, Jingxi Chen, Sakshum Kulshrestha, Chahat Deep Singh, Yiannis Aloimonos, Christopher Metzler. Engineer the PSF so event cameras encode depth for free. - [Assessor360: Multi-sequence Network for Blind Omnidirectional Image Quality Assessment](https://www.hm-cai.com/papers/assessor360/) — NeurIPS'23 (2023). Authors: Tianhe Wu, Shuwei Shi, Haoming Cai, Mingdeng Cao, Jing Xiao, Yinqiang Zheng, Yujiu Yang. Assess 360° image quality the way a viewer actually explores the sphere. PDF: https://arxiv.org/abs/2305.10983 Project: https://github.com/TianheWu/Assessor360 - [Snow Removal in Video: A New Dataset and A Novel Method](https://www.hm-cai.com/papers/snow/) — ICCV 2023. Authors: Haoyu Chen, Jingjing Ren, Jinjin Gu, Hongtao Wu, Xuequan Lu, Haoming Cai, Lei Zhu. Remove snow from video consistently across frames, learned from real snowfall rather than synthetic. PDF: https://openaccess.thecvf.com/content/ICCV2023/html/Chen_Snow_Removal_in_Video_A_New_Dataset_and_A_Novel_ICCV_2023_paper.html Project: https://haoyuchen.com/VideoDesnowing - [Super-resolution by predicting offsets: An ultra-efficient super-resolution network for rasterized images](https://www.hm-cai.com/papers/srpo/) — ECCV 2022. Authors: Jinjin Gu, Haoming Cai, Chenyu Dong, Ruofan Zhang, Yulun Zhang, Wenming Yang, Chun Yuan. Super-resolve by predicting where to sample, not what to generate. PDF: https://arxiv.org/pdf/2210.04198 Project: https://github.com/HaomingCai/SRPO - [Efficient image super-resolution using vast-receptive-field attention](https://www.hm-cai.com/papers/vapsr/) — AIM Challenge @ ECCV 2022. Authors: Haoming Cai*, Lin Zhou*, Jinjin Gu, Zheyuan Li, Yingqi Liu, Xiangyu Chen, Yu Qiao, Chao Dong. Widen the receptive field with attention instead of depth. PDF: https://arxiv.org/pdf/2210.05960 Project: https://github.com/zhoumumu/VapSR - [Blueprint separable residual network for efficient image super-resolution](https://www.hm-cai.com/papers/bsrn/) — NTIRE Challenge @ CVPR 2022. Authors: Zheyuan Li, Yingqi Liu, Xiangyu Chen, Haoming Cai, Jinjin Gu, Yu Qiao, Chao Dong. Super-resolve on a tight compute budget — winner of the NTIRE’22 efficient SR track. PDF: https://openaccess.thecvf.com/content/CVPR2022W/NTIRE/papers/Li_Blueprint_Separable_Residual_Network_for_Efficient_Image_Super-Resolution_CVPRW_2022_paper.pdf Project: https://github.com/xiaom233/BSRN - [Toward interactive modulation for photo-realistic image restoration](https://www.hm-cai.com/papers/cugan/) — CVPRW'21 (2021). Authors: Haoming Cai, Jingwen He, Yu Qiao, Chao Dong. Dial deblurring and denoising continuously, on a GAN, without retraining. PDF: https://openaccess.thecvf.com/content/CVPR2021W/NTIRE/papers/Cai_Toward_Interactive_Modulation_for_Photo-Realistic_Image_Restoration_CVPRW_2021_paper.pdf Project: https://github.com/HaomingCai/CUGAN - [PIPAL: A Large-Scale Image Quality Assessment Dataset for Perceptual Image Restoration](https://www.hm-cai.com/papers/pipal/) — ECCV 2020. Authors: Jinjin Gu, Haoming Cai, Haoyu Chen, Xiaoxing Ye, Jimmy S. Ren, Chao Dong. Score image quality on the artifacts that modern restoration models actually produce. PDF: https://arxiv.org/pdf/2007.12142 ## Optional - Interactive figures from the two featured papers are on the homepage (drag across them to move the light, and to pull focus). - Off-hours photography: https://www.flickr.com/photos/201416778@N04/