Portrait of Jingxi Chen

Jingxi Chen

Ph.D. student in Computer Science, University of Maryland, College Park

Research Intern at Google, Amazon, Dolby

Email CV Google Scholar LinkedIn GitHub

My research focuses on image, video, and world generative models, computational photography, and robotic vision.

I am on the job market for full-time research scientist/engineer roles in Image/Video Generation and World Models. Feel free to reach out. ianchen@umd.edu

Bio

Welcome! I am a Ph.D. student in the Computer Science Department at the University of Maryland, College Park, where I work with Prof. Yiannis Aloimonos and Dr. Cornelia Fermüller in the Perception and Robotics Group. I also work closely with Prof. Christopher Metzler.

During my Ph.D., I have been a Student Researcher at Google Research, an Applied Scientist Intern at Amazon, and a Ph.D. Research Intern at Dolby Laboratories.

Selected Publications

For a complete publication list, please see Google Scholar.

First Frame Is the Place to Go for Video Content Customization
Jingxi Chen*, Zongxia Li*, Zhichao Liu, Guangyao Shi, Xiyang Wu, Fuxiao Liu, Cornelia Fermüller, Brandon Y. Feng, Yiannis Aloimonos.
CVPR 2026.
This work studies controlled video generation by composing content from multiple reference images. We propose a novel and efficient formulation that leverages the in-context capabilities of pre-trained video generation models.

From Inpainting to Layer Decomposition: Repurposing Generative Inpainting Models for Image Layer Decomposition
Jingxi Chen, Yixiao Zhang, Xiaoye Qian, Zongxia Li, Cornelia Fermüller, Caren Chen, Yiannis Aloimonos.
CVPR 2026.
Diffusion-based Image Layer Decomposition with a unified token-to-token model. Work done during my summer internship at Amazon Prime Video team.

Repurposing Pre-trained Video Diffusion Models for Event-based Video Interpolation
Jingxi Chen, Brandon Y. Feng, Haoming Cai, Tianfu Wang, Levi Burner, Dehao Yuan, Cornelia Fermüller, Christopher A. Metzler, Yiannis Aloimonos.
CVPR 2025.
In this work, we adapt pre-trained video diffusion models trained on internet-scale datasets to solve the specialized real-world video task of event-based video interpolation.

Temporally Consistent Atmospheric Turbulence Mitigation with Neural Representations
Haoming Cai*, Jingxi Chen*, Brandon Y. Feng, Weiyun Jiang, Mingyang Xie, Kevin W. Zhang, Cornelia Fermüller, Yiannis Aloimonos, Ashok Veeraraghavan, Christopher A. Metzler.
NeurIPS 2024.
ConVRT is an efficient INR framework for video-based turbulence mitigation that operates in test-time optimization manner.

Active Human Pose Estimation via an Autonomous UAV Agent

Active Human Pose Estimation via an Autonomous UAV Agent
Jingxi Chen, Botao He, Chahat D. Singh, Cornelia Fermüller, Yiannis Aloimonos.
IROS 2024.
We leverage radiance fields to imagine different human views to find the best drone pose for aerial cinematography.

Microsaccade-inspired Event Camera for Robotics
Botao He, Ze Wang, Yuan Zhou, Jingxi Chen, Chahat D. Singh, Haojia Li, Yuman Gao, Shaojie Shen, Kaiwei Wang, Yanjun Cao, Chao Xu, Yiannis Aloimonos, Fei Gao, Cornelia Fermüller.
Science Robotics, 2024.
Inspired by microsaccades, we designed an event-based perception system capable of simultaneously maintaining low reaction time and stable texture.

CodedEvents: Optimal Point-Spread-Function Engineering for 3D-Tracking with Event Cameras

CodedEvents: Optimal Point-Spread-Function Engineering for 3D-Tracking with Event Cameras
Sachin Shah, Matthew Chan, Haoming Cai, Jingxi Chen, Sakshum Kulshrestha, Chahat D. Singh, Christopher A. Metzler, Yiannis Aloimonos.
CVPR 2024.
CodedEvents is a novel method for optimal point-spread-function engineering for 3D-tracking with event cameras.

ProxMaP: Proximal Occupancy Map Prediction for Efficient Indoor Robot Navigation

ProxMaP: Proximal Occupancy Map Prediction for Efficient Indoor Robot Navigation
Vishnu D. Sharma, Jingxi Chen, Pratap Tokekar.
IROS 2023.
We present a self-supervised occupancy prediction technique, ProxMaP, to predict the occupancy within the proximity of the robot to enable faster navigation.

Multi-Agent Reinforcement Learning for Visibility-based Persistent Monitoring
Jingxi Chen*, Amrish Baskaran*, Zhongshun Zhang, Pratap Tokekar.
IROS 2021.
We present a Multi-Agent Reinforcement Learning (MARL) algorithm for the Visibility-based Persistent Monitoring (VPM) problem.

* Equal contribution.

Experiences

  • Google Research Summer 2026
    Student Researcher
    Conducted fundamental research on frontier video and world generation models, focusing on 3D geometry, motion dynamics, and long-term world consistency, with large-scale distributed training using JAX and TPUs.
  • Amazon Summer 2025
    Applied Scientist Intern
    Diffusion-based image layer decomposition (CVPR 2026).
    Strongly Inclined (Top 1% of all interns).
  • Dolby Laboratories Summer 2024
    Ph.D. Research Intern
    Neural event data compression and neural video codecs; two patents submitted.
  • University of Maryland Fall 2022 – Present
    Ph.D. in Computer Science
    Perception and Robotics Group.
  • Brain Corp Jun – Aug 2022
    Robotics Software Engineer Intern
  • University of Maryland 2020 – 2022
    M.S. in Computer Science
  • University of Maryland 2017 – 2020
    B.S. in Computer Science

Honors & Service

Fellowships & Awards
  • NeuroPAC Fellowship 2024
    Supported by the NSF grant "AccelNet: Accelerating Research on Neuromorphic Perception, Action, and Cognition."
  • Ph.D. Dean Fellowship 2022 – 2023
    University of Maryland, College Park.
  • John D. Gannon Endowed Scholarship
  • Capital One Bank Dean's Scholarship Fund in Computer Science
Service
  • Conference Reviewer
    CVPR 2025, ICCV 2025, ICRA 2021, 2023, 2024, IROS 2021, 2022, 2024.
  • Journal Reviewer
    IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI).