I am a Founding Research Scientist at Nuance Labs, where we build visual conversational AI that feels human.
I received my PhD from Texas A&M University, where I was advised by Nima Kalantari in the Aggie Graphics Group. My research focused on 3D Computer Vision and Computational Photography, with an emphasis on 3D scene reconstruction and generation.
Previously, I was a Founding AI Researcher at Morphic, specializing in 3D Computer Vision and Generative Video, focusing on controllable generation using techniques like 3D Gaussian Splatting, large diffusion models, and 3D/4D reconstruction. Before that, I was a Research Scientist Intern at Meta Reality Labs (2023), a Research Intern at Leia Inc. (2021), and a SDE Intern at Amazon (2019).
Two independent random-walk crops of one monocular video become a source-target pair, turning internet-scale footage into free supervision for camera-controlled reshooting of dynamic scenes.
Posing single-image 360° generation as joint panorama-and-depth optimization, solved by alternating minimization, keeps the lifted 3D scene globally coherent.
Two specialized diffusion priors, one repairing visible regions and one inpainting unseen ones, reconstruct a detailed 3D scene from just a handful of images.
Matching content and style Gaussian distributions with a Wasserstein-2 (Earth-Mover's) distance transfers scene detail directly, with no generative latent-space losses.
Decomposing wide stereo baselines into multi-plane disparities, while warping time non-uniformly, enables joint view-and-time interpolation of stereo video.
Encoding bidirectional flow into a coordinate network via a hypernetwork yields continuous intermediate flows that stay robust even when brightness changes between frames.
Aligning and fusing frames in stages avoids distant-frame misalignment, while a gradient-mask-conditioned discriminator stops the GAN from hallucinating noise in smooth regions.
A low-res, high-speed auxiliary camera supplies the true motion to reconstruct 1080p slow-motion video, capturing non-linear motion that interpolation alone misses.