Patents by Inventor Deqing Sun

Deqing Sun has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Publication number: 20260203923
    Abstract: Improved methods are provided for generating, via a noise-diffusion iterative process, depth maps or optical flow maps from input images. Also provided are improved methods for training the machine learning model(s) employed in the iterative process and for augmenting die set of training data used to train such models. By translating the depth or optical flow map prediction process into the noise diffusion context, improved performance with respect to compute cost, training data, requirements, model size, and output quality are obtained. Additionally, the noise diffusion context allows models trained as described herein to generate maps de novo from target color images and/or to begin from initial ‘guess’ maps (e.g., noisy maps, maps containing holes) when generating improved output maps, natively incorporating the imperfect prior information represented by such initial maps.
    Type: Application
    Filed: January 26, 2024
    Publication date: July 16, 2026
    Inventors: Saurabh SAXENA, Mohammad NOROUZI, David James FLEET, Abhishek KAR, Charles Irwin HERRMANN, Junhwa HUR, Deqing SUN
  • Patent number: 12632974
    Abstract: Apparatuses, systems, and techniques are presented to determine distance and/or motion of one or more objects represented in one or more images. In at least one embodiment, distance and motion are determined using segmentations, single-view depth, and optical flow data inferred for one or more images captured in sequence using a monocular camera.
    Type: Grant
    Filed: February 10, 2020
    Date of Patent: May 19, 2026
    Assignee: NVIDIA Corporation
    Inventors: Zhaoyang Lv, Kihwan Kim, Alejandro Troccoli, Deqing Sun, Jan Kautz
  • Publication number: 20260118049
    Abstract: A method of controlling a refrigerator includes instructing at least one camera to capture images forward of a main body opening of a main body of the refrigerator of items being placed into or removed from at least one storage compartment. The images are analyzed to determine whether the images contain an object of interest. The method further includes assigning a direction of movement for the object of interest being at least one of into or out of the storage compartment. An inventory for the storage compartment is updated based on the object of interest and the direction of movement for the object of interest. Instructions are then provided to at least one environmental control device. The instructions control the environmental control device to change a storage condition within the storage compartment based on the inventory in the storage compartment.
    Type: Application
    Filed: October 31, 2024
    Publication date: April 30, 2026
    Inventors: John Anderson, Tyler Sanborn, Boris Kontorovich, Bryce Copenhaver, Deqing Sun
  • Patent number: 12573161
    Abstract: A computing system and method can be used to render a 3D shape from one or more images. In particular, the present disclosure provides a general pipeline for learning articulated shape reconstruction from images (LASR). The pipeline can reconstruct rigid or nonrigid 3D shapes. In particular, the pipeline can automatically decompose non-rigidly deforming shapes into rigid motions near rigid-bones. This pipeline incorporates an analysis-by-synthesis strategy and forward-renders silhouette, optical flow, and color images which can be compared against the video observations to adjust the internal parameters of the model. By inverting a rendering pipeline and incorporating optical flow, the pipeline can recover a mesh of a 3D model from the one or more images input by a user.
    Type: Grant
    Filed: December 21, 2020
    Date of Patent: March 10, 2026
    Assignee: GOOGLE LLC
    Inventors: Deqing Sun, Varun Jampani, Gengshan Yang, Daniel Vlasic, Huiwen Chang, Forrester H. Cole, Ce Liu, William Tafel Freeman
  • Patent number: 12566963
    Abstract: Systems and methods to train a convolutional neural network having two or more filter layers having different filtering parameters corresponding to respective different portions of a digital representation of an image. A processor comprising one or more arithmetic logic units (ALUs) to be configured to identify one or more features within an image based, at least in part, on a convolutional neural network having two or more filter layers having different filtering parameters corresponding to respective different portions of a digital representation of the image.
    Type: Grant
    Filed: March 20, 2019
    Date of Patent: March 3, 2026
    Assignee: NVIDIA Corporation
    Inventors: Varun Jampani, Hang Su, Deqing Sun, Orazio Gallo, Erik G. Learned-Miller, Jan Kautz
  • Publication number: 20260032794
    Abstract: A refrigerator system includes a main body and a door mounted to and movable with respect to the main body, an ambient light sensor, a light source, at least one camera positioned to have a field of view that includes an entrance opening leading to the at least one storage compartment, and at least one computing device in operable connection with the at least one camera, the ambient light sensor, and the light source.
    Type: Application
    Filed: July 29, 2024
    Publication date: January 29, 2026
    Inventors: Boris Kontorovich, Bryce Copenhaver, Tyler Sanborn, John Anderson, Deqing Sun
  • Publication number: 20250363590
    Abstract: Despite recent progress, existing frame interpolation methods still struggle with extremely high resolution images and challenging cases such as repetitive textures, thin objects, and fast motion. To address these issues, provided is a cascaded diffusion frame interpolation approach that excels in these scenarios while achieving competitive performance on standard benchmarks.
    Type: Application
    Filed: May 22, 2025
    Publication date: November 27, 2025
    Inventors: Deqing Sun, Junhwa Hur, Charles Irwin Herrmann, Saurabh Saxena, David James Fleet, Janne Matias Kontkanen, Wei-Sheng Lai, Yichang Shih, Michael Rubinstein
  • Publication number: 20250310646
    Abstract: This document describes systems and techniques directed at fusing optically zoomed images into one digitally zoomed image. In aspects, a computing device having at least two cameras and an image-processing manager is configured to receive, from a first camera, a first image at a first optical zoom and, from a second camera, a second image at a second optical zoom different from the first optical zoom. The first and second cameras capture a same scene from different fields of view and different points of view. The image-processing manager receives a desired digital zoom between the first optical zoom and the second optical zoom. Based on the first and second images, the image-processing manager determines an overlap region of the first image in which the second image overlaps the first image.
    Type: Application
    Filed: May 17, 2022
    Publication date: October 2, 2025
    Inventors: Xiaotong Wu, Chia-Kai Liang, Wei-Sheng Lai, Yichang Shih, Deqing Sun, Michael Krainin, Lun-Cheng Chu
  • Publication number: 20250238905
    Abstract: Provided is a video generation model for performing text-to-video (T2V) or other video generation techniques. The proposed model reduces the computational costs associated with video generation. In particular, unlike traditional T2V methods, the disclosed technology can generate the full temporal duration of a video clip at once, bypassing the need for extensive computation. As one example, a machine-learned denoising diffusion model can simultaneously process a plurality of noisy inputs that correspond to various timestamps spanning the temporal dimension of a video to simultaneously generate synthetic frames for the video that match the timestamps.
    Type: Application
    Filed: January 22, 2025
    Publication date: July 24, 2025
    Inventors: Inbar Mosseri, Omer Bar Tal, Hila Chefer-Livshen, Omer Tov, Charles Irwin Herrmann, Rony Paiss, Shiran Elyahu Zada, Ariel Ephrat, Junhwa Hur, Guanghui Liu, Amit Raj, Yuanzhen Li, Michael Rubinstein, Tomer Michaeli, Oliver Wang, Deqing Sun, Tali Dekel
  • Publication number: 20250166135
    Abstract: Methods, systems, and apparatus, including computer programs encoded on computer storage media, for controllable video generation. One of the methods includes receiving a text prompt that specifies an object; receiving a control input that comprises an image that depicts a particular instance of the object; generating a video that comprises a respective video frame at each of a plurality of time steps in the video and that depicts the particular instance of the object. Generating the video includes, at each of the plurality of time steps: obtaining a text prompt embedding; obtaining a control input embedding; and generating the respective video frame at the time step using a video generation neural network while the video generation neural network is conditioned on the text prompt embedding and on the control input embedding.
    Type: Application
    Filed: November 18, 2024
    Publication date: May 22, 2025
    Inventors: Yu-Chuan Su, Hsin-Ping Huang, Ming-Hsuan Yang, Deqing Sun, Lu Jiang, Yukun Zhu, Xuhui Jia
  • Publication number: 20240265490
    Abstract: Provided is a computer system that includes one or more processors and one or more non-transitory computer-readable media that collectively store a machine-learned image interpolation model. The machine-learned image interpolation model is configured to: extract, for each of multiple different scales, a respective set of feature values from each of a pair of input images; generate, for each of the multiple different scales, a respective flow estimate for each of the pair of input images that indicates a respective flow from the interpolation time to the respective capture time; warp, for each of the multiple different scales, the respective set of feature values for each of the pair of input images according to the respective flow estimate to generate respective warped sets of features; and generate a interpolated image based on the respective warped sets of features for the pair of input images and the multiple different scales.
    Type: Application
    Filed: February 8, 2024
    Publication date: August 8, 2024
    Inventors: Janne Matias Kontkanen, Eric Tabellion, Brian Lee Curless, Fitsum Reda, Deqing Sun, Caroline Rebecca Pantofaru
  • Publication number: 20240212325
    Abstract: Systems and methods for training models to predict dense correspondences across images such as human images. A model may be trained using synthetic training data created from one or more 3D computer models of a subject. In addition, one or more geodesic distances derived from the surfaces of one or more of the 3D models may be used to generate one or more loss values, which may in turn be used in modifying the model's parameters during training.
    Type: Application
    Filed: March 6, 2024
    Publication date: June 27, 2024
    Inventors: Yinda Zhang, Feitong Tan, Danhang Tang, Mingsong Dou, Kaiwen Guo, Sean Ryan Francesco Fanello, Sofien Bouaziz, Cem Keskin, Ruofei Du, Rohit Kumar Pandey, Deqing Sun
  • Patent number: 11954899
    Abstract: Systems and methods for training models to predict dense correspondences across images such as human images. A model may be trained using synthetic training data created from one or more 3D computer models of a subject. In addition, one or more geodesic distances derived from the surfaces of one or more of the 3D models may be used to generate one or more loss values, which may in turn be used in modifying the model's parameters during training.
    Type: Grant
    Filed: March 11, 2021
    Date of Patent: April 9, 2024
    Assignee: GOOGLE LLC
    Inventors: Yinda Zhang, Feitong Tan, Danhang Tang, Mingsong Dou, Kaiwen Guo, Sean Ryan Francesco Fanello, Sofien Bouaziz, Cem Keskin, Ruofei Du, Rohit Kumar Pandey, Deqing Sun
  • Publication number: 20240046618
    Abstract: Systems and methods for training models to predict dense correspondences across images such as human images. A model may be trained using synthetic training data created from one or more 3D computer models of a subject. In addition, one or more geodesic distances derived from the surfaces of one or more of the 3D models may be used to generate one or more loss values, which may in turn be used in modifying the model's parameters during training.
    Type: Application
    Filed: March 11, 2021
    Publication date: February 8, 2024
    Inventors: Yinda Zhang, Feitong Tan, Danhang Tang, Mingsong Dou, Kaiwen Guo, Sean Ryan Francesco Fanello, Sofien Bouaziz, Cem Keskin, Ruofei Du, Rohit Kumar Pandey, Deqing Sun
  • Publication number: 20240013497
    Abstract: A computing system and method can be used to render a 3D shape from one or more images. In particular, the present disclosure provides a general pipeline for learning articulated shape reconstruction from images (LASR). The pipeline can reconstruct rigid or nonrigid 3D shapes. In particular, the pipeline can automatically decompose non-rigidly deforming shapes into rigid motions near rigid-bones. This pipeline incorporates an analysis-by-synthesis strategy and forward-renders silhouette, optical flow, and color images which can be compared against the video observations to adjust the internal parameters of the model. By inverting a rendering pipeline and incorporating optical flow, the pipeline can recover a mesh of a 3D model from the one or more images input by a user.
    Type: Application
    Filed: December 21, 2020
    Publication date: January 11, 2024
    Inventors: Deqing Sun, Varun Jampani, Gengshan Yang, Daniel Vlasic, Huiwen Chang, Forrester H. Cole, Ce Liu, William Tafel Freeman
  • Patent number: 11790550
    Abstract: A method includes obtaining a first plurality of feature vectors associated with a first image and a second plurality of feature vectors associated with a second image. The method also includes generating a plurality of transformed feature vectors by transforming each respective feature vector of the first plurality of feature vectors by a kernel matrix trained to define an elliptical inner product space. The method additionally includes generating a cost volume by determining, for each respective transformed feature vector of the plurality of transformed feature vectors, a plurality of inner products, wherein each respective inner product of the plurality of inner products is between the respective transformed feature vector and a corresponding candidate feature vector of a corresponding subset of the second plurality of feature vectors. The method further includes determining, based on the cost volume, a pixel correspondence between the first image and the second image.
    Type: Grant
    Filed: July 8, 2020
    Date of Patent: October 17, 2023
    Assignee: Google LLC
    Inventors: Taihong Xiao, Deqing Sun, Ming-Hsuan Yang, Qifei Wang, Jinwei Yuan
  • Patent number: 11636668
    Abstract: A method includes filtering a point cloud transformation of a 3D object to generate a 3D lattice and processing the 3D lattice through a series of bilateral convolution networks (BCL), each BCL in the series having a lower lattice feature scale than a preceding BCL in the series. The output of each BCL in the series is concatenated to generate an intermediate 3D lattice. Further filtering of the intermediate 3D lattice generates a first prediction of features of the 3D object.
    Type: Grant
    Filed: May 22, 2018
    Date of Patent: April 25, 2023
    Inventors: Varun Jampani, Hang Su, Deqing Sun, Ming-Hsuan Yang, Jan Kautz
  • Patent number: 11508076
    Abstract: A neural network model receives color data for a sequence of images corresponding to a dynamic scene in three-dimensional (3D) space. Motion of objects in the image sequence results from a combination of a dynamic camera orientation and motion or a change in the shape of an object in the 3D space. The neural network model generates two components that are used to produce a 3D motion field representing the dynamic (non-rigid) part of the scene. The two components are information identifying dynamic and static portions of each image and the camera orientation. The dynamic portions of each image contain motion in the 3D space that is independent of the camera orientation. In other words, the motion in the 3D space (estimated 3D scene flow data) is separated from the motion of the camera.
    Type: Grant
    Filed: January 22, 2021
    Date of Patent: November 22, 2022
    Assignee: NVIDIA Corporation
    Inventors: Zhaoyang Lv, Kihwan Kim, Deqing Sun, Alejandro Jose Troccoli, Jan Kautz
  • Patent number: 11496773
    Abstract: A method, computer readable medium, and system are disclosed for identifying residual video data. This data describes data that is lost during a compression of original video data. For example, the original video data may be compressed and then decompressed, and this result may be compared to the original video data to determine the residual video data. This residual video data is transformed into a smaller format by means of encoding, binarizing, and compressing, and is sent to a destination. At the destination, the residual video data is transformed back into its original format and is used during the decompression of the compressed original video data to improve a quality of the decompressed original video data.
    Type: Grant
    Filed: June 18, 2021
    Date of Patent: November 8, 2022
    Assignee: NVIDIA CORPORATION
    Inventors: Yi-Hsuan Tsai, Ming-Yu Liu, Deqing Sun, Ming-Hsuan Yang, Jan Kautz
  • Publication number: 20220189051
    Abstract: A method includes obtaining a first plurality of feature vectors associated with a first image and a second plurality of feature vectors associated with a second image. The method also includes generating a plurality of transformed feature vectors by transforming each respective feature vector of the first plurality of feature vectors by a kernel matrix trained to define an elliptical inner product space. The method additionally includes generating a cost volume by determining, for each respective transformed feature vector of the plurality of transformed feature vectors, a plurality of inner products, wherein each respective inner product of the plurality of inner products is between the respective transformed feature vector and a corresponding candidate feature vector of a corresponding subset of the second plurality of feature vectors. The method further includes determining, based on the cost volume, a pixel correspondence between the first image and the second image.
    Type: Application
    Filed: July 8, 2020
    Publication date: June 16, 2022
    Inventors: Taihong Xiao, Deqing Sun, Ming-Hsuan Yang, Qifei Wang, Jinwei Yuan