Patents by Inventor David James FLEET

David James FLEET has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Publication number: 20260203923
    Abstract: Improved methods are provided for generating, via a noise-diffusion iterative process, depth maps or optical flow maps from input images. Also provided are improved methods for training the machine learning model(s) employed in the iterative process and for augmenting die set of training data used to train such models. By translating the depth or optical flow map prediction process into the noise diffusion context, improved performance with respect to compute cost, training data, requirements, model size, and output quality are obtained. Additionally, the noise diffusion context allows models trained as described herein to generate maps de novo from target color images and/or to begin from initial ‘guess’ maps (e.g., noisy maps, maps containing holes) when generating improved output maps, natively incorporating the imperfect prior information represented by such initial maps.
    Type: Application
    Filed: January 26, 2024
    Publication date: July 16, 2026
    Inventors: Saurabh SAXENA, Mohammad NOROUZI, David James FLEET, Abhishek KAR, Charles Irwin HERRMANN, Junhwa HUR, Deqing SUN
  • Publication number: 20260203996
    Abstract: A central problem in training NeRF models is addressed, namely, optimization in the presence of distractors, such as transient or moving objects and photometric phenomena that are not persistent throughout the capture session. Example techniques formulate training as a form of iteratively re-weighted least squares, with a variant of trimmed LS, and an inductive bias on the smoothness of the outlier process.
    Type: Application
    Filed: December 7, 2023
    Publication date: July 16, 2026
    Inventors: Daniel Christopher Duckworth, Sara Sabour Rouh Aghdam, Ivan Mikhaylovich Krasin, Andrea Tagliasacchi, David James Fleet, Suhani Deepak-Ranu Vora
  • Publication number: 20260148449
    Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating images. In one aspect, a method includes: receiving an input text prompt including a sequence of text tokens in a natural language; processing the input text prompt using a text encoder neural network to generate a set of contextual embeddings of the input text prompt; and processing the contextual embeddings through a sequence of generative neural networks to generate a final output image that depicts a scene that is described by the input text prompt.
    Type: Application
    Filed: October 24, 2025
    Publication date: May 28, 2026
    Inventors: Chitwan Saharia, William Chan, Mohammad Norouzi, Saurabh Saxena, Yi Li, Jay Ha Whang, David James Fleet, Jonathan Ho
  • Publication number: 20260134543
    Abstract: Provided are systems and methods for performing panoptic segmentation of images and videos using a denoising diffusion model. The panoptic segmentation task is formulated as a conditional discrete data generation problem. This is achieved by learning a generative model for panoptic masks, for example treated as an array of discrete tokens, conditioned on an input image. The generative model can also be applied to video data by including predictions from past frames as an additional conditioning signal. This enables the model to learn to track and segment objects automatically across video frames.
    Type: Application
    Filed: October 12, 2023
    Publication date: May 14, 2026
    Inventors: Ting Chen, Yi Li, Saurabh Saxena, Geoffrey Everest Hinton, David James Fleet
  • Patent number: 12555367
    Abstract: Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating an output video conditioned on an input. In one aspect, a method comprises receiving the input; initializing a current intermediate representation; generating an output video by updating the current intermediate representation at each of a plurality of iterations, wherein the updating comprises, at each iteration: processing an intermediate input for the iteration comprising the current intermediate representation using a diffusion model that is configured to process the intermediate input to generate a noise output; and updating the current intermediate representation using the noise output for the iteration.
    Type: Grant
    Filed: April 6, 2023
    Date of Patent: February 17, 2026
    Assignee: Google LLC
    Inventors: Jonathan Ho, Tim Salimans, Alexey Alexeevich Gritsenko, William Chan, Mohammad Norouzi, David James Fleet
  • Publication number: 20250371345
    Abstract: A method includes receiving training data comprising a plurality of pairs of images. Each pair comprises a noisy image and a denoised version of the noisy image. The method also includes training a multi-task diffusion model to perform a plurality of image-to-image translation tasks, wherein the training comprises iteratively generating a forward diffusion process by predicting, at each iteration in a sequence of iterations and based on a current noisy estimate of the denoised version of the noisy image, noise data for a next noisy estimate of the denoised version of the noisy image, updating, at each iteration, the current noisy estimate to the next noisy estimate by combining the current noisy estimate with the predicted noise data, and determining a reverse diffusion process by inverting the forward diffusion process to predict the denoised version of the noisy image. The method additionally includes providing the trained diffusion model.
    Type: Application
    Filed: August 12, 2025
    Publication date: December 4, 2025
    Inventors: Chitwan Saharia, Mohammad Norouzi, William Chan, Huiwen Chang, David James Fleet, Christopher Albert Lee, Jonathan Ho, Tim Salimans
  • Publication number: 20250363590
    Abstract: Despite recent progress, existing frame interpolation methods still struggle with extremely high resolution images and challenging cases such as repetitive textures, thin objects, and fast motion. To address these issues, provided is a cascaded diffusion frame interpolation approach that excels in these scenarios while achieving competitive performance on standard benchmarks.
    Type: Application
    Filed: May 22, 2025
    Publication date: November 27, 2025
    Inventors: Deqing Sun, Junhwa Hur, Charles Irwin Herrmann, Saurabh Saxena, David James Fleet, Janne Matias Kontkanen, Wei-Sheng Lai, Yichang Shih, Michael Rubinstein
  • Patent number: 12482160
    Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating images. In one aspect, a method includes: receiving an input text prompt including a sequence of text tokens in a natural language; processing the input text prompt using a text encoder neural network to generate a set of contextual embeddings of the input text prompt; and processing the contextual embeddings through a sequence of generative neural networks to generate a final output image that depicts a scene that is described by the input text prompt.
    Type: Grant
    Filed: April 2, 2024
    Date of Patent: November 25, 2025
    Assignee: Google LLC
    Inventors: Chitwan Saharia, William Chan, Mohammad Norouzi, Saurabh Saxena, Yi Li, Jay Ha Whang, David James Fleet, Jonathan Ho
  • Patent number: 12387096
    Abstract: A method includes receiving training data comprising a plurality of pairs of images. Each pair comprises a noisy image and a denoised version of the noisy image. The method also includes training a multi-task diffusion model to perform a plurality of image-to-image translation tasks, wherein the training comprises iteratively generating a forward diffusion process by predicting, at each iteration in a sequence of iterations and based on a current noisy estimate of the denoised version of the noisy image, noise data for a next noisy estimate of the denoised version of the noisy image, updating, at each iteration, the current noisy estimate to the next noisy estimate by combining the current noisy estimate with the predicted noise data, and determining a reverse diffusion process by inverting the forward diffusion process to predict the denoised version of the noisy image. The method additionally includes providing the trained diffusion model.
    Type: Grant
    Filed: October 5, 2022
    Date of Patent: August 12, 2025
    Assignee: Google LLC
    Inventors: Chitwan Saharia, Mohammad Norouzi, William Chan, Huiwen Chang, David James Fleet, Christopher Albert Lee, Jonathan Ho, Tim Salimans
  • Publication number: 20250218168
    Abstract: Methods, systems, and apparatus, including computer programs encoded on computer storage media, for performing multiple computer vision tasks using a shared computer vision neural network. In one aspect, one of the methods includes obtaining an input image; processing the input image and a prompt sequence using a shared computer vision neural network to generate an output sequence that comprises respective token at each of a plurality of time steps, wherein each token is selected from a shared vocabulary of tokens that is shared between the plurality of computer vision tasks, wherein the shared vocabulary comprises (i) a first set of tokens that each represent a respective discrete number from a set of discretized numbers and (ii) a second set of tokens that each represent a natural language text token.
    Type: Application
    Filed: May 19, 2023
    Publication date: July 3, 2025
    Inventors: Ting Chen, David James Fleet, Geoffrey E. Hinton, Yi Li, Saurabh Saxena, Tsung-Yi Lin
  • Publication number: 20250139959
    Abstract: Methods, systems, and apparatus, including computer programs encoded on computer storage media, for object detection using neural networks. In one aspect, one of the methods includes obtaining an input image; processing the input image using an object detection neural network to generate an output sequence that comprises respective token at each of a plurality of time steps, wherein each token is selected from a vocabulary of tokens that comprises (i) a first set of tokens that each represent a respective discrete number from a set of discretized numbers and (ii) a second set of tokens that each represent a respective object category from a set of object categories; and generating, from the tokens in the output sequence, an object detection output for the input image.
    Type: Application
    Filed: September 19, 2022
    Publication date: May 1, 2025
    Inventors: Ting Chen, Saurabh Saxena, Yi Li, Geoffrey E. Hinton, David James Fleet
  • Publication number: 20240338936
    Abstract: Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating an output video conditioned on an input. In one aspect, a method comprises receiving the input; initializing a current intermediate representation; generating an output video by updating the current intermediate representation at each of a plurality of iterations, wherein the updating comprises, at each iteration: processing an intermediate input for the iteration comprising the current intermediate representation using a diffusion model that is configured to process the intermediate input to generate a noise output; and updating the current intermediate representation using the noise output for the iteration.
    Type: Application
    Filed: April 6, 2023
    Publication date: October 10, 2024
    Inventors: Jonathan Ho, Tim Salimans, Alexey Alexeevich Gritsenko, William Chan, Mohammad Norouzi, David James Fleet
  • Publication number: 20240249456
    Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating images. In one aspect, a method includes: receiving an input text prompt including a sequence of text tokens in a natural language; processing the input text prompt using a text encoder neural network to generate a set of contextual embeddings of the input text prompt; and processing the contextual embeddings through a sequence of generative neural networks to generate a final output image that depicts a scene that is described by the input text prompt.
    Type: Application
    Filed: April 2, 2024
    Publication date: July 25, 2024
    Inventors: Chitwan Saharia, William Chan, Mohammad Norouzi, Saurabh Saxena, Yi Li, Jay Ha Whang, David James Fleet, Jonathan Ho
  • Patent number: 11978141
    Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating images. In one aspect, a method includes: receiving an input text prompt including a sequence of text tokens in a natural language; processing the input text prompt using a text encoder neural network to generate a set of contextual embeddings of the input text prompt; and processing the contextual embeddings through a sequence of generative neural networks to generate a final output image that depicts a scene that is described by the input text prompt.
    Type: Grant
    Filed: May 19, 2023
    Date of Patent: May 7, 2024
    Assignee: Google LLC
    Inventors: Chitwan Saharia, William Chan, Mohammad Norouzi, Saurabh Saxena, Yi Li, Jay Ha Whang, David James Fleet, Jonathan Ho
  • Publication number: 20230377226
    Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating images. In one aspect, a method includes: receiving an input text prompt including a sequence of text tokens in a natural language; processing the input text prompt using a text encoder neural network to generate a set of contextual embeddings of the input text prompt; and processing the contextual embeddings through a sequence of generative neural networks to generate a final output image that depicts a scene that is described by the input text prompt.
    Type: Application
    Filed: May 19, 2023
    Publication date: November 23, 2023
    Inventors: Chitwan Saharia, William Chan, Mohammad Norouzi, Saurabh Saxena, Yi Li, Jay Ha Whang, David James Fleet, Jonathan Ho
  • Publication number: 20230103638
    Abstract: A method includes receiving training data comprising a plurality of pairs of images. Each pair comprises a noisy image and a denoised version of the noisy image. The method also includes training a multi-task diffusion model to perform a plurality of image-to-image translation tasks, wherein the training comprises iteratively generating a forward diffusion process by predicting, at each iteration in a sequence of iterations and based on a current noisy estimate of the denoised version of the noisy image, noise data for a next noisy estimate of the denoised version of the noisy image, updating, at each iteration, the current noisy estimate to the next noisy estimate by combining the current noisy estimate with the predicted noise data, and determining a reverse diffusion process by inverting the forward diffusion process to predict the denoised version of the noisy image. The method additionally includes providing the trained diffusion model.
    Type: Application
    Filed: October 5, 2022
    Publication date: April 6, 2023
    Inventors: Chitwan Saharia, Mohammad Norouzi, William Chan, Huiwen Chang, David James Fleet, Christopher Albert Lee, Jonathan Ho, Tim Salimans
  • Patent number: 11515002
    Abstract: Disclosed herein are systems and methods for efficient 3D structure estimation from images of a transmissive object, including cryo-EM images. The method generally comprises, receiving a set of 2D images of a target specimen from an electron microscope, carrying out a reconstruction technique to determine a likely molecular structure, and outputting the estimated 3D structure of the specimen. The described reconstruction technique comprises: establishing a probabilistic model of the target structure; optimizing using stochastic optimization to determine which structure is most likely; and, optionally utilizing importance sampling to minimize computational burden.
    Type: Grant
    Filed: February 28, 2019
    Date of Patent: November 29, 2022
    Inventors: Marcus Anthony Brubaker, Ali Punjani, David James Fleet
  • Publication number: 20200066371
    Abstract: Disclosed herein are systems and methods for efficient 3D structure estimation from images of a transmissive object, including cryo-EM images. The method generally comprises, receiving a set of 2D images of a target specimen from an electron microscope, carrying out a reconstruction technique to determine a likely molecular structure, and outputting the estimated 3D structure of the specimen. The described reconstruction technique comprises: establishing a probabilistic model of the target structure; optimizing using stochastic optimization to determine which structure is most likely; and, optionally utilizing importance sampling to minimize computational burden.
    Type: Application
    Filed: February 28, 2019
    Publication date: February 27, 2020
    Inventors: Marcus Anthony BRUBAKER, Ali PUNJANI, David James FLEET
  • Patent number: 10282513
    Abstract: Disclosed herein are systems and methods for efficient 3D structure estimation from images of a transmissive object, including cryo-EM images. The method generally comprises, receiving a set of 2D images of a target specimen from an electron microscope, carrying out a reconstruction technique to determine a likely molecular structure, and outputting the estimated 3D structure of the specimen. The described reconstruction technique comprises: establishing a probabilistic model of the target structure; optimizing using stochastic optimization to determine which structure is most likely; and, optionally utilizing importance sampling to minimize computational burden.
    Type: Grant
    Filed: October 13, 2016
    Date of Patent: May 7, 2019
    Inventors: Marcus Anthony Brubaker, Ali Punjani, David James Fleet
  • Patent number: 10242483
    Abstract: A system and a method for image alignment between at least two images to a three-dimensional model. The method including: determining a lower bound and an upper bound of an acceptable likelihood of mismatch between the at least two images; evaluating the likelihood of mismatch between the at least two images over a set of poses (r), shifts (t), or both poses (r) and shifts (t); and discarding those evaluations resulting beyond the lower bound and upper bound.
    Type: Grant
    Filed: August 14, 2017
    Date of Patent: March 26, 2019
    Inventors: Ali Punjani, Marcus Anthony Brubaker, David James Fleet