Patents by Inventor David James FLEET
David James FLEET has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Publication number: 20260203923Abstract: Improved methods are provided for generating, via a noise-diffusion iterative process, depth maps or optical flow maps from input images. Also provided are improved methods for training the machine learning model(s) employed in the iterative process and for augmenting die set of training data used to train such models. By translating the depth or optical flow map prediction process into the noise diffusion context, improved performance with respect to compute cost, training data, requirements, model size, and output quality are obtained. Additionally, the noise diffusion context allows models trained as described herein to generate maps de novo from target color images and/or to begin from initial ‘guess’ maps (e.g., noisy maps, maps containing holes) when generating improved output maps, natively incorporating the imperfect prior information represented by such initial maps.Type: ApplicationFiled: January 26, 2024Publication date: July 16, 2026Inventors: Saurabh SAXENA, Mohammad NOROUZI, David James FLEET, Abhishek KAR, Charles Irwin HERRMANN, Junhwa HUR, Deqing SUN
-
Publication number: 20260203996Abstract: A central problem in training NeRF models is addressed, namely, optimization in the presence of distractors, such as transient or moving objects and photometric phenomena that are not persistent throughout the capture session. Example techniques formulate training as a form of iteratively re-weighted least squares, with a variant of trimmed LS, and an inductive bias on the smoothness of the outlier process.Type: ApplicationFiled: December 7, 2023Publication date: July 16, 2026Inventors: Daniel Christopher Duckworth, Sara Sabour Rouh Aghdam, Ivan Mikhaylovich Krasin, Andrea Tagliasacchi, David James Fleet, Suhani Deepak-Ranu Vora
-
Publication number: 20260148449Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating images. In one aspect, a method includes: receiving an input text prompt including a sequence of text tokens in a natural language; processing the input text prompt using a text encoder neural network to generate a set of contextual embeddings of the input text prompt; and processing the contextual embeddings through a sequence of generative neural networks to generate a final output image that depicts a scene that is described by the input text prompt.Type: ApplicationFiled: October 24, 2025Publication date: May 28, 2026Inventors: Chitwan Saharia, William Chan, Mohammad Norouzi, Saurabh Saxena, Yi Li, Jay Ha Whang, David James Fleet, Jonathan Ho
-
Publication number: 20260134543Abstract: Provided are systems and methods for performing panoptic segmentation of images and videos using a denoising diffusion model. The panoptic segmentation task is formulated as a conditional discrete data generation problem. This is achieved by learning a generative model for panoptic masks, for example treated as an array of discrete tokens, conditioned on an input image. The generative model can also be applied to video data by including predictions from past frames as an additional conditioning signal. This enables the model to learn to track and segment objects automatically across video frames.Type: ApplicationFiled: October 12, 2023Publication date: May 14, 2026Inventors: Ting Chen, Yi Li, Saurabh Saxena, Geoffrey Everest Hinton, David James Fleet
-
Patent number: 12555367Abstract: Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating an output video conditioned on an input. In one aspect, a method comprises receiving the input; initializing a current intermediate representation; generating an output video by updating the current intermediate representation at each of a plurality of iterations, wherein the updating comprises, at each iteration: processing an intermediate input for the iteration comprising the current intermediate representation using a diffusion model that is configured to process the intermediate input to generate a noise output; and updating the current intermediate representation using the noise output for the iteration.Type: GrantFiled: April 6, 2023Date of Patent: February 17, 2026Assignee: Google LLCInventors: Jonathan Ho, Tim Salimans, Alexey Alexeevich Gritsenko, William Chan, Mohammad Norouzi, David James Fleet
-
Publication number: 20250371345Abstract: A method includes receiving training data comprising a plurality of pairs of images. Each pair comprises a noisy image and a denoised version of the noisy image. The method also includes training a multi-task diffusion model to perform a plurality of image-to-image translation tasks, wherein the training comprises iteratively generating a forward diffusion process by predicting, at each iteration in a sequence of iterations and based on a current noisy estimate of the denoised version of the noisy image, noise data for a next noisy estimate of the denoised version of the noisy image, updating, at each iteration, the current noisy estimate to the next noisy estimate by combining the current noisy estimate with the predicted noise data, and determining a reverse diffusion process by inverting the forward diffusion process to predict the denoised version of the noisy image. The method additionally includes providing the trained diffusion model.Type: ApplicationFiled: August 12, 2025Publication date: December 4, 2025Inventors: Chitwan Saharia, Mohammad Norouzi, William Chan, Huiwen Chang, David James Fleet, Christopher Albert Lee, Jonathan Ho, Tim Salimans
-
Publication number: 20250363590Abstract: Despite recent progress, existing frame interpolation methods still struggle with extremely high resolution images and challenging cases such as repetitive textures, thin objects, and fast motion. To address these issues, provided is a cascaded diffusion frame interpolation approach that excels in these scenarios while achieving competitive performance on standard benchmarks.Type: ApplicationFiled: May 22, 2025Publication date: November 27, 2025Inventors: Deqing Sun, Junhwa Hur, Charles Irwin Herrmann, Saurabh Saxena, David James Fleet, Janne Matias Kontkanen, Wei-Sheng Lai, Yichang Shih, Michael Rubinstein
-
Patent number: 12482160Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating images. In one aspect, a method includes: receiving an input text prompt including a sequence of text tokens in a natural language; processing the input text prompt using a text encoder neural network to generate a set of contextual embeddings of the input text prompt; and processing the contextual embeddings through a sequence of generative neural networks to generate a final output image that depicts a scene that is described by the input text prompt.Type: GrantFiled: April 2, 2024Date of Patent: November 25, 2025Assignee: Google LLCInventors: Chitwan Saharia, William Chan, Mohammad Norouzi, Saurabh Saxena, Yi Li, Jay Ha Whang, David James Fleet, Jonathan Ho
-
Patent number: 12387096Abstract: A method includes receiving training data comprising a plurality of pairs of images. Each pair comprises a noisy image and a denoised version of the noisy image. The method also includes training a multi-task diffusion model to perform a plurality of image-to-image translation tasks, wherein the training comprises iteratively generating a forward diffusion process by predicting, at each iteration in a sequence of iterations and based on a current noisy estimate of the denoised version of the noisy image, noise data for a next noisy estimate of the denoised version of the noisy image, updating, at each iteration, the current noisy estimate to the next noisy estimate by combining the current noisy estimate with the predicted noise data, and determining a reverse diffusion process by inverting the forward diffusion process to predict the denoised version of the noisy image. The method additionally includes providing the trained diffusion model.Type: GrantFiled: October 5, 2022Date of Patent: August 12, 2025Assignee: Google LLCInventors: Chitwan Saharia, Mohammad Norouzi, William Chan, Huiwen Chang, David James Fleet, Christopher Albert Lee, Jonathan Ho, Tim Salimans
-
Publication number: 20250218168Abstract: Methods, systems, and apparatus, including computer programs encoded on computer storage media, for performing multiple computer vision tasks using a shared computer vision neural network. In one aspect, one of the methods includes obtaining an input image; processing the input image and a prompt sequence using a shared computer vision neural network to generate an output sequence that comprises respective token at each of a plurality of time steps, wherein each token is selected from a shared vocabulary of tokens that is shared between the plurality of computer vision tasks, wherein the shared vocabulary comprises (i) a first set of tokens that each represent a respective discrete number from a set of discretized numbers and (ii) a second set of tokens that each represent a natural language text token.Type: ApplicationFiled: May 19, 2023Publication date: July 3, 2025Inventors: Ting Chen, David James Fleet, Geoffrey E. Hinton, Yi Li, Saurabh Saxena, Tsung-Yi Lin
-
Publication number: 20250139959Abstract: Methods, systems, and apparatus, including computer programs encoded on computer storage media, for object detection using neural networks. In one aspect, one of the methods includes obtaining an input image; processing the input image using an object detection neural network to generate an output sequence that comprises respective token at each of a plurality of time steps, wherein each token is selected from a vocabulary of tokens that comprises (i) a first set of tokens that each represent a respective discrete number from a set of discretized numbers and (ii) a second set of tokens that each represent a respective object category from a set of object categories; and generating, from the tokens in the output sequence, an object detection output for the input image.Type: ApplicationFiled: September 19, 2022Publication date: May 1, 2025Inventors: Ting Chen, Saurabh Saxena, Yi Li, Geoffrey E. Hinton, David James Fleet
-
Publication number: 20240338936Abstract: Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating an output video conditioned on an input. In one aspect, a method comprises receiving the input; initializing a current intermediate representation; generating an output video by updating the current intermediate representation at each of a plurality of iterations, wherein the updating comprises, at each iteration: processing an intermediate input for the iteration comprising the current intermediate representation using a diffusion model that is configured to process the intermediate input to generate a noise output; and updating the current intermediate representation using the noise output for the iteration.Type: ApplicationFiled: April 6, 2023Publication date: October 10, 2024Inventors: Jonathan Ho, Tim Salimans, Alexey Alexeevich Gritsenko, William Chan, Mohammad Norouzi, David James Fleet
-
Publication number: 20240249456Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating images. In one aspect, a method includes: receiving an input text prompt including a sequence of text tokens in a natural language; processing the input text prompt using a text encoder neural network to generate a set of contextual embeddings of the input text prompt; and processing the contextual embeddings through a sequence of generative neural networks to generate a final output image that depicts a scene that is described by the input text prompt.Type: ApplicationFiled: April 2, 2024Publication date: July 25, 2024Inventors: Chitwan Saharia, William Chan, Mohammad Norouzi, Saurabh Saxena, Yi Li, Jay Ha Whang, David James Fleet, Jonathan Ho
-
Patent number: 11978141Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating images. In one aspect, a method includes: receiving an input text prompt including a sequence of text tokens in a natural language; processing the input text prompt using a text encoder neural network to generate a set of contextual embeddings of the input text prompt; and processing the contextual embeddings through a sequence of generative neural networks to generate a final output image that depicts a scene that is described by the input text prompt.Type: GrantFiled: May 19, 2023Date of Patent: May 7, 2024Assignee: Google LLCInventors: Chitwan Saharia, William Chan, Mohammad Norouzi, Saurabh Saxena, Yi Li, Jay Ha Whang, David James Fleet, Jonathan Ho
-
Publication number: 20230377226Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating images. In one aspect, a method includes: receiving an input text prompt including a sequence of text tokens in a natural language; processing the input text prompt using a text encoder neural network to generate a set of contextual embeddings of the input text prompt; and processing the contextual embeddings through a sequence of generative neural networks to generate a final output image that depicts a scene that is described by the input text prompt.Type: ApplicationFiled: May 19, 2023Publication date: November 23, 2023Inventors: Chitwan Saharia, William Chan, Mohammad Norouzi, Saurabh Saxena, Yi Li, Jay Ha Whang, David James Fleet, Jonathan Ho
-
Publication number: 20230103638Abstract: A method includes receiving training data comprising a plurality of pairs of images. Each pair comprises a noisy image and a denoised version of the noisy image. The method also includes training a multi-task diffusion model to perform a plurality of image-to-image translation tasks, wherein the training comprises iteratively generating a forward diffusion process by predicting, at each iteration in a sequence of iterations and based on a current noisy estimate of the denoised version of the noisy image, noise data for a next noisy estimate of the denoised version of the noisy image, updating, at each iteration, the current noisy estimate to the next noisy estimate by combining the current noisy estimate with the predicted noise data, and determining a reverse diffusion process by inverting the forward diffusion process to predict the denoised version of the noisy image. The method additionally includes providing the trained diffusion model.Type: ApplicationFiled: October 5, 2022Publication date: April 6, 2023Inventors: Chitwan Saharia, Mohammad Norouzi, William Chan, Huiwen Chang, David James Fleet, Christopher Albert Lee, Jonathan Ho, Tim Salimans
-
Patent number: 11515002Abstract: Disclosed herein are systems and methods for efficient 3D structure estimation from images of a transmissive object, including cryo-EM images. The method generally comprises, receiving a set of 2D images of a target specimen from an electron microscope, carrying out a reconstruction technique to determine a likely molecular structure, and outputting the estimated 3D structure of the specimen. The described reconstruction technique comprises: establishing a probabilistic model of the target structure; optimizing using stochastic optimization to determine which structure is most likely; and, optionally utilizing importance sampling to minimize computational burden.Type: GrantFiled: February 28, 2019Date of Patent: November 29, 2022Inventors: Marcus Anthony Brubaker, Ali Punjani, David James Fleet
-
Publication number: 20200066371Abstract: Disclosed herein are systems and methods for efficient 3D structure estimation from images of a transmissive object, including cryo-EM images. The method generally comprises, receiving a set of 2D images of a target specimen from an electron microscope, carrying out a reconstruction technique to determine a likely molecular structure, and outputting the estimated 3D structure of the specimen. The described reconstruction technique comprises: establishing a probabilistic model of the target structure; optimizing using stochastic optimization to determine which structure is most likely; and, optionally utilizing importance sampling to minimize computational burden.Type: ApplicationFiled: February 28, 2019Publication date: February 27, 2020Inventors: Marcus Anthony BRUBAKER, Ali PUNJANI, David James FLEET
-
Patent number: 10282513Abstract: Disclosed herein are systems and methods for efficient 3D structure estimation from images of a transmissive object, including cryo-EM images. The method generally comprises, receiving a set of 2D images of a target specimen from an electron microscope, carrying out a reconstruction technique to determine a likely molecular structure, and outputting the estimated 3D structure of the specimen. The described reconstruction technique comprises: establishing a probabilistic model of the target structure; optimizing using stochastic optimization to determine which structure is most likely; and, optionally utilizing importance sampling to minimize computational burden.Type: GrantFiled: October 13, 2016Date of Patent: May 7, 2019Inventors: Marcus Anthony Brubaker, Ali Punjani, David James Fleet
-
Patent number: 10242483Abstract: A system and a method for image alignment between at least two images to a three-dimensional model. The method including: determining a lower bound and an upper bound of an acceptable likelihood of mismatch between the at least two images; evaluating the likelihood of mismatch between the at least two images over a set of poses (r), shifts (t), or both poses (r) and shifts (t); and discarding those evaluations resulting beyond the lower bound and upper bound.Type: GrantFiled: August 14, 2017Date of Patent: March 26, 2019Inventors: Ali Punjani, Marcus Anthony Brubaker, David James Fleet