Patents by Inventor Bryan Catanzaro

Bryan Catanzaro has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Publication number: 20260205558
    Abstract: Apparatuses, systems, and techniques to enhance video are disclosed. In at least one embodiment, one or more neural networks are used to create, from a first video, a second video having one or more additional video frames.
    Type: Application
    Filed: March 6, 2026
    Publication date: July 16, 2026
    Inventors: Kevin Shih, Aysegul Dundar, Animesh Garg, Robert Pottorff, Andrew Tao, Bryan Catanzaro
  • Patent number: 12670895
    Abstract: Systems and methods to help synthesize a second audio signal based, at least in part, on one or more neural networks trained using one or more characteristics of a first audio signal. Systems and methods to train one or more neural networks to synthesize a second audio signal based, at least in part, on one or more characteristics of a first audio signal.
    Type: Grant
    Filed: June 12, 2019
    Date of Patent: June 30, 2026
    Assignee: NVIDIA Corporation
    Inventors: Ryan Prenger, Rafael Valle, Bryan Catanzaro
  • Publication number: 20260179180
    Abstract: Apparatuses, systems, and techniques are presented to generate images with one or more visual effects applied. In at least one embodiment, one or more visual effects are applied to one or more images having a resolution that is less than a first resolution and those visual effects approximated for one or more images having a resolution that is greater than or equal to the first resolution.
    Type: Application
    Filed: July 17, 2025
    Publication date: June 25, 2026
    Inventors: Robert Pottorff, David Tarjan, Andrew Tao, Bryan Catanzaro
  • Publication number: 20260170228
    Abstract: The disclosed method for training a multimodal model includes performing one or more first operations to train a connector disposed between one or more vision encoders and a language model included in the multimodal model; performing one or more second operations to train the multimodal model using a first dataset; and performing one or more third operations to train the multimodal model using a second dataset to generate a trained multimodal model, where the second dataset is smaller than the first dataset, and where the trained multimodal model processes at least one of an input image or an input text to generate an output text.
    Type: Application
    Filed: September 8, 2025
    Publication date: June 18, 2026
    Inventors: Zhiding YU, Zhiqi LI, Guo CHEN, Shilong LIU, Shihao WANG, Vibashan VISHNUKUMAR SHARMINI, Shiyi LAN, Hao ZHANG, Yilin ZHAO, Subhashree RADHAKRISHNAN, Nai Chen CHANG, Karan SAPRA, Amala Sanjay DESHMUKH, Tuomas RINTAMAKI, Matthieu LE, De-An HUANG, Jose Manuel ALVAREZ LOPEZ, Bryan CATANZARO, Jan KAUTZ, Andrew J. TAO, Guilin LIU
  • Publication number: 20260170269
    Abstract: The disclosed method for training a multimodal model includes performing one or more first operations to train a connector disposed between one or more vision encoders and a language model included in the multimodal model; performing one or more second operations to train the multimodal model using a first dataset; and performing one or more third operations to train the multimodal model using a second dataset to generate a trained multimodal model, where the second dataset is smaller than the first dataset, and where the trained multimodal model processes at least one of an input image or an input text to generate an output text.
    Type: Application
    Filed: September 8, 2025
    Publication date: June 18, 2026
    Inventors: Zhiding YU, Zhiqi LI, Guo CHEN, Shilong LIU, Shihao WANG, Vibashan VISHNUKUMAR SHARMINI, Shiyi LAN, Hao ZHANG, Yilin ZHAO, Subhashree RADHAKRISHNAN, Nai Chen CHANG, Karan SAPRA, Amala Sanjay DESHMUKH, Tuomas RINTAMAKI, Matthieu LE, De-An HUANG, Jose Manuel ALVAREZ LOPEZ, Bryan CATANZARO, Jan KAUTZ, Andrew J. TAO, Guilin LIU
  • Publication number: 20260154862
    Abstract: Apparatuses, systems, and techniques are presented to generate one or more images. In at least one embodiment, two or more pixels from two or more images are blended based, at least in part, on a distance of the two or more pixels from a region of the two or more images, in which pixel colors are substantially similar.
    Type: Application
    Filed: August 18, 2025
    Publication date: June 4, 2026
    Inventors: Robert Pottorff, Karan Sapra, Andrew Tao, Bryan Catanzaro, Jarmo Lunden
  • Publication number: 20260134187
    Abstract: Apparatuses, systems, and techniques for designing a data path circuit such as a parallel prefix circuit with reinforcement learning are described. A method can include receiving a first design state of a data path circuit, inputting the first design state of the data path circuit into a machine learning model, and performing reinforcement learning using the machine learning model to output a final design state of the data path circuit, wherein the final design state of the data path circuit has decreased area, power consumption and/or delay as compared to conventionally designed data path circuits.
    Type: Application
    Filed: August 8, 2025
    Publication date: May 14, 2026
    Inventors: Rajarshi Roy, Saad Godil, Jonathan Raiman, Neel Kant, Ilyas Elkin, Ming Y. Siu, Robert Kirby, Stuart Oberman, Bryan Catanzaro
  • Publication number: 20260093459
    Abstract: Disclosed are apparatuses, systems, and techniques for automated iterative generation and debugging of a computer code (CC) using a language model (LM). The techniques include causing the LM to perform, responsive to a task prompt, iterative generation of the CC, an individual iteration causing the LM to (i) produce multiple evaluations of a previous faulty version of the CC, (ii) generate, responsive to the multiple evaluations, multiple modified versions of the CC, and (iii) automatically select, as an output of the individual iteration, a best performing, in view of one or more tests, version of the CC from the multiple modified versions of the CC.
    Type: Application
    Filed: September 27, 2024
    Publication date: April 2, 2026
    Inventors: Jialin Song, Jonathan Raiman, Bryan Catanzaro
  • Publication number: 20260080858
    Abstract: Disclosed are apparatuses, systems, and techniques that may use machine learning for implementing generative text-to-speech models. The techniques include identifying a mapping of speech characteristics (SC) on a target distribution of a latent variable using a non-linear transformation for at least a subset of the SC. Parameters of the non-linear transformation are determined using a neural network that approximates a statistics of the SC with a statistics predicted for the SC based on the identified mapping and the target distribution of the latent variable.
    Type: Application
    Filed: November 21, 2025
    Publication date: March 19, 2026
    Inventors: Kevin Shih, José Rafael Valle Gomes da Costa, Rohan Badlani, João Felipe Santos, Bryan Catanzaro
  • Patent number: 12574471
    Abstract: Apparatuses, systems, and techniques to enhance video are disclosed. In at least one embodiment, one or more neural networks are used to create, from a first video, a second video having one or more additional video frames.
    Type: Grant
    Filed: January 23, 2024
    Date of Patent: March 10, 2026
    Assignee: NVIDIA Corporation
    Inventors: Kevin Shih, Aysegul Dundar, Animesh Garg, Robert Pottorff, Andrew Tao, Bryan Catanzaro
  • Patent number: 12573000
    Abstract: Apparatuses, systems, and techniques are presented to generate images. In at least one embodiment, one or more neural networks are used to generate one or more images using one or more pixel weights determined based, at least in part, on one or more sub-pixel offset values.
    Type: Grant
    Filed: October 8, 2020
    Date of Patent: March 10, 2026
    Assignee: NVIDIA Corporation
    Inventors: Shiqiu Liu, Robert Pottorff, Guilin Liu, Karan Sapra, Jon Barker, David Tarjan, Pekka Janis, Edvard Fagerholm, Lei Yang, Kevin Shih, Marco Salvi, Timo Roman, Andrew Tao, Bryan Catanzaro
  • Publication number: 20260052256
    Abstract: A neural network architecture is disclosed for performing image in-painting using partial convolution operations. The neural network processes an image and a corresponding mask that identifies holes in the image utilizing partial convolution operations, where the mask is used by the partial convolution operation to zero out coefficients of the convolution kernel corresponding to invalid pixel data for the holes. The mask is updated after each partial convolution operation is performed in an encoder section of the neural network. In one embodiment, the neural network is implemented using an encoder-decoder framework with skip links to forward representations of the features at different sections of the encoder to corresponding sections of the decoder.
    Type: Application
    Filed: August 22, 2025
    Publication date: February 19, 2026
    Inventors: Guilin Liu, Fitsum A. Reda, Kevin Shih, Ting-Chun Wang, Andrew Tao, Bryan Catanzaro
  • Patent number: 12555186
    Abstract: Apparatuses, systems, and techniques are presented to generate images. In at least one embodiment, one or more neural networks are used to generate one or more images using one or more pixel weights determined based, at least in part, on one or more sub-pixel offset values.
    Type: Grant
    Filed: February 10, 2021
    Date of Patent: February 17, 2026
    Assignee: NVIDIA Corporation
    Inventors: Shiqiu Liu, Robert Pottorff, Guilin Liu, Karan Sapra, Jon Barker, David Tarjan, Pekka Janis, Edvard Fagerholm, Lei Yang, Kevin Shih, Marco Salvi, Timo Roman, Andrew Tao, Bryan Catanzaro
  • Patent number: 12548113
    Abstract: Apparatuses, systems, and techniques are presented to generate images with one or more visual effects applied. In at least one embodiment, one or more visual effects are applied to one or more images having a resolution that is less than a first resolution and those visual effects approximated for one or more images having a resolution that is greater than or equal to the first resolution.
    Type: Grant
    Filed: September 3, 2020
    Date of Patent: February 10, 2026
    Assignee: NVIDIA Corporation
    Inventors: Robert Pottorff, David Tarjan, Andrew Tao, Bryan Catanzaro
  • Publication number: 20250384268
    Abstract: The disclosed method for training multimodal models includes performing one or more operations to train a plurality of vision language models to generate a plurality of trained vision language models, where each trained vision language model included in the plurality of trained vision language models comprises a different vision encoder and a first language model, and performing one or more operations to train a multimodal model to generate a trained multimodal model, where the trained multimodal model comprises the different vision encoders and a second language model.
    Type: Application
    Filed: April 7, 2025
    Publication date: December 18, 2025
    Inventors: Guilin LIU, Zhiding YU, Min SHI, Fuxiao LIU, Shihao WANG, Shijia LIAO, Subhashree RADHAKRISHNAN, De-An HUANG, Hongxu YIN, Karan SAPRA, Bryan CATANZARO, Andrew J. TAO, Jan KAUTZ
  • Publication number: 20250384295
    Abstract: The disclosed method for training multimodal models includes performing one or more operations to train a plurality of vision language models to generate a plurality of trained vision language models, where each trained vision language model included in the plurality of trained vision language models comprises a different vision encoder and a first language model, and performing one or more operations to train a multimodal model to generate a trained multimodal model, where the trained multimodal model comprises the different vision encoders and a second language model.
    Type: Application
    Filed: April 7, 2025
    Publication date: December 18, 2025
    Inventors: Guilin LIU, Zhiding YU, Min SHI, Fuxiao LIU, Shihao WANG, Shijia LIAO, Subhashree RADHAKRISHNAN, De-An HUANG, Hongxu YIN, Karan SAPRA, Bryan CATANZARO, Andrew J. TAO, Jan KAUTZ
  • Patent number: 12488778
    Abstract: Disclosed are apparatuses, systems, and techniques that may use machine learning for implementing generative text-to-speech models. The techniques include identifying a mapping of speech characteristics (SC) on a target distribution of a latent variable using a non-linear transformation for at least a subset of the SC. Parameters of the non-linear transformation are determined using a neural network that approximates a statistics of the SC with a statistics predicted for the SC based on the identified mapping and the target distribution of the latent variable.
    Type: Grant
    Filed: January 20, 2023
    Date of Patent: December 2, 2025
    Assignee: NVIDIA Corporation
    Inventors: Kevin Shih, José Rafael Valle Gomes da Costa, Rohan Badlani, João Felipe Santos, Bryan Catanzaro
  • Patent number: 12425605
    Abstract: A neural network architecture is disclosed for performing image in-painting using partial convolution operations. The neural network processes an image and a corresponding mask that identifies holes in the image utilizing partial convolution operations, where the mask is used by the partial convolution operation to zero out coefficients of the convolution kernel corresponding to invalid pixel data for the holes. The mask is updated after each partial convolution operation is performed in an encoder section of the neural network. In one embodiment, the neural network is implemented using an encoder-decoder framework with skip links to forward representations of the features at different sections of the encoder to corresponding sections of the decoder.
    Type: Grant
    Filed: March 21, 2019
    Date of Patent: September 23, 2025
    Assignee: NVIDIA Corporation
    Inventors: Guilin Liu, Fitsum A. Reda, Kevin Shih, Ting-Chun Wang, Andrew Tao, Bryan Catanzaro
  • Patent number: 12400291
    Abstract: Apparatuses, systems, and techniques are presented to generate images with one or more visual effects applied. In at least one embodiment, one or more visual effects are applied to one or more images having a resolution that is less than a first resolution and those visual effects approximated for one or more images having a resolution that is greater than or equal to the first resolution.
    Type: Grant
    Filed: August 19, 2021
    Date of Patent: August 26, 2025
    Assignee: NVIDIA CORPORATION
    Inventors: Robert Pottorff, David Tarjan, Andrew Tao, Bryan Catanzaro
  • Patent number: 12394113
    Abstract: Apparatuses, systems, and techniques are presented to generate one or more images. In at least one embodiment, two or more pixels from two or more images are blended based, at least in part, on a distance of the two or more pixels from a region of the two or more images, in which pixel colors are substantially similar.
    Type: Grant
    Filed: June 18, 2021
    Date of Patent: August 19, 2025
    Assignee: NVIDIA Corporation
    Inventors: Robert Pottorff, Karan Sapra, Andrew Tao, Bryan Catanzaro, Jarmo Lunden