Patents by Inventor Subhashree Radhakrishnan

Subhashree Radhakrishnan has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Publication number: 20260204274
    Abstract: Approaches presented herein include an audio generation system that incorporates a two-part generator having an encoder and decoder structure. An input mel-spectrogram is provided to the encoder to generate embeddings for the decoder to upsample and then produce one or more waveforms. During training, a discriminator may be used to evaluate the one or more waveforms to update weights of the encoder and/or the decoder. To conserve memory consumption, gradient checkpoints associated with the decoder may be deleted upon generation of the outputs and then, during backpropagation, gradients may be recomputed.
    Type: Application
    Filed: January 16, 2025
    Publication date: July 16, 2026
    Inventors: Shijia Liao, Shiyi Lan, Arun George Zachariah, Subhashree Radhakrishnan
  • Publication number: 20260187127
    Abstract: Generating a response to a query from a long input can be difficult because of the number of data units to be analyzed, where a data unit can be a video frame, an image, a page, a slide, a word count, or other data unit. The FRAG process can first reduce the number of data units to be analyzed, such as using down-sampling, temporal proximity, visual similarity, or other algorithms. The reduced number of data units can be further reduced by scoring each data unit and selecting the data units that have the highest likelihood of addressing the query, such as using a Top-K or scoring threshold parameter. The data units that satisfy the scoring algorithm can be processed by an LMM to generate a response to the query. The focusing of the data units for the LMM can improve the quality of the response that the LMM generates.
    Type: Application
    Filed: July 18, 2025
    Publication date: July 2, 2026
    Inventors: De-An Huang, Subhashree Radhakrishnan, Zhiding Yu, Jan Kautz
  • Publication number: 20260170228
    Abstract: The disclosed method for training a multimodal model includes performing one or more first operations to train a connector disposed between one or more vision encoders and a language model included in the multimodal model; performing one or more second operations to train the multimodal model using a first dataset; and performing one or more third operations to train the multimodal model using a second dataset to generate a trained multimodal model, where the second dataset is smaller than the first dataset, and where the trained multimodal model processes at least one of an input image or an input text to generate an output text.
    Type: Application
    Filed: September 8, 2025
    Publication date: June 18, 2026
    Inventors: Zhiding YU, Zhiqi LI, Guo CHEN, Shilong LIU, Shihao WANG, Vibashan VISHNUKUMAR SHARMINI, Shiyi LAN, Hao ZHANG, Yilin ZHAO, Subhashree RADHAKRISHNAN, Nai Chen CHANG, Karan SAPRA, Amala Sanjay DESHMUKH, Tuomas RINTAMAKI, Matthieu LE, De-An HUANG, Jose Manuel ALVAREZ LOPEZ, Bryan CATANZARO, Jan KAUTZ, Andrew J. TAO, Guilin LIU
  • Publication number: 20260170269
    Abstract: The disclosed method for training a multimodal model includes performing one or more first operations to train a connector disposed between one or more vision encoders and a language model included in the multimodal model; performing one or more second operations to train the multimodal model using a first dataset; and performing one or more third operations to train the multimodal model using a second dataset to generate a trained multimodal model, where the second dataset is smaller than the first dataset, and where the trained multimodal model processes at least one of an input image or an input text to generate an output text.
    Type: Application
    Filed: September 8, 2025
    Publication date: June 18, 2026
    Inventors: Zhiding YU, Zhiqi LI, Guo CHEN, Shilong LIU, Shihao WANG, Vibashan VISHNUKUMAR SHARMINI, Shiyi LAN, Hao ZHANG, Yilin ZHAO, Subhashree RADHAKRISHNAN, Nai Chen CHANG, Karan SAPRA, Amala Sanjay DESHMUKH, Tuomas RINTAMAKI, Matthieu LE, De-An HUANG, Jose Manuel ALVAREZ LOPEZ, Bryan CATANZARO, Jan KAUTZ, Andrew J. TAO, Guilin LIU
  • Publication number: 20260141696
    Abstract: Multimodal large language models (MLLMs) have evolved to interpret visual elements, progressing from text prompts for holistic image understanding to sophisticated approaches for region-level understanding. However, a key limitation of existing methods is the reliance on representations that may not consistently capture regions across frames, particularly when aiming for a unified solution for both images and videos. The present disclosure unifies image and video region-level understanding by an LLM via token marks.
    Type: Application
    Filed: June 27, 2025
    Publication date: May 21, 2026
    Inventors: Ryo Hachiuma, Min-Hung Chen, Miran Heo, De-An Huang, Sifei Liu, Subhashree Radhakrishnan, Yu-Chiang Wang
  • Publication number: 20260127901
    Abstract: Approaches presented herein may be used to generate captions using raw caption information. Raw caption information may be used, with an associated image, to generate a detailed image caption. Object lists may then be generated from the image and/or the detailed image caption to produce an image including boxing box proposals for objects within the image. One or more trained machine learning systems may then be used to generate region of interest captions that infuse the global caption context associated with the raw caption information.
    Type: Application
    Filed: November 7, 2024
    Publication date: May 7, 2026
    Inventors: Subhashree Radhakrishnan, Shijia Liao, Charul Verma, Zhiding Yu, Sifei Liu, Sean Cha
  • Patent number: 12614284
    Abstract: Class agnostic object mask generation uses a vision transformer-based auto-labeling framework requiring only images and object bounding boxes to generate object (segmentation) masks. The generated object masks, images, and object labels may then be used to train instance segmentation models or other neural networks to localize and segment objects with pixel-level accuracy.
    Type: Grant
    Filed: July 20, 2023
    Date of Patent: April 28, 2026
    Assignee: NVIDIA Corporation
    Inventors: Shiyi Lan, Zhiding Yu, Subhashree Radhakrishnan, Jose Manuel Alvarez Lopez, Animashree Anandkumar
  • Publication number: 20260099548
    Abstract: In various examples, a technique for performing conditional data sourcing and curation includes retrieving, by a plurality of processing nodes, a plurality of content items from one or more data sources, wherein each processing node retrieves a different subset of the plurality of content items from the one or more data sources. The technique also includes applying a first set of filters to metadata associated with the content items to generate a plurality of filtered content items. The technique further includes generating, based on a subset of the metadata associated with the filtered content items and a second set of filters, mappings between the filtered content items and text descriptions for the filtered content items and storing, based on the mappings, the filtered content items in association with the text descriptions in one or more data stores.
    Type: Application
    Filed: October 3, 2024
    Publication date: April 9, 2026
    Inventors: Subhashree RADHAKRISHNAN, Shijia LIAO, Charul VERMA, Farzin AGHDASI
  • Publication number: 20250384268
    Abstract: The disclosed method for training multimodal models includes performing one or more operations to train a plurality of vision language models to generate a plurality of trained vision language models, where each trained vision language model included in the plurality of trained vision language models comprises a different vision encoder and a first language model, and performing one or more operations to train a multimodal model to generate a trained multimodal model, where the trained multimodal model comprises the different vision encoders and a second language model.
    Type: Application
    Filed: April 7, 2025
    Publication date: December 18, 2025
    Inventors: Guilin LIU, Zhiding YU, Min SHI, Fuxiao LIU, Shihao WANG, Shijia LIAO, Subhashree RADHAKRISHNAN, De-An HUANG, Hongxu YIN, Karan SAPRA, Bryan CATANZARO, Andrew J. TAO, Jan KAUTZ
  • Publication number: 20250384295
    Abstract: The disclosed method for training multimodal models includes performing one or more operations to train a plurality of vision language models to generate a plurality of trained vision language models, where each trained vision language model included in the plurality of trained vision language models comprises a different vision encoder and a first language model, and performing one or more operations to train a multimodal model to generate a trained multimodal model, where the trained multimodal model comprises the different vision encoders and a second language model.
    Type: Application
    Filed: April 7, 2025
    Publication date: December 18, 2025
    Inventors: Guilin LIU, Zhiding YU, Min SHI, Fuxiao LIU, Shihao WANG, Shijia LIAO, Subhashree RADHAKRISHNAN, De-An HUANG, Hongxu YIN, Karan SAPRA, Bryan CATANZARO, Andrew J. TAO, Jan KAUTZ
  • Publication number: 20250349122
    Abstract: Embodiments of the present disclosure relate to language instructed temporal localization in videos, and provide multimodal large language models (LLMs) for performing language instructed temporal localization in video, as well as methods for training and implementing such models. In contrast to conventional systems, models according to embodiments of the present disclosure are designed to answer “when?” questions, while simultaneously improving other relevant capabilities of multimodal LLMs. Additionally, and/or alternatively, embodiments of the present disclosure may utilize a soft cross entropy loss and/or a dynamic sampling strategy to further improve the model, which allows the model to better understand temporal information and perform event localization tasks. For example, embodiments of the present disclosure may perform a dynamic sampling strategy and utilize video tokens and image tokens and/or utilize a soft cross entropy loss that applies a Gaussian distribution to the loss.
    Type: Application
    Filed: October 17, 2024
    Publication date: November 13, 2025
    Inventors: De-An Huang, Shijia Liao, Subhashree Radhakrishnan, Zaid Pervaiz Bhat, Zhiding Yu, Parthasarathy Sriram, Jan Kautz
  • Publication number: 20250191369
    Abstract: Embodiments of the present disclosure relate to language instructed temporal localization in videos, and provide multimodal large language models (LLMs) for performing language instructed temporal localization in video, as well as methods for training and implementing such models. In contrast to conventional systems, models according to embodiments of the present disclosure are designed to answer “when?” questions, while simultaneously improving other relevant capabilities of multimodal LLMs.
    Type: Application
    Filed: July 29, 2024
    Publication date: June 12, 2025
    Inventors: De-An Huang, Shijia Liao, Subhashree Radhakrishnan, Hongxu Yin, Pavlo Molchanov, Zhiding Yu, Jan Kautz
  • Publication number: 20250029409
    Abstract: Approaches are disclosed herein for an automatic segmentation labeling system that identifies objects for potential open-class categories and generates segmentation masks for objects. The disclosed system may use a training pipeline that trains two segmentation models. The training pipeline may take, as input, a set of images with bounding boxes and class labels. The set of images may be fed into a first segmentation network with the bounding boxes used as ground truth for weak supervision. The first segmentation network may be trained to generate pseudo segmentation masks. In a second stage, the trained first segmentation network is used to generate pseudo masks for a set of input images. The generated pseudo masks are provided as input, along with the corresponding images, to a second segmentation network to be used as a type of ground truth data for training the second segmentation network to generate high-quality segmentation masks.
    Type: Application
    Filed: July 18, 2023
    Publication date: January 23, 2025
    Inventors: Subhashree Radhakrishnan, Ramanathan Arunachahalam, Farzin Aghdasi, Zhiding Yu, Shiyi Lan
  • Publication number: 20250005956
    Abstract: In various examples, sensor data—such as masked sensor data—may be used as input to a machine learning model to determine a confidence for object to person associations. The masked sensor data may focus the machine learning model on particular regions of the image that correspond to persons, objects, or some combination thereof. In some embodiments, coordinates corresponding to persons, objects, or combinations thereof, in addition to area ratios between various regions of the image corresponding to the persons, objects, or combinations thereof, may be used to further aid the machine learning model in focusing on important regions of the image for determining the object to person associations.
    Type: Application
    Filed: September 9, 2024
    Publication date: January 2, 2025
    Inventors: Parthasarathy Sriram, Fnu Ratnesh Kumar, Anil Ubale, Farzin Aghdasi, Yan Zhai, Subhashree Radhakrishnan
  • Publication number: 20240386586
    Abstract: In various examples, systems and methods are disclosed relating to using neural networks for object detection or instance/semantic segmentation for, without limitation, autonomous or semi-autonomous systems and applications. In some implementations, one or more neural networks receive an image (or other sensor data representation) and a bounding shape corresponding to at least a portion of an object in the image. The bounding shape can include or be labeled with an identifier, class, and/or category of the object. The neural network can determine a mask for the object based at least on processing the image and the bounding shape. The mask can be used for various applications, such as annotating masks for vehicle or machine perception and navigation processes.
    Type: Application
    Filed: May 19, 2023
    Publication date: November 21, 2024
    Applicant: NVIDIA Corporation
    Inventors: Alperen DEGIRMENCI, Jiwoong CHOI, Zhiding YU, Ke CHEN, Shubhranshu SINGH, Yashar ASGARIEH, Subhashree RADHAKRISHNAN, James SKINNER, Jose Manuel ALVAREZ LOPEZ
  • Patent number: 12087077
    Abstract: In various examples, sensor data—such as masked sensor data—may be used as input to a machine learning model to determine a confidence for object to person associations. The masked sensor data may focus the machine learning model on particular regions of the image that correspond to persons, objects, or some combination thereof. In some embodiments, coordinates corresponding to persons, objects, or combinations thereof, in addition to area ratios between various regions of the image corresponding to the persons, objects, or combinations thereof, may be used to further aid the machine learning model in focusing on important regions of the image for determining the object to person associations.
    Type: Grant
    Filed: July 5, 2023
    Date of Patent: September 10, 2024
    Assignee: NVIDIA Corporation
    Inventors: Parthasarathy Sriram, Fnu Ratnesh Kumar, Anil Ubale, Farzin Aghdasi, Yan Zhai, Subhashree Radhakrishnan
  • Publication number: 20240221166
    Abstract: Video instance segmentation is a computer vision task that aims to detect, segment, and track objects continuously in videos. It can be used in numerous real-world applications, such as video editing, three-dimensional (3D) reconstruction, 3D navigation (e.g. for autonomous driving and/or robotics), and view point estimation. However, current machine learning-based processes employed for video instance segmentation are lacking, particularly because the densely annotated videos needed for supervised training of high-quality models are not readily available and are not easily generated. To address the issues in the prior art, the present disclosure provides point-level supervision for video instance segmentation in a manner that allows the resulting machine learning model to handle any object category.
    Type: Application
    Filed: December 22, 2023
    Publication date: July 4, 2024
    Inventors: Zhiding Yu, Shuaiyi Huang, De-An Huang, Shiyi Lan, Subhashree Radhakrishnan, Jose M. Alvarez Lopez, Anima Anandkumar
  • Publication number: 20240169545
    Abstract: Class agnostic object mask generation uses a vision transformer-based auto-labeling framework requiring only images and object bounding boxes to generate object (segmentation) masks. The generated object masks, images, and object labels may then be used to train instance segmentation models or other neural networks to localize and segment objects with pixel-level accuracy. The generated object masks may supplement or replace conventional human generated annotations. The human generated annotations may be misaligned compared with the object boundaries, resulting in poor quality labeled segmentation masks. In contrast with conventional techniques, the generated object masks are class agnostic and are automatically generated based only on a bounding box image region without relying on either labels or semantic information.
    Type: Application
    Filed: July 20, 2023
    Publication date: May 23, 2024
    Inventors: Shiyi Lan, Zhiding Yu, Subhashree Radhakrishnan, Jose Manuel Alvarez Lopez, Animashree Anandkumar
  • Patent number: 11899749
    Abstract: In various examples, training methods as described to generate a trained neural network that is robust to various environmental features. In an embodiment, training includes modifying images of a dataset and generating boundary boxes and/or other segmentation information for the modified images which is used to train a neural network.
    Type: Grant
    Filed: March 15, 2021
    Date of Patent: February 13, 2024
    Assignee: NVIDIA CORPORATION
    Inventors: Subhashree Radhakrishnan, Partha Sriram, Farzin Aghdasi, Seunghwan Cha, Zhiding Yu
  • Publication number: 20230351795
    Abstract: In various examples, sensor data—such as masked sensor data—may be used as input to a machine learning model to determine a confidence for object to person associations. The masked sensor data may focus the machine learning model on particular regions of the image that correspond to persons, objects, or some combination thereof. In some embodiments, coordinates corresponding to persons, objects, or combinations thereof, in addition to area ratios between various regions of the image corresponding to the persons, objects, or combinations thereof, may be used to further aid the machine learning model in focusing on important regions of the image for determining the object to person associations.
    Type: Application
    Filed: July 5, 2023
    Publication date: November 2, 2023
    Inventors: Parthasarathy Sriram, Fnu Ratnesh Kumar, Anil Ubale, Farzin Aghdasi, Yan Zhai, Subhashree Radhakrishnan