Patents by Inventor Subhashree Radhakrishnan
Subhashree Radhakrishnan has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Publication number: 20260204274Abstract: Approaches presented herein include an audio generation system that incorporates a two-part generator having an encoder and decoder structure. An input mel-spectrogram is provided to the encoder to generate embeddings for the decoder to upsample and then produce one or more waveforms. During training, a discriminator may be used to evaluate the one or more waveforms to update weights of the encoder and/or the decoder. To conserve memory consumption, gradient checkpoints associated with the decoder may be deleted upon generation of the outputs and then, during backpropagation, gradients may be recomputed.Type: ApplicationFiled: January 16, 2025Publication date: July 16, 2026Inventors: Shijia Liao, Shiyi Lan, Arun George Zachariah, Subhashree Radhakrishnan
-
Publication number: 20260187127Abstract: Generating a response to a query from a long input can be difficult because of the number of data units to be analyzed, where a data unit can be a video frame, an image, a page, a slide, a word count, or other data unit. The FRAG process can first reduce the number of data units to be analyzed, such as using down-sampling, temporal proximity, visual similarity, or other algorithms. The reduced number of data units can be further reduced by scoring each data unit and selecting the data units that have the highest likelihood of addressing the query, such as using a Top-K or scoring threshold parameter. The data units that satisfy the scoring algorithm can be processed by an LMM to generate a response to the query. The focusing of the data units for the LMM can improve the quality of the response that the LMM generates.Type: ApplicationFiled: July 18, 2025Publication date: July 2, 2026Inventors: De-An Huang, Subhashree Radhakrishnan, Zhiding Yu, Jan Kautz
-
Publication number: 20260170228Abstract: The disclosed method for training a multimodal model includes performing one or more first operations to train a connector disposed between one or more vision encoders and a language model included in the multimodal model; performing one or more second operations to train the multimodal model using a first dataset; and performing one or more third operations to train the multimodal model using a second dataset to generate a trained multimodal model, where the second dataset is smaller than the first dataset, and where the trained multimodal model processes at least one of an input image or an input text to generate an output text.Type: ApplicationFiled: September 8, 2025Publication date: June 18, 2026Inventors: Zhiding YU, Zhiqi LI, Guo CHEN, Shilong LIU, Shihao WANG, Vibashan VISHNUKUMAR SHARMINI, Shiyi LAN, Hao ZHANG, Yilin ZHAO, Subhashree RADHAKRISHNAN, Nai Chen CHANG, Karan SAPRA, Amala Sanjay DESHMUKH, Tuomas RINTAMAKI, Matthieu LE, De-An HUANG, Jose Manuel ALVAREZ LOPEZ, Bryan CATANZARO, Jan KAUTZ, Andrew J. TAO, Guilin LIU
-
Publication number: 20260170269Abstract: The disclosed method for training a multimodal model includes performing one or more first operations to train a connector disposed between one or more vision encoders and a language model included in the multimodal model; performing one or more second operations to train the multimodal model using a first dataset; and performing one or more third operations to train the multimodal model using a second dataset to generate a trained multimodal model, where the second dataset is smaller than the first dataset, and where the trained multimodal model processes at least one of an input image or an input text to generate an output text.Type: ApplicationFiled: September 8, 2025Publication date: June 18, 2026Inventors: Zhiding YU, Zhiqi LI, Guo CHEN, Shilong LIU, Shihao WANG, Vibashan VISHNUKUMAR SHARMINI, Shiyi LAN, Hao ZHANG, Yilin ZHAO, Subhashree RADHAKRISHNAN, Nai Chen CHANG, Karan SAPRA, Amala Sanjay DESHMUKH, Tuomas RINTAMAKI, Matthieu LE, De-An HUANG, Jose Manuel ALVAREZ LOPEZ, Bryan CATANZARO, Jan KAUTZ, Andrew J. TAO, Guilin LIU
-
Publication number: 20260141696Abstract: Multimodal large language models (MLLMs) have evolved to interpret visual elements, progressing from text prompts for holistic image understanding to sophisticated approaches for region-level understanding. However, a key limitation of existing methods is the reliance on representations that may not consistently capture regions across frames, particularly when aiming for a unified solution for both images and videos. The present disclosure unifies image and video region-level understanding by an LLM via token marks.Type: ApplicationFiled: June 27, 2025Publication date: May 21, 2026Inventors: Ryo Hachiuma, Min-Hung Chen, Miran Heo, De-An Huang, Sifei Liu, Subhashree Radhakrishnan, Yu-Chiang Wang
-
Publication number: 20260127901Abstract: Approaches presented herein may be used to generate captions using raw caption information. Raw caption information may be used, with an associated image, to generate a detailed image caption. Object lists may then be generated from the image and/or the detailed image caption to produce an image including boxing box proposals for objects within the image. One or more trained machine learning systems may then be used to generate region of interest captions that infuse the global caption context associated with the raw caption information.Type: ApplicationFiled: November 7, 2024Publication date: May 7, 2026Inventors: Subhashree Radhakrishnan, Shijia Liao, Charul Verma, Zhiding Yu, Sifei Liu, Sean Cha
-
Patent number: 12614284Abstract: Class agnostic object mask generation uses a vision transformer-based auto-labeling framework requiring only images and object bounding boxes to generate object (segmentation) masks. The generated object masks, images, and object labels may then be used to train instance segmentation models or other neural networks to localize and segment objects with pixel-level accuracy.Type: GrantFiled: July 20, 2023Date of Patent: April 28, 2026Assignee: NVIDIA CorporationInventors: Shiyi Lan, Zhiding Yu, Subhashree Radhakrishnan, Jose Manuel Alvarez Lopez, Animashree Anandkumar
-
Publication number: 20260099548Abstract: In various examples, a technique for performing conditional data sourcing and curation includes retrieving, by a plurality of processing nodes, a plurality of content items from one or more data sources, wherein each processing node retrieves a different subset of the plurality of content items from the one or more data sources. The technique also includes applying a first set of filters to metadata associated with the content items to generate a plurality of filtered content items. The technique further includes generating, based on a subset of the metadata associated with the filtered content items and a second set of filters, mappings between the filtered content items and text descriptions for the filtered content items and storing, based on the mappings, the filtered content items in association with the text descriptions in one or more data stores.Type: ApplicationFiled: October 3, 2024Publication date: April 9, 2026Inventors: Subhashree RADHAKRISHNAN, Shijia LIAO, Charul VERMA, Farzin AGHDASI
-
Publication number: 20250384268Abstract: The disclosed method for training multimodal models includes performing one or more operations to train a plurality of vision language models to generate a plurality of trained vision language models, where each trained vision language model included in the plurality of trained vision language models comprises a different vision encoder and a first language model, and performing one or more operations to train a multimodal model to generate a trained multimodal model, where the trained multimodal model comprises the different vision encoders and a second language model.Type: ApplicationFiled: April 7, 2025Publication date: December 18, 2025Inventors: Guilin LIU, Zhiding YU, Min SHI, Fuxiao LIU, Shihao WANG, Shijia LIAO, Subhashree RADHAKRISHNAN, De-An HUANG, Hongxu YIN, Karan SAPRA, Bryan CATANZARO, Andrew J. TAO, Jan KAUTZ
-
Publication number: 20250384295Abstract: The disclosed method for training multimodal models includes performing one or more operations to train a plurality of vision language models to generate a plurality of trained vision language models, where each trained vision language model included in the plurality of trained vision language models comprises a different vision encoder and a first language model, and performing one or more operations to train a multimodal model to generate a trained multimodal model, where the trained multimodal model comprises the different vision encoders and a second language model.Type: ApplicationFiled: April 7, 2025Publication date: December 18, 2025Inventors: Guilin LIU, Zhiding YU, Min SHI, Fuxiao LIU, Shihao WANG, Shijia LIAO, Subhashree RADHAKRISHNAN, De-An HUANG, Hongxu YIN, Karan SAPRA, Bryan CATANZARO, Andrew J. TAO, Jan KAUTZ
-
Publication number: 20250349122Abstract: Embodiments of the present disclosure relate to language instructed temporal localization in videos, and provide multimodal large language models (LLMs) for performing language instructed temporal localization in video, as well as methods for training and implementing such models. In contrast to conventional systems, models according to embodiments of the present disclosure are designed to answer “when?” questions, while simultaneously improving other relevant capabilities of multimodal LLMs. Additionally, and/or alternatively, embodiments of the present disclosure may utilize a soft cross entropy loss and/or a dynamic sampling strategy to further improve the model, which allows the model to better understand temporal information and perform event localization tasks. For example, embodiments of the present disclosure may perform a dynamic sampling strategy and utilize video tokens and image tokens and/or utilize a soft cross entropy loss that applies a Gaussian distribution to the loss.Type: ApplicationFiled: October 17, 2024Publication date: November 13, 2025Inventors: De-An Huang, Shijia Liao, Subhashree Radhakrishnan, Zaid Pervaiz Bhat, Zhiding Yu, Parthasarathy Sriram, Jan Kautz
-
Publication number: 20250191369Abstract: Embodiments of the present disclosure relate to language instructed temporal localization in videos, and provide multimodal large language models (LLMs) for performing language instructed temporal localization in video, as well as methods for training and implementing such models. In contrast to conventional systems, models according to embodiments of the present disclosure are designed to answer “when?” questions, while simultaneously improving other relevant capabilities of multimodal LLMs.Type: ApplicationFiled: July 29, 2024Publication date: June 12, 2025Inventors: De-An Huang, Shijia Liao, Subhashree Radhakrishnan, Hongxu Yin, Pavlo Molchanov, Zhiding Yu, Jan Kautz
-
Publication number: 20250029409Abstract: Approaches are disclosed herein for an automatic segmentation labeling system that identifies objects for potential open-class categories and generates segmentation masks for objects. The disclosed system may use a training pipeline that trains two segmentation models. The training pipeline may take, as input, a set of images with bounding boxes and class labels. The set of images may be fed into a first segmentation network with the bounding boxes used as ground truth for weak supervision. The first segmentation network may be trained to generate pseudo segmentation masks. In a second stage, the trained first segmentation network is used to generate pseudo masks for a set of input images. The generated pseudo masks are provided as input, along with the corresponding images, to a second segmentation network to be used as a type of ground truth data for training the second segmentation network to generate high-quality segmentation masks.Type: ApplicationFiled: July 18, 2023Publication date: January 23, 2025Inventors: Subhashree Radhakrishnan, Ramanathan Arunachahalam, Farzin Aghdasi, Zhiding Yu, Shiyi Lan
-
Publication number: 20250005956Abstract: In various examples, sensor data—such as masked sensor data—may be used as input to a machine learning model to determine a confidence for object to person associations. The masked sensor data may focus the machine learning model on particular regions of the image that correspond to persons, objects, or some combination thereof. In some embodiments, coordinates corresponding to persons, objects, or combinations thereof, in addition to area ratios between various regions of the image corresponding to the persons, objects, or combinations thereof, may be used to further aid the machine learning model in focusing on important regions of the image for determining the object to person associations.Type: ApplicationFiled: September 9, 2024Publication date: January 2, 2025Inventors: Parthasarathy Sriram, Fnu Ratnesh Kumar, Anil Ubale, Farzin Aghdasi, Yan Zhai, Subhashree Radhakrishnan
-
Publication number: 20240386586Abstract: In various examples, systems and methods are disclosed relating to using neural networks for object detection or instance/semantic segmentation for, without limitation, autonomous or semi-autonomous systems and applications. In some implementations, one or more neural networks receive an image (or other sensor data representation) and a bounding shape corresponding to at least a portion of an object in the image. The bounding shape can include or be labeled with an identifier, class, and/or category of the object. The neural network can determine a mask for the object based at least on processing the image and the bounding shape. The mask can be used for various applications, such as annotating masks for vehicle or machine perception and navigation processes.Type: ApplicationFiled: May 19, 2023Publication date: November 21, 2024Applicant: NVIDIA CorporationInventors: Alperen DEGIRMENCI, Jiwoong CHOI, Zhiding YU, Ke CHEN, Shubhranshu SINGH, Yashar ASGARIEH, Subhashree RADHAKRISHNAN, James SKINNER, Jose Manuel ALVAREZ LOPEZ
-
Patent number: 12087077Abstract: In various examples, sensor data—such as masked sensor data—may be used as input to a machine learning model to determine a confidence for object to person associations. The masked sensor data may focus the machine learning model on particular regions of the image that correspond to persons, objects, or some combination thereof. In some embodiments, coordinates corresponding to persons, objects, or combinations thereof, in addition to area ratios between various regions of the image corresponding to the persons, objects, or combinations thereof, may be used to further aid the machine learning model in focusing on important regions of the image for determining the object to person associations.Type: GrantFiled: July 5, 2023Date of Patent: September 10, 2024Assignee: NVIDIA CorporationInventors: Parthasarathy Sriram, Fnu Ratnesh Kumar, Anil Ubale, Farzin Aghdasi, Yan Zhai, Subhashree Radhakrishnan
-
Publication number: 20240221166Abstract: Video instance segmentation is a computer vision task that aims to detect, segment, and track objects continuously in videos. It can be used in numerous real-world applications, such as video editing, three-dimensional (3D) reconstruction, 3D navigation (e.g. for autonomous driving and/or robotics), and view point estimation. However, current machine learning-based processes employed for video instance segmentation are lacking, particularly because the densely annotated videos needed for supervised training of high-quality models are not readily available and are not easily generated. To address the issues in the prior art, the present disclosure provides point-level supervision for video instance segmentation in a manner that allows the resulting machine learning model to handle any object category.Type: ApplicationFiled: December 22, 2023Publication date: July 4, 2024Inventors: Zhiding Yu, Shuaiyi Huang, De-An Huang, Shiyi Lan, Subhashree Radhakrishnan, Jose M. Alvarez Lopez, Anima Anandkumar
-
Publication number: 20240169545Abstract: Class agnostic object mask generation uses a vision transformer-based auto-labeling framework requiring only images and object bounding boxes to generate object (segmentation) masks. The generated object masks, images, and object labels may then be used to train instance segmentation models or other neural networks to localize and segment objects with pixel-level accuracy. The generated object masks may supplement or replace conventional human generated annotations. The human generated annotations may be misaligned compared with the object boundaries, resulting in poor quality labeled segmentation masks. In contrast with conventional techniques, the generated object masks are class agnostic and are automatically generated based only on a bounding box image region without relying on either labels or semantic information.Type: ApplicationFiled: July 20, 2023Publication date: May 23, 2024Inventors: Shiyi Lan, Zhiding Yu, Subhashree Radhakrishnan, Jose Manuel Alvarez Lopez, Animashree Anandkumar
-
Patent number: 11899749Abstract: In various examples, training methods as described to generate a trained neural network that is robust to various environmental features. In an embodiment, training includes modifying images of a dataset and generating boundary boxes and/or other segmentation information for the modified images which is used to train a neural network.Type: GrantFiled: March 15, 2021Date of Patent: February 13, 2024Assignee: NVIDIA CORPORATIONInventors: Subhashree Radhakrishnan, Partha Sriram, Farzin Aghdasi, Seunghwan Cha, Zhiding Yu
-
Publication number: 20230351795Abstract: In various examples, sensor data—such as masked sensor data—may be used as input to a machine learning model to determine a confidence for object to person associations. The masked sensor data may focus the machine learning model on particular regions of the image that correspond to persons, objects, or some combination thereof. In some embodiments, coordinates corresponding to persons, objects, or combinations thereof, in addition to area ratios between various regions of the image corresponding to the persons, objects, or combinations thereof, may be used to further aid the machine learning model in focusing on important regions of the image for determining the object to person associations.Type: ApplicationFiled: July 5, 2023Publication date: November 2, 2023Inventors: Parthasarathy Sriram, Fnu Ratnesh Kumar, Anil Ubale, Farzin Aghdasi, Yan Zhai, Subhashree Radhakrishnan