Patents by Inventor Utkarsh Vaidya

Utkarsh Vaidya has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Publication number: 20260204261
    Abstract: Disclosed are apparatuses, systems, and techniques for generating multi-format transcriptions of speech. The techniques include processing, using an encoder, audio frames representative of a speech to generate embeddings encoding the speech and processing, using multiple decoders, the embeddings to generate multiple transcriptions of the speech. An individual transcription is generated by a respective decoder and conforms to a respective text format that differs from other text formats in capitalization, punctuation, use of non-alphabet characters, and/or identification of individual utterances of the speech.
    Type: Application
    Filed: January 16, 2025
    Publication date: July 16, 2026
    Inventors: Harishchandra Dubey, Myungjong Kim, Oluwatobi Olabiyi, Utkarsh Vaidya
  • Publication number: 20260162654
    Abstract: A textual transcript and one or more language indicators are determined using a multilingual speech-to-text (STT) model of a multilingual automatic speech recognition (ASR) system and using an audio sample as input to the multilingual STT model. The textual transcript is associated with the audio sample, and the one or more language indicators are each associated with a respective grammatical unit of one or more grammatical units of the textual transcript. A monolingual language model (LM) of a plurality of monolingual LMs of the ASR system is identified using a language indicator of the one or more language indicators. The textual transcript associated with the audio sample is caused to be refined using the identified LM and using a subset of the textual transcript as input to the identified LM.
    Type: Application
    Filed: December 9, 2024
    Publication date: June 11, 2026
    Inventors: Myungjong Kim, Mayank Jain, Yitagessu Gebremedhin, Utkarsh Vaidya, Oluwatobi Olabiyi
  • Publication number: 20260154514
    Abstract: In various examples, techniques are described for adapting a multilingual Large Language Model (LLM) into a bilingual Small Language Model (SLM) that exhibits model capacity to understand, process, and generate content in both English and a Low-Resource Language (LRL). The techniques include compressing the LLM to generate a multilingual SLM and performing continued pre-training on the multilingual SLM to generate the bilingual SLM. The techniques also include performing one or more alignment techniques on the bilingual SLM to adapt the SLM's outputs to human values and expectations regarding, e.g., profanity, privacy, politeness, bias, and/or conversational style. The techniques may generate various training corpora, each including one or more of natural English content, natural LRL content, synthetic LRL content generated via translation from English sources, and transliterated synthetic LRL content based on transliterations of natural and/or synthetic LRL content.
    Type: Application
    Filed: September 30, 2025
    Publication date: June 4, 2026
    Inventors: Raviraj JOSHI, Kanishk SINGLA, Anusha KAMATH, Raunak KALANI, Utkarsh VAIDYA, Sanjay Singh CHAUHAN, Niranjan WARTIKAR, Eileen Margaret Peters LONG
  • Publication number: 20260064994
    Abstract: Approaches presented herein provide for the generation of relatively small language models that are optimized for target languages. A multilingual large language model (LLM) can be reduced in size using a process such as language-aware pruning, where individual network parameters have importance scores calculated with respect to the target language and then an appropriate number of lower-importance score parameters are removed from the network. Continued pretraining can be performed using a set of training data including real and/or synthesized text in the target language, to obtain a high performing language model with a limited number of parameters optimized for a target language, as may correspond to a lower-resource language that may otherwise not have enough training data available to sufficiently train a language model from scratch.
    Type: Application
    Filed: August 29, 2024
    Publication date: March 5, 2026
    Inventors: Raviraj Joshi, Utkarsh Vaidya, Sanjay Singh Chauhan
  • Publication number: 20260057884
    Abstract: In various examples, generating unified text using speech recognition models for AI systems and applications is described herein. Systems and methods are disclosed that use a machine learning model that is trained to generate unified text associated with user speech, where the unified text includes punction marks, capitalizations of words, inverse text normalization formatting, end of sentence (EOS) detections, and/or end of utterance (EOU) detections. For instance, the machine learning model may receive audio data representing speech as input. The machine learning model may then process the audio data and, based at least on the processing, generate output data associated with the speech. In some examples, the output data may represent tokens, such as tokens associated with automatic speech recognition processing, punctuation and capitalization processing, EOS and/or EOU processing, and/or inverse text normalization processing. In such examples, the tokens may then be processed to generate the unified text.
    Type: Application
    Filed: August 21, 2024
    Publication date: February 26, 2026
    Inventors: Harishchandra Dubey, Myungjong Kim, Utkarsh Vaidya, Oluwatobi Olabiyi, Nourchene Ferchichi
  • Publication number: 20260010706
    Abstract: Approaches presented herein provide for the generation of text transcripts of speech represented in audio data. In particular, an automatic speech recognition (ASR) model can be used together with a retrieval augmented generation (RAG) pipeline to provide for improvement of transcripts that include terminology related, or specific, to a specific knowledge domain. A knowledge base for a given domain can include a number of files or documents in a number of different formats (e.g., documents, images, and webpages) that do not need to be cleaned, classified, or curated. When an ASR generates a transcript where at least one word has a confidence level that falls below a confidence threshold, that transcript can be passed to a language model of the RAG pipeline which can use the retrieved domain-specific data to attempt to identify the appropriate words or terms to use to replace the words tagged as having low confidence.
    Type: Application
    Filed: July 12, 2024
    Publication date: January 8, 2026
    Inventors: Mayank Jain, Fan Qian, Utkarsh Vaidya, Niranjan Wartikar, Sanjay Singh Chauhan, Eileen Margaret Peters Long, Myungjong Kim
  • Publication number: 20260004070
    Abstract: In various examples, detecting breaks in speech for conversational AI systems and applications is described herein. Systems and methods are disclosed herein that use both end of sentence detection and end of utterance detection associated with words from text (e.g., tokens) to determine when to further process various portions of the text. For instance, one or more models may process text data associated with the text, where the text data may be generated using an automatic speech recognition (ARS) model based on audio data representing speech. Based at least on processing the text data, the model(s) may generate and/or output data representing first indicators that the words are associated with ends of sentences, second indicators that the words are associated with ends of utterances, and third indicators that the words are not associated with either ends of sentences or ends of utterances.
    Type: Application
    Filed: June 26, 2024
    Publication date: January 1, 2026
    Inventors: Myungjong Kim, Harishchandra Dubey, Utkarsh Vaidya, Oluwatobi Olabiyi
  • Patent number: 12353793
    Abstract: The disclosure is directed to a process that can predict and prevent an audio artifact from occurring. The process can monitor the systems, processes, and execution threads on a larger system/device, such as a mobile or in-vehicle device. Using a learning algorithm, such as deep neural network (DNN), the information collected can generate a prediction of whether an audio artifact is likely to occur. The process can use a second learning algorithm, which also can be a DNN, to generate recommended system adjustments that can attempt to prevent the audio glitch from occurring. The recommendations can be for various systems and components on the device, such as changing the processing system frequency, the memory frequency, and the audio buffer size. After the audio artifact has been prevented, the system adjustments can be reversed fully or in steps to return the system to its state prior to the system adjustments.
    Type: Grant
    Filed: May 28, 2024
    Date of Patent: July 8, 2025
    Assignee: NVIDIA Corporation
    Inventors: Utkarsh Vaidya, Sumit Bhattacharya
  • Publication number: 20240311080
    Abstract: The disclosure is directed to a process that can predict and prevent an audio artifact from occurring. The process can monitor the systems, processes, and execution threads on a larger system/device, such as a mobile or in-vehicle device. Using a learning algorithm, such as deep neural network (DNN), the information collected can generate a prediction of whether an audio artifact is likely to occur. The process can use a second learning algorithm, which also can be a DNN, to generate recommended system adjustments that can attempt to prevent the audio glitch from occurring. The recommendations can be for various systems and components on the device, such as changing the processing system frequency, the memory frequency, and the audio buffer size. After the audio artifact has been prevented, the system adjustments can be reversed fully or in steps to return the system to its state prior to the system adjustments.
    Type: Application
    Filed: May 28, 2024
    Publication date: September 19, 2024
    Inventors: Utkarsh Vaidya, Sumit Bhattacharya
  • Patent number: 11995378
    Abstract: The disclosure is directed to a process that can predict and prevent an audio artifact from occurring. The process can monitor the systems, processes, and execution threads on a larger system/device, such as a mobile or in-vehicle device. Using a learning algorithm, such as deep neural network (DNN), the information collected can generate a prediction of whether an audio artifact is likely to occur. The process can use a second learning algorithm, which also can be a DNN, to generate recommended system adjustments that can attempt to prevent the audio glitch from occurring. The recommendations can be for various systems and components on the device, such as changing the processing system frequency, the memory frequency, and the audio buffer size. After the audio artifact has been prevented, the system adjustments can be reversed fully or in steps to return the system to its state prior to the system adjustments.
    Type: Grant
    Filed: January 30, 2023
    Date of Patent: May 28, 2024
    Assignee: NVIDIA Corporation
    Inventors: Utkarsh Vaidya, Sumit Bhattacharya
  • Patent number: 11817117
    Abstract: In various examples, end of speech (EOS) for an audio signal is determined based at least in part on a rate of speech for a speaker. For a segment of the audio signal, EOS is indicated based at least in part on an EOS threshold determined based at least in part on the rate of speech for the speaker.
    Type: Grant
    Filed: January 29, 2021
    Date of Patent: November 14, 2023
    Assignee: NVIDIA CORPORATION
    Inventors: Utkarsh Vaidya, Ravindra Yeshwant Lokhande, Viraj Gangadhar Karandikar, Niranjan Rajendra Wartikar, Sumit Kumar Bhattacharya
  • Publication number: 20230298579
    Abstract: Apparatuses, systems, and techniques are presented to recognize speech in an audio signal. In particular, various embodiments can indicate an end of one or more speech segments based, at least in part, on one or more characters predicted to be within these one or more speech segments.
    Type: Application
    Filed: May 25, 2023
    Publication date: September 21, 2023
    Inventors: Utkarsh Vaidya, Sumit Bhattacharya, Viraj Karandikar, Niranjan Wartikar
  • Publication number: 20230168857
    Abstract: The disclosure is directed to a process that can predict and prevent an audio artifact from occurring. The process can monitor the systems, processes, and execution threads on a larger system/ device, such as a mobile or in-vehicle device. Using a learning algorithm, such as deep neural network (DNN), the information collected can generate a prediction of whether an audio artifact is likely to occur. The process can use a second learning algorithm, which also can be a DNN, to generate recommended system adjustments that can attempt to prevent the audio glitch from occurring. The recommendations can be for various systems and components on the device, such as changing the processing system frequency, the memory frequency, and the audio buffer size. After the audio artifact has been prevented, the system adjustments can be reversed fully or in steps to return the system to its state prior to the system adjustments.
    Type: Application
    Filed: January 30, 2023
    Publication date: June 1, 2023
    Inventors: Utkarsh Vaidya, Sumit Bhattacharya
  • Patent number: 11567728
    Abstract: The disclosure is directed to a process that can predict and prevent an audio artifact from occurring. The process can monitor the systems, processes, and execution threads on a larger system/device, such as a mobile or in-vehicle device. Using a learning algorithm, such as deep neural network (DNN), the information collected can generate a prediction of whether an audio artifact is likely to occur. The process can use a second learning algorithm, which also can be a DNN, to generate recommended system adjustments that can attempt to prevent the audio glitch from occurring. The recommendations can be for various systems and components on the device, such as changing the processing system frequency, the memory frequency, and the audio buffer size. After the audio artifact has been prevented, the system adjustments can be reversed fully or in steps to return the system to its state prior to the system adjustments.
    Type: Grant
    Filed: December 14, 2020
    Date of Patent: January 31, 2023
    Assignee: NVIDIA Corporation
    Inventors: Utkarsh Vaidya, Sumit Bhattacharya
  • Publication number: 20220246167
    Abstract: In various examples, end of speech (EOS) for an audio signal is determined based at least in part on a rate of speech for a speaker. For a segment of the audio signal, EOS is indicated based at least in part on an EOS threshold determined based at least in part on the rate of speech for the speaker.
    Type: Application
    Filed: January 29, 2021
    Publication date: August 4, 2022
    Inventors: Utkarsh Vaidya, Ravindra Yeshwant Lokhande, Viraj Gangadhar Karandikar, Niranjan Rajendra Wartikar, Sumit Kumar Bhattacharya
  • Publication number: 20210358490
    Abstract: Apparatuses, systems, and techniques are presented to recognize speech in an audio signal. In particular, various embodiments can indicate an end of one or more speech segments based, at least in part, on one or more characters predicted to be within these one or more speech segments.
    Type: Application
    Filed: May 18, 2020
    Publication date: November 18, 2021
    Inventors: Utkarsh Vaidya, Sumit Bhattacharya, Viraj Karandikar, Niranjan Wartikar
  • Publication number: 20210103425
    Abstract: The disclosure is directed to a process that can predict and prevent an audio artifact from occurring. The process can monitor the systems, processes, and execution threads on a larger system/device, such as a mobile or in-vehicle device. Using a learning algorithm, such as deep neural network (DNN), the information collected can generate a prediction of whether an audio artifact is likely to occur. The process can use a second learning algorithm, which also can be a DNN, to generate recommended system adjustments that can attempt to prevent the audio glitch from occurring. The recommendations can be for various systems and components on the device, such as changing the processing system frequency, the memory frequency, and the audio buffer size. After the audio artifact has been prevented, the system adjustments can be reversed fully or in steps to return the system to its state prior to the system adjustments.
    Type: Application
    Filed: December 14, 2020
    Publication date: April 8, 2021
    Inventors: Utkarsh Vaidya, Sumit Bhattacharya
  • Patent number: 10896021
    Abstract: The disclosure is directed to a process that can predict an audio glitch, and then attempt to preempt the audio glitch. The process can monitor the systems, processes, and execution threads on a larger system or device, such as a mobile device or an in-vehicle device. Using a learning algorithm, such as deep neural network (DNN), the information collected can generate a prediction of whether an audio glitch is likely to occur. An audio glitch can be an audio underrun condition. The process can use a second learning algorithm, which also can be a DNN, to generate recommended system adjustments that can attempt to prevent the audio glitch from occurring. The recommendations can be for various systems and components on the device, such as changing the processing system frequency, the memory frequency, and the audio buffer size. After the audio underrun condition has abated, the system adjustments can be reversed fully or in steps to return the system to its state prior to the system adjustments.
    Type: Grant
    Filed: February 26, 2019
    Date of Patent: January 19, 2021
    Assignee: Nvidia Corporation
    Inventors: Utkarsh Vaidya, Sumit Bhattacharya
  • Publication number: 20200272409
    Abstract: The disclosure is directed to a process that can predict an audio glitch, and then attempt to preempt the audio glitch. The process can monitor the systems, processes, and execution threads on a larger system or device, such as a mobile device or an in-vehicle device. Using a learning algorithm, such as deep neural network (DNN), the information collected can generate a prediction of whether an audio glitch is likely to occur. An audio glitch can be an audio underrun condition. The process can use a second learning algorithm, which also can be a DNN, to generate recommended system adjustments that can attempt to prevent the audio glitch from occurring. The recommendations can be for various systems and components on the device, such as changing the processing system frequency, the memory frequency, and the audio buffer size. After the audio underrun condition has abated, the system adjustments can be reversed fully or in steps to return the system to its state prior to the system adjustments.
    Type: Application
    Filed: February 26, 2019
    Publication date: August 27, 2020
    Inventors: Utkarsh Vaidya, Sumit Bhattacharya