Patents by Inventor Utkarsh Vaidya
Utkarsh Vaidya has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Publication number: 20260204261Abstract: Disclosed are apparatuses, systems, and techniques for generating multi-format transcriptions of speech. The techniques include processing, using an encoder, audio frames representative of a speech to generate embeddings encoding the speech and processing, using multiple decoders, the embeddings to generate multiple transcriptions of the speech. An individual transcription is generated by a respective decoder and conforms to a respective text format that differs from other text formats in capitalization, punctuation, use of non-alphabet characters, and/or identification of individual utterances of the speech.Type: ApplicationFiled: January 16, 2025Publication date: July 16, 2026Inventors: Harishchandra Dubey, Myungjong Kim, Oluwatobi Olabiyi, Utkarsh Vaidya
-
Publication number: 20260162654Abstract: A textual transcript and one or more language indicators are determined using a multilingual speech-to-text (STT) model of a multilingual automatic speech recognition (ASR) system and using an audio sample as input to the multilingual STT model. The textual transcript is associated with the audio sample, and the one or more language indicators are each associated with a respective grammatical unit of one or more grammatical units of the textual transcript. A monolingual language model (LM) of a plurality of monolingual LMs of the ASR system is identified using a language indicator of the one or more language indicators. The textual transcript associated with the audio sample is caused to be refined using the identified LM and using a subset of the textual transcript as input to the identified LM.Type: ApplicationFiled: December 9, 2024Publication date: June 11, 2026Inventors: Myungjong Kim, Mayank Jain, Yitagessu Gebremedhin, Utkarsh Vaidya, Oluwatobi Olabiyi
-
Publication number: 20260154514Abstract: In various examples, techniques are described for adapting a multilingual Large Language Model (LLM) into a bilingual Small Language Model (SLM) that exhibits model capacity to understand, process, and generate content in both English and a Low-Resource Language (LRL). The techniques include compressing the LLM to generate a multilingual SLM and performing continued pre-training on the multilingual SLM to generate the bilingual SLM. The techniques also include performing one or more alignment techniques on the bilingual SLM to adapt the SLM's outputs to human values and expectations regarding, e.g., profanity, privacy, politeness, bias, and/or conversational style. The techniques may generate various training corpora, each including one or more of natural English content, natural LRL content, synthetic LRL content generated via translation from English sources, and transliterated synthetic LRL content based on transliterations of natural and/or synthetic LRL content.Type: ApplicationFiled: September 30, 2025Publication date: June 4, 2026Inventors: Raviraj JOSHI, Kanishk SINGLA, Anusha KAMATH, Raunak KALANI, Utkarsh VAIDYA, Sanjay Singh CHAUHAN, Niranjan WARTIKAR, Eileen Margaret Peters LONG
-
Publication number: 20260064994Abstract: Approaches presented herein provide for the generation of relatively small language models that are optimized for target languages. A multilingual large language model (LLM) can be reduced in size using a process such as language-aware pruning, where individual network parameters have importance scores calculated with respect to the target language and then an appropriate number of lower-importance score parameters are removed from the network. Continued pretraining can be performed using a set of training data including real and/or synthesized text in the target language, to obtain a high performing language model with a limited number of parameters optimized for a target language, as may correspond to a lower-resource language that may otherwise not have enough training data available to sufficiently train a language model from scratch.Type: ApplicationFiled: August 29, 2024Publication date: March 5, 2026Inventors: Raviraj Joshi, Utkarsh Vaidya, Sanjay Singh Chauhan
-
Publication number: 20260057884Abstract: In various examples, generating unified text using speech recognition models for AI systems and applications is described herein. Systems and methods are disclosed that use a machine learning model that is trained to generate unified text associated with user speech, where the unified text includes punction marks, capitalizations of words, inverse text normalization formatting, end of sentence (EOS) detections, and/or end of utterance (EOU) detections. For instance, the machine learning model may receive audio data representing speech as input. The machine learning model may then process the audio data and, based at least on the processing, generate output data associated with the speech. In some examples, the output data may represent tokens, such as tokens associated with automatic speech recognition processing, punctuation and capitalization processing, EOS and/or EOU processing, and/or inverse text normalization processing. In such examples, the tokens may then be processed to generate the unified text.Type: ApplicationFiled: August 21, 2024Publication date: February 26, 2026Inventors: Harishchandra Dubey, Myungjong Kim, Utkarsh Vaidya, Oluwatobi Olabiyi, Nourchene Ferchichi
-
Publication number: 20260010706Abstract: Approaches presented herein provide for the generation of text transcripts of speech represented in audio data. In particular, an automatic speech recognition (ASR) model can be used together with a retrieval augmented generation (RAG) pipeline to provide for improvement of transcripts that include terminology related, or specific, to a specific knowledge domain. A knowledge base for a given domain can include a number of files or documents in a number of different formats (e.g., documents, images, and webpages) that do not need to be cleaned, classified, or curated. When an ASR generates a transcript where at least one word has a confidence level that falls below a confidence threshold, that transcript can be passed to a language model of the RAG pipeline which can use the retrieved domain-specific data to attempt to identify the appropriate words or terms to use to replace the words tagged as having low confidence.Type: ApplicationFiled: July 12, 2024Publication date: January 8, 2026Inventors: Mayank Jain, Fan Qian, Utkarsh Vaidya, Niranjan Wartikar, Sanjay Singh Chauhan, Eileen Margaret Peters Long, Myungjong Kim
-
Publication number: 20260004070Abstract: In various examples, detecting breaks in speech for conversational AI systems and applications is described herein. Systems and methods are disclosed herein that use both end of sentence detection and end of utterance detection associated with words from text (e.g., tokens) to determine when to further process various portions of the text. For instance, one or more models may process text data associated with the text, where the text data may be generated using an automatic speech recognition (ARS) model based on audio data representing speech. Based at least on processing the text data, the model(s) may generate and/or output data representing first indicators that the words are associated with ends of sentences, second indicators that the words are associated with ends of utterances, and third indicators that the words are not associated with either ends of sentences or ends of utterances.Type: ApplicationFiled: June 26, 2024Publication date: January 1, 2026Inventors: Myungjong Kim, Harishchandra Dubey, Utkarsh Vaidya, Oluwatobi Olabiyi
-
Patent number: 12353793Abstract: The disclosure is directed to a process that can predict and prevent an audio artifact from occurring. The process can monitor the systems, processes, and execution threads on a larger system/device, such as a mobile or in-vehicle device. Using a learning algorithm, such as deep neural network (DNN), the information collected can generate a prediction of whether an audio artifact is likely to occur. The process can use a second learning algorithm, which also can be a DNN, to generate recommended system adjustments that can attempt to prevent the audio glitch from occurring. The recommendations can be for various systems and components on the device, such as changing the processing system frequency, the memory frequency, and the audio buffer size. After the audio artifact has been prevented, the system adjustments can be reversed fully or in steps to return the system to its state prior to the system adjustments.Type: GrantFiled: May 28, 2024Date of Patent: July 8, 2025Assignee: NVIDIA CorporationInventors: Utkarsh Vaidya, Sumit Bhattacharya
-
Publication number: 20240311080Abstract: The disclosure is directed to a process that can predict and prevent an audio artifact from occurring. The process can monitor the systems, processes, and execution threads on a larger system/device, such as a mobile or in-vehicle device. Using a learning algorithm, such as deep neural network (DNN), the information collected can generate a prediction of whether an audio artifact is likely to occur. The process can use a second learning algorithm, which also can be a DNN, to generate recommended system adjustments that can attempt to prevent the audio glitch from occurring. The recommendations can be for various systems and components on the device, such as changing the processing system frequency, the memory frequency, and the audio buffer size. After the audio artifact has been prevented, the system adjustments can be reversed fully or in steps to return the system to its state prior to the system adjustments.Type: ApplicationFiled: May 28, 2024Publication date: September 19, 2024Inventors: Utkarsh Vaidya, Sumit Bhattacharya
-
Patent number: 11995378Abstract: The disclosure is directed to a process that can predict and prevent an audio artifact from occurring. The process can monitor the systems, processes, and execution threads on a larger system/device, such as a mobile or in-vehicle device. Using a learning algorithm, such as deep neural network (DNN), the information collected can generate a prediction of whether an audio artifact is likely to occur. The process can use a second learning algorithm, which also can be a DNN, to generate recommended system adjustments that can attempt to prevent the audio glitch from occurring. The recommendations can be for various systems and components on the device, such as changing the processing system frequency, the memory frequency, and the audio buffer size. After the audio artifact has been prevented, the system adjustments can be reversed fully or in steps to return the system to its state prior to the system adjustments.Type: GrantFiled: January 30, 2023Date of Patent: May 28, 2024Assignee: NVIDIA CorporationInventors: Utkarsh Vaidya, Sumit Bhattacharya
-
Patent number: 11817117Abstract: In various examples, end of speech (EOS) for an audio signal is determined based at least in part on a rate of speech for a speaker. For a segment of the audio signal, EOS is indicated based at least in part on an EOS threshold determined based at least in part on the rate of speech for the speaker.Type: GrantFiled: January 29, 2021Date of Patent: November 14, 2023Assignee: NVIDIA CORPORATIONInventors: Utkarsh Vaidya, Ravindra Yeshwant Lokhande, Viraj Gangadhar Karandikar, Niranjan Rajendra Wartikar, Sumit Kumar Bhattacharya
-
Publication number: 20230298579Abstract: Apparatuses, systems, and techniques are presented to recognize speech in an audio signal. In particular, various embodiments can indicate an end of one or more speech segments based, at least in part, on one or more characters predicted to be within these one or more speech segments.Type: ApplicationFiled: May 25, 2023Publication date: September 21, 2023Inventors: Utkarsh Vaidya, Sumit Bhattacharya, Viraj Karandikar, Niranjan Wartikar
-
Publication number: 20230168857Abstract: The disclosure is directed to a process that can predict and prevent an audio artifact from occurring. The process can monitor the systems, processes, and execution threads on a larger system/ device, such as a mobile or in-vehicle device. Using a learning algorithm, such as deep neural network (DNN), the information collected can generate a prediction of whether an audio artifact is likely to occur. The process can use a second learning algorithm, which also can be a DNN, to generate recommended system adjustments that can attempt to prevent the audio glitch from occurring. The recommendations can be for various systems and components on the device, such as changing the processing system frequency, the memory frequency, and the audio buffer size. After the audio artifact has been prevented, the system adjustments can be reversed fully or in steps to return the system to its state prior to the system adjustments.Type: ApplicationFiled: January 30, 2023Publication date: June 1, 2023Inventors: Utkarsh Vaidya, Sumit Bhattacharya
-
Patent number: 11567728Abstract: The disclosure is directed to a process that can predict and prevent an audio artifact from occurring. The process can monitor the systems, processes, and execution threads on a larger system/device, such as a mobile or in-vehicle device. Using a learning algorithm, such as deep neural network (DNN), the information collected can generate a prediction of whether an audio artifact is likely to occur. The process can use a second learning algorithm, which also can be a DNN, to generate recommended system adjustments that can attempt to prevent the audio glitch from occurring. The recommendations can be for various systems and components on the device, such as changing the processing system frequency, the memory frequency, and the audio buffer size. After the audio artifact has been prevented, the system adjustments can be reversed fully or in steps to return the system to its state prior to the system adjustments.Type: GrantFiled: December 14, 2020Date of Patent: January 31, 2023Assignee: NVIDIA CorporationInventors: Utkarsh Vaidya, Sumit Bhattacharya
-
Publication number: 20220246167Abstract: In various examples, end of speech (EOS) for an audio signal is determined based at least in part on a rate of speech for a speaker. For a segment of the audio signal, EOS is indicated based at least in part on an EOS threshold determined based at least in part on the rate of speech for the speaker.Type: ApplicationFiled: January 29, 2021Publication date: August 4, 2022Inventors: Utkarsh Vaidya, Ravindra Yeshwant Lokhande, Viraj Gangadhar Karandikar, Niranjan Rajendra Wartikar, Sumit Kumar Bhattacharya
-
Publication number: 20210358490Abstract: Apparatuses, systems, and techniques are presented to recognize speech in an audio signal. In particular, various embodiments can indicate an end of one or more speech segments based, at least in part, on one or more characters predicted to be within these one or more speech segments.Type: ApplicationFiled: May 18, 2020Publication date: November 18, 2021Inventors: Utkarsh Vaidya, Sumit Bhattacharya, Viraj Karandikar, Niranjan Wartikar
-
Publication number: 20210103425Abstract: The disclosure is directed to a process that can predict and prevent an audio artifact from occurring. The process can monitor the systems, processes, and execution threads on a larger system/device, such as a mobile or in-vehicle device. Using a learning algorithm, such as deep neural network (DNN), the information collected can generate a prediction of whether an audio artifact is likely to occur. The process can use a second learning algorithm, which also can be a DNN, to generate recommended system adjustments that can attempt to prevent the audio glitch from occurring. The recommendations can be for various systems and components on the device, such as changing the processing system frequency, the memory frequency, and the audio buffer size. After the audio artifact has been prevented, the system adjustments can be reversed fully or in steps to return the system to its state prior to the system adjustments.Type: ApplicationFiled: December 14, 2020Publication date: April 8, 2021Inventors: Utkarsh Vaidya, Sumit Bhattacharya
-
Patent number: 10896021Abstract: The disclosure is directed to a process that can predict an audio glitch, and then attempt to preempt the audio glitch. The process can monitor the systems, processes, and execution threads on a larger system or device, such as a mobile device or an in-vehicle device. Using a learning algorithm, such as deep neural network (DNN), the information collected can generate a prediction of whether an audio glitch is likely to occur. An audio glitch can be an audio underrun condition. The process can use a second learning algorithm, which also can be a DNN, to generate recommended system adjustments that can attempt to prevent the audio glitch from occurring. The recommendations can be for various systems and components on the device, such as changing the processing system frequency, the memory frequency, and the audio buffer size. After the audio underrun condition has abated, the system adjustments can be reversed fully or in steps to return the system to its state prior to the system adjustments.Type: GrantFiled: February 26, 2019Date of Patent: January 19, 2021Assignee: Nvidia CorporationInventors: Utkarsh Vaidya, Sumit Bhattacharya
-
Publication number: 20200272409Abstract: The disclosure is directed to a process that can predict an audio glitch, and then attempt to preempt the audio glitch. The process can monitor the systems, processes, and execution threads on a larger system or device, such as a mobile device or an in-vehicle device. Using a learning algorithm, such as deep neural network (DNN), the information collected can generate a prediction of whether an audio glitch is likely to occur. An audio glitch can be an audio underrun condition. The process can use a second learning algorithm, which also can be a DNN, to generate recommended system adjustments that can attempt to prevent the audio glitch from occurring. The recommendations can be for various systems and components on the device, such as changing the processing system frequency, the memory frequency, and the audio buffer size. After the audio underrun condition has abated, the system adjustments can be reversed fully or in steps to return the system to its state prior to the system adjustments.Type: ApplicationFiled: February 26, 2019Publication date: August 27, 2020Inventors: Utkarsh Vaidya, Sumit Bhattacharya