Patents by Inventor Vladimir Bataev
Vladimir Bataev has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Patent number: 12731577Abstract: Systems and methods provide for a machine learning system to train a machine learning model to output a penalty-free emission when processing an auditory input. For example, as the system generates paths through a probability lattice, one or more paths may include a penalty-free emission that skips at least one frame associated with the probability lattice, but that does not add a cost to a final path cost. The use of the penalty-free emissions may be represented through one or more graphical representations used for training in order to develop loss functions for models. One or more of these frameworks may be incorporated into automatic speech recognition pipelines to improve training while also reducing coding requirements to simplify debugging operations.Type: GrantFiled: July 20, 2023Date of Patent: September 8, 2026Assignee: Nvidia CorporationInventors: Aleksandr Laptev, Vladimir Bataev, Igor Gitman, Boris Ginsburg
-
Publication number: 20260038488Abstract: In various examples, first textual data may be applied to a first MLM to generate an intermediate speech representation (e.g., a frequency-domain representation), the intermediate audio representation and a second MLM may be used to generate output data indicating second textual data, and parameters of the second MLM may be updated using the output data and ground truth data associated with the first textual data. The first MLM may include a trained Text-To-Speech (TTS) model and the second MLM may include an Automatic Speech Recognition (ASR) model. A generator from a generative adversarial networks may be used to enhance an initial intermediate audio representation generated using the first MLM and the enhanced intermediate audio representation may be provided to the second MLM. The generator may include generator blocks that receive the initial intermediate audio representation to sequentially generate the enhanced intermediate audio representation.Type: ApplicationFiled: October 13, 2025Publication date: February 5, 2026Inventors: Vladimir Bataev, Roman Korostik, Evgenii Shabalin, Vitaly Sergeyevich Lavrukhin, Boris Ginsburg
-
Publication number: 20260038487Abstract: In various examples, first textual data may be applied to a first MLM to generate an intermediate speech representation (e.g., a frequency-domain representation), the intermediate audio representation and a second MLM may be used to generate output data indicating second textual data, and parameters of the second MLM may be updated using the output data and ground truth data associated with the first textual data. The first MLM may include a trained Text-To-Speech (TTS) model and the second MLM may include an Automatic Speech Recognition (ASR) model. A generator from a generative adversarial networks may be used to enhance an initial intermediate audio representation generated using the first MLM and the enhanced intermediate audio representation may be provided to the second MLM. The generator may include generator blocks that receive the initial intermediate audio representation to sequentially generate the enhanced intermediate audio representation.Type: ApplicationFiled: October 13, 2025Publication date: February 5, 2026Inventors: Vladimir Bataev, Roman Korostik, Evgenii Shabalin, Vitaly Sergeyevich Lavrukhin, Boris Ginsburg
-
Publication number: 20260004767Abstract: Disclosed are apparatuses, systems, and techniques that use a text-to-speech (TTS) transducer to perform TTS operations. The techniques include generating an initial input for a second model using an output of a first model. The techniques include generating, using the second model and the initial input, a first set of audio codes. The techniques include iteratively generating subsequent sets of audio codes using, at each iteration, the second model and a respective subsequent input for the second model. The respective subsequent input can reflect at least one previous set of audio codes generated by the second model.Type: ApplicationFiled: March 26, 2025Publication date: January 1, 2026Inventors: Vladimir Bataev, Subhankar Ghosh, Vitaly Lavrukhin, Boris Ginsburg
-
Patent number: 12444409Abstract: In various examples, first textual data may be applied to a first MLM to generate an intermediate speech representation (e.g., a frequency-domain representation), the intermediate audio representation and a second MLM may be used to generate output data indicating second textual data, and parameters of the second MLM may be updated using the output data and ground truth data associated with the first textual data. The first MLM may include a trained Text-To-Speech (TTS) model and the second MLM may include an Automatic Speech Recognition (ASR) model. A generator from a generative adversarial networks may be used to enhance an initial intermediate audio representation generated using the first MLM and the enhanced intermediate audio representation may be provided to the second MLM. The generator may include generator blocks that receive the initial intermediate audio representation to sequentially generate the enhanced intermediate audio representation.Type: GrantFiled: September 15, 2023Date of Patent: October 14, 2025Assignee: NVIDIA CorporationInventors: Vladimir Bataev, Roman Korostik, Evgenii Shabalin, Vitaly Sergeyevich Lavrukhin, Boris Ginsburg
-
Publication number: 20250279091Abstract: Disclosed are apparatuses, systems, and techniques that use label-looping processing for efficient automatic speech recognition (ASR). The techniques include performing a plurality of iterations of an outer processing loop to identify content units (CUs) of a media item having multiple frames. An individual iteration of the outer processing loop includes updating, using a first neural network (NN) and identified non-blank CU, a state of the media item and performing one or more iterations of an inner processing loop. An individual iteration of the inner processing loop includes processing, using a second NN, the state of the media item and an individual frame to predict a CU associated with the individual frame. The iterations of the inner processing loop are performed until the predicted CU corresponds to a non-blank CU. The identified plurality of CUs is used to generate a representation of the media item.Type: ApplicationFiled: August 29, 2024Publication date: September 4, 2025Inventors: Vladimir Bataev, Hainan Xu, Vitaly Lavrukhin, Boris Ginsburg
-
Publication number: 20250218433Abstract: Disclosed are apparatuses, systems, and techniques that may use machine learning for implementing automatic speech recognition (ASR) facilitated with a search for target words. The techniques include applying an ASR model to audio data to generate an ASR output representative of a likelihood that the audio data comprises one or more spoken speech units (SUs), generating, using the ASR output, a first score characterizing a likelihood that the audio data comprises a first word, wherein the first word comprises a dictionary word, generating, using the ASR output, a second score characterizing a likelihood that the audio data comprises a second word, wherein the second word comprises a word of a plurality of target words, wherein the plurality of target words is identified based at least on a context of the audio data, and predicting, using the first score and the second score, a spoken word associated with the audio data.Type: ApplicationFiled: March 25, 2024Publication date: July 3, 2025Inventors: Andrei Andrusenko, Aleksandr Laptev, Vladimir Bataev, Vitaly Lavrukhin, Boris Ginsburg
-
Publication number: 20240265913Abstract: Systems and methods provide for a machine learning system to train a machine learning model to output a penalty-free emission when processing an auditory input. For example, as the system generates paths through a probability lattice, one or more paths may include a penalty-free emission that skips at least one frame associated with the probability lattice, but that does not add a cost to a final path cost. The use of the penalty-free emissions may be represented through one or more graphical representations used for training in order to develop loss functions for models. One or more of these frameworks may be incorporated into automatic speech recognition pipelines to improve training while also reducing coding requirements to simplify debugging operations.Type: ApplicationFiled: July 20, 2023Publication date: August 8, 2024Inventors: Aleksandr Laptev, Vladimir Bataev, Igor Gitman, Boris Ginsburg
-
Publication number: 20240265912Abstract: Systems and methods provide for a machine learning system to train a machine learning model to output a penalty-free emission when processing an auditory input. For example, as the system generates paths through a probability lattice, one or more paths may include a penalty-free emission that skips at least one frame associated with the probability lattice, but that does not add a cost to a final path cost. The use of the penalty-free emissions may be represented through one or more graphical representations used for training in order to develop loss functions for models. One or more of these frameworks may be incorporated into automatic speech recognition pipelines to improve training while also reducing coding requirements to simplify debugging operations.Type: ApplicationFiled: July 20, 2023Publication date: August 8, 2024Inventors: Aleksandr Laptev, Vladimir Bataev, Igor Gitman, Boris Ginsburg
-
Publication number: 20240233714Abstract: In various examples, first textual data may be applied to a first MLM to generate an intermediate speech representation (e.g., a frequency-domain representation), the intermediate audio representation and a second MLM may be used to generate output data indicating second textual data, and parameters of the second MLM may be updated using the output data and ground truth data associated with the first textual data. The first MLM may include a trained Text-To-Speech (TTS) model and the second MLM may include an Automatic Speech Recognition (ASR) model. A generator from a generative adversarial networks may be used to enhance an initial intermediate audio representation generated using the first MLM and the enhanced intermediate audio representation may be provided to the second MLM. The generator may include generator blocks that receive the initial intermediate audio representation to sequentially generate the enhanced intermediate audio representation.Type: ApplicationFiled: September 15, 2023Publication date: July 11, 2024Inventors: Vladimir Bataev, Roman Korostik, Evgenii Shabalin, Vitaly Sergeyevich Lavrukhin, Boris Ginsburg
-
Publication number: 20240135920Abstract: In various examples, first textual data may be applied to a first MLM to generate an intermediate speech representation (e.g., a frequency-domain representation), the intermediate audio representation and a second MLM may be used to generate output data indicating second textual data, and parameters of the second MLM may be updated using the output data and ground truth data associated with the first textual data. The first MLM may include a trained Text-To-Speech (TTS) model and the second MLM may include an Automatic Speech Recognition (ASR) model. A generator from a generative adversarial networks may be used to enhance an initial intermediate audio representation generated using the first MLM and the enhanced intermediate audio representation may be provided to the second MLM. The generator may include generator blocks that receive the initial intermediate audio representation to sequentially generate the enhanced intermediate audio representation.Type: ApplicationFiled: September 14, 2023Publication date: April 25, 2024Inventors: Vladimir Bataev, Roman Korostik, Evgenii Shabalin, Vitaly Sergeyevich Lavrukhin, Boris Ginsburg