Patents by Inventor Aleksandr Laptev

Aleksandr Laptev has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Publication number: 20250218433
    Abstract: Disclosed are apparatuses, systems, and techniques that may use machine learning for implementing automatic speech recognition (ASR) facilitated with a search for target words. The techniques include applying an ASR model to audio data to generate an ASR output representative of a likelihood that the audio data comprises one or more spoken speech units (SUs), generating, using the ASR output, a first score characterizing a likelihood that the audio data comprises a first word, wherein the first word comprises a dictionary word, generating, using the ASR output, a second score characterizing a likelihood that the audio data comprises a second word, wherein the second word comprises a word of a plurality of target words, wherein the plurality of target words is identified based at least on a context of the audio data, and predicting, using the first score and the second score, a spoken word associated with the audio data.
    Type: Application
    Filed: March 25, 2024
    Publication date: July 3, 2025
    Inventors: Andrei Andrusenko, Aleksandr Laptev, Vladimir Bataev, Vitaly Lavrukhin, Boris Ginsburg
  • Publication number: 20250029618
    Abstract: Disclosed are apparatuses, systems, and techniques that may use machine learning for implementing speaker recognition, verification, and/or diarization. The techniques include receiving a first set of audio data channels (ADCs) jointly capturing a speech produced by one or more speakers and obtaining, using the first set of ADCs, a second set of one or more ADCs. Individual ADCs of the second set of ADCs represent one or more channels of the first set of ADCs, and at least one channel of the second set of ADCs represents a cluster of two or more ADCs of the first set of ADCs, the two of more ADCs being selected based on similarity of audio data of the two or more ADCs. The techniques further include processing, using an audio processing neural network model, the second set of ADCs to obtain an association of the speech to the one or more speakers.
    Type: Application
    Filed: January 11, 2024
    Publication date: January 23, 2025
    Inventors: Taejin Park, Ante Jukic, He Huang, Venkata Naga Krishna Chaitanya Puvvada, Kunal Dhawan, Nithin Rao Koluguri, Nikolay Karpov, Aleksandr Laptev, Jagadeesh Balam
  • Publication number: 20250029632
    Abstract: Disclosed are apparatuses, systems, and techniques that may use machine learning for implementing speaker recognition, verification, and/or diarization. The techniques include processing audio data channels (ADCs) using a voice detection model to determine voice activity likelihoods (VALs) that individual ADCs include speech, obtaining, using VALs, a second set of ADC(s), and processing, using an audio processing a neural network (NN) model, the second set of ADCs to obtain association of the speech to the one or more speakers. The techniques also include generating a plurality of embeddings associated with the ADCs, processing the plurality of embeddings to obtain aggregated embedding(s) that represent audio data of multiple ADCs, and processing the aggregated embedding(s), using the audio processing NN model, to obtain association of the speech to the one or more speakers.
    Type: Application
    Filed: January 11, 2024
    Publication date: January 23, 2025
    Inventors: Taejin Park, Ante Jukic, He Huang, Venkata Naga Krishna Chaitanya Puvvada, Kunal Dhawan, Nithin Rao Koluguri, Nikolay Karpov, Aleksandr Laptev, Jagadeesh Balam
  • Publication number: 20240265913
    Abstract: Systems and methods provide for a machine learning system to train a machine learning model to output a penalty-free emission when processing an auditory input. For example, as the system generates paths through a probability lattice, one or more paths may include a penalty-free emission that skips at least one frame associated with the probability lattice, but that does not add a cost to a final path cost. The use of the penalty-free emissions may be represented through one or more graphical representations used for training in order to develop loss functions for models. One or more of these frameworks may be incorporated into automatic speech recognition pipelines to improve training while also reducing coding requirements to simplify debugging operations.
    Type: Application
    Filed: July 20, 2023
    Publication date: August 8, 2024
    Inventors: Aleksandr Laptev, Vladimir Bataev, Igor Gitman, Boris Ginsburg
  • Publication number: 20240265912
    Abstract: Systems and methods provide for a machine learning system to train a machine learning model to output a penalty-free emission when processing an auditory input. For example, as the system generates paths through a probability lattice, one or more paths may include a penalty-free emission that skips at least one frame associated with the probability lattice, but that does not add a cost to a final path cost. The use of the penalty-free emissions may be represented through one or more graphical representations used for training in order to develop loss functions for models. One or more of these frameworks may be incorporated into automatic speech recognition pipelines to improve training while also reducing coding requirements to simplify debugging operations.
    Type: Application
    Filed: July 20, 2023
    Publication date: August 8, 2024
    Inventors: Aleksandr Laptev, Vladimir Bataev, Igor Gitman, Boris Ginsburg