Patents by Inventor Ariya Rastrow
Ariya Rastrow has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Patent number: 12658191Abstract: A system that incorporates contextual entity information when performing automatic speech processing (ASR) using a neural network architecture. The system identifies entities that may be related to the context of an utterance. Text information and pronunciation information related to those entities are encoded and used to determine biasing data that is applied to encoded audio data. The resulting adjusted encoded audio data is processed by the existing neural network architecture to determine ASR data representing a transcription of the utterance.Type: GrantFiled: March 28, 2023Date of Patent: June 16, 2026Assignee: Amazon Technologies, Inc.Inventors: Jing Liu, Qi Luo, Xinyu Ren, Ariya Rastrow, Ankur Gandhe, Denis Filimonov, Grant Strimel, Andreas Stolcke, Ivan Bulyko, Rahul Pandey
-
Patent number: 12614543Abstract: Techniques for performing spoken language understanding (SLU) processing on a device are described. Example embodiments involve a device determining whether a spoken input corresponds to a supported spoken input class, a supported spoken input with dynamic content class or an unsupported spoken input class. For a spoken input corresponding to the supported spoken input with dynamic content class, the device may determine an entity corresponding to the spoken input from a set of entities, which may be determined based on device context data and/or user profile data. For a spoken input corresponding to the supported spoken input class, the device may determine an intent and entity using stored data. For a spoken input corresponding to the unsupported spoken input class, the device may send the audio data to a system for processing.Type: GrantFiled: February 14, 2022Date of Patent: April 28, 2026Assignee: Amazon Technologies, Inc.Inventors: Anastasios Alexandridis, Kanthashree Mysore Sathyendra, Grant Strimel, Pavel Kveton, Jon A. Webb, Ariya Rastrow
-
Patent number: 12586565Abstract: Techniques for biasing for entities during automatic speech recognition (ASR) processing are described. In some embodiments, a system implements a gating component that is configured to switch on and off entity biasing on an audio frame basis when processing a spoken input. The gating component processes an audio frame to determine whether the audio frame likely includes a representation of a custom entity. Based on the determination, a biasing component, which is configured to generate entity embeddings, may be turned on or off. In this manner, entity biasing does not run on every audio frame, but only on the audio frames where it can be helpful in increasing ASR accuracy.Type: GrantFiled: March 29, 2023Date of Patent: March 24, 2026Assignee: Amazon Technologies, Inc.Inventors: Anastasios Alexandridis, Kanthashree Mysore Sathyendra, Grant Strimel, Feng-Ju Chang, Ariya Rastrow, Nathan Anthony Susanj, Athanasios Mouchtaris
-
Patent number: 12531070Abstract: Some speech processing systems may handle some commands on-device rather than sending the audio data to a second device or system for processing. The first device may have limited speech processing capabilities sufficient for handling common language and/or commands, while the second device (e.g., an edge device and/or a remote system) may call on additional language models, entity libraries, skill components, etc. to perform additional tasks. An intermediate data generator may facilitate dividing speech processing operations between devices by generating a stream of data that includes a first-pass ASR output (e.g., a word or sub-word lattice) and other characteristics of the audio data such as whisper detection, speaker identification, media signatures, etc. The second device can perform the additional processing using the data stream; e.g., without using the audio data. Thus, privacy may be enhanced by processing the audio data locally without sending it to other devices/systems.Type: GrantFiled: June 6, 2023Date of Patent: January 20, 2026Assignee: Amazon Technologies, Inc.Inventors: Stanislaw Ignacy Pasko, Pawel Zelazko, Cagdas Bak, Eli Joshua Fidler, Michal Kowalczuk, Andrew Oberlin, Ariya Rastrow
-
Patent number: 12531056Abstract: Techniques for ASR processing using language model (LM)-generated context are described. A LM is prompted to generate words that are relevant for/may be included in a future user input. The prompt to the LM can include words from user interaction history, dialog history, dialog topic, user preferences, etc. The information included in the prompt may focus on rare or unique words rather than words that the ASR model is already confident in recognizing. The techniques can be plugged into an existing/pretrained ASR model and can be used with any existing/pretrained LM, thus saving resources needed to implement and maintain the components.Type: GrantFiled: December 15, 2023Date of Patent: January 20, 2026Assignee: Amazon Technologies, Inc.Inventors: Jing Liu, Mingzhi Yu, Sunwoo Kim, Grant Strimel, Ross William McGowan, Kanthashree Mysore Sathyendra, Andreas Stolcke, Ariya Rastrow
-
Patent number: 12512096Abstract: Described herein is a system for rescoring automatic speech recognition hypotheses for conversational devices that have multi-turn dialogs with a user. The system leverages dialog context by incorporating data related to past user utterances and data related to the system generated response corresponding to the past user utterance. Incorporation of this data improves recognition of a particular user utterance within the dialog.Type: GrantFiled: June 7, 2021Date of Patent: December 30, 2025Assignee: Amazon Technologies, Inc.Inventors: Behnam Hedayatnia, Anirudh Raju, Ankur Gandhe, Chandra Prakash Khatri, Ariya Rastrow, Anushree Venkatesh, Arindam Mandal, Raefer Christopher Gabriel, Ahmad Shikib Mehri
-
Patent number: 12488798Abstract: A first neural network (NN) model may generate labels for training a second NN model. The second NN model may represent instances of a NN model operating on multiple different devices (e.g., decentralized user and/or edge devices). The system may include using a “teacher” model to process data received by one or more of the devices to generate a labeled dataset. The system may use the labeled dataset and a “student” model to calculate gradient data for updating the student model. The student model may be the same or similar to NN model instances operating on the devices. The system may validate the updated student model to determine, for example, whether it exhibits improved performance when processing the newly received data and/or historical data. The system may distribute the validated update to the devices.Type: GrantFiled: June 29, 2022Date of Patent: December 2, 2025Assignee: Amazon Technologies, Inc.Inventors: Aaron Eakin, Ariya Rastrow, James Garnet Droppo, Gurpreet Singh Chadha, Pankaj Sitpure, Buddha Puneeth Nandanoor, Santosh Kumar Patro, Milind Mukesh Rao, Gopinath Chennupati, Andrew Oberlin, Prahalad Venkataramanan, Frederick Weber, Venkata Kishore Nandury, Anand Mohan
-
Patent number: 12412567Abstract: Techniques for reducing latency in processing of audio data, where the latency may be caused in detecting audio of interest in the audio data, are described. A device that captures audio data may include a detection component to determine when the audio data includes audio of interest (e.g., device-directed speech), and an audio embedding generator to generate embedding vectors for the captured audio data while the detection component processes the audio data. The device may generate an embedding vector for audio data captured at the device for a duration of time; determine, at the end of the duration of time, that the audio data represents audio of interest; and send the embedding vector to an audio processing component (e.g., an automatic speech recognition component) for processing.Type: GrantFiled: May 5, 2021Date of Patent: September 9, 2025Assignee: Amazon Technologies, Inc.Inventors: Bjorn Hoffmeister, Ariya Rastrow, Grant Strimel
-
Publication number: 20250174231Abstract: A speech interface device is configured to detect an interrupt event and process a voice command without detecting a wakeword. The device includes on-device interrupt architecture configured to detect when device-directed speech is present and send audio data to a remote system for speech processing. This architecture includes an interrupt detector that detects an interrupt event (e.g., device-directed speech) with low latency, enabling the device to quickly lower a volume of output audio and/or perform other actions in response to a potential voice command. In addition, the architecture includes a device directed classifier that processes an entire utterance and corresponding semantic information and detects device-directed speech with high accuracy. Using the device directed classifier, the device may reject the interrupt event and increase a volume of the output audio or may accept the interrupt event, causing the output audio to end and performing speech processing on the audio data.Type: ApplicationFiled: January 21, 2025Publication date: May 29, 2025Inventors: Ariya Rastrow, Eli Joshua Fidler, Roland Maximilian Rolf Maas, Nikko Strom, Aaron Eakin, Diamond Bishop, Bjorn Hoffmeister, Sanjeev Mishra
-
Patent number: 12236950Abstract: A speech interface device is configured to detect an interrupt event and process a voice command without detecting a wakeword. The device includes on-device interrupt architecture configured to detect when device-directed speech is present and send audio data to a remote system for speech processing. This architecture includes an interrupt detector that detects an interrupt event (e.g., device-directed speech) with low latency, enabling the device to quickly lower a volume of output audio and/or perform other actions in response to a potential voice command. In addition, the architecture includes a device directed classifier that processes an entire utterance and corresponding semantic information and detects device-directed speech with high accuracy. Using the device directed classifier, the device may reject the interrupt event and increase a volume of the output audio or may accept the interrupt event, causing the output audio to end and performing speech processing on the audio data.Type: GrantFiled: January 3, 2023Date of Patent: February 25, 2025Assignee: Amazon Technologies, Inc.Inventors: Ariya Rastrow, Eli Joshua Fidler, Roland Maximilian Rolf Maas, Nikko Strom, Aaron Eakin, Diamond Bishop, Bjorn Hoffmeister, Sanjeev Mishra
-
Patent number: 12211517Abstract: A speech-processing system may determine potential endpoints in a user's speech. Such endpoint prediction may include determining a potential endpoint in a stream of audio data, and may additionally including determining an endpoint score representing a likelihood that the potential endpoint represents an end of speech representing a complete user input. When the potential endpoint has been determined, the system may publish a transcript of speech that preceded the potential endpoint, and send it to downstream components. The system may continue to transcribe audio data and determine additional potential endpoints while the downstream components process the transcript. The downstream components may determine whether the transcript is complete; e.g., represents the entirety of the user input. Final endpoint determinations may be made based on the results of the downstream processing including automatic speech recognition, natural language understanding, etc.Type: GrantFiled: September 15, 2021Date of Patent: January 28, 2025Assignee: Amazon Technologies, Inc.Inventors: Roland Maximilian Rolf Maas, Bjorn Hoffmeister, Ariya Rastrow, James Garnet Droppo, Veerdhawal Pande, Maarten Van Segbroeck, Gautam Tiwari, Andrew Smith, Eli Joshua Fidler
-
Patent number: 12205574Abstract: Techniques for using multiple machine learning (ML) models, with varying compute costs, for ASR processing is described. The system may include an arbitrator component configured to determine which ML model is to be used to process an audio frame from a sequence of audio frames representing a spoken natural language input. The arbitrator component may switch between the ML models, on a frame-by-frame basis, to reduce an overall compute cost for the entire spoken natural language input. The outputs of the different ML models may be combined to determine the final output for the entire spoken natural language input.Type: GrantFiled: March 22, 2021Date of Patent: January 21, 2025Assignee: Amazon Technologies, Inc.Inventors: Grant Strimel, Ariya Rastrow, Jonathan Jenner Macoskey
-
Patent number: 12039975Abstract: A natural language system may be configured to act as a participant in a conversation between two users. The system may determine when a user expression such as speech, a gesture, or the like is directed from one user to the other. The system may processing input data related the expression (such as audio data, input data, language processing result data, conversation context data, etc.) to determine if the system should interject a response to the user-to-user expression. If so, the system may process the input data to determine a response and output it. The system may track that response as part of the data related to the ongoing conversation.Type: GrantFiled: December 4, 2020Date of Patent: July 16, 2024Assignee: Amazon Technologies, Inc.Inventors: Prakash Krishnan, Arindam Mandal, Siddhartha Reddy Jonnalagadda, Nikko Strom, Ariya Rastrow, Shiv Naga Prasad Vitaladevuni, Angeliki Metallinou, Vincent Auvray, Minmin Shen, Josey Diego Sandoval, Rohit Prasad, Thomas Taylor, Amotz Maimon
-
Patent number: 12014726Abstract: Exemplary embodiments relate to adapting a generic language model during runtime using domain-specific language model data. The system performs an audio frame-level analysis, to determine if the utterance corresponds to a particular domain and whether the ASR hypothesis needs to be rescored. The system processes, using a trained classifier, the ASR hypothesis (a partial hypothesis) generated for the audio data processed so far. The system determines whether to rescore the hypothesis after every few audio frames (representing a word in the utterance) are processed by the speech recognition system.Type: GrantFiled: March 28, 2022Date of Patent: June 18, 2024Assignee: Amazon Technologies, Inc.Inventors: Ankur Gandhe, Ariya Rastrow, Roland Maximilian Rolf Maas, Bjorn Hoffmeister
-
Patent number: 11908468Abstract: A system that is capable of resolving anaphora using timing data received by a local device. A local device outputs audio representing a list of entries. The audio may represent synthesized speech of the list of entries. A user can interrupt the device to select an entry in the list, such as by saying “that one.” The local device can determine an offset time representing the time between when audio playback began and when the user interrupted. The local device sends the offset time and audio data representing the utterance to a speech processing system which can then use the offset time and stored data to identify which entry on the list was most recently output by the local device when the user interrupted. The system can then resolve anaphora to match that entry and can perform additional processing based on the referred to item.Type: GrantFiled: December 4, 2020Date of Patent: February 20, 2024Assignee: Amazon Technologies, Inc.Inventors: Prakash Krishnan, Arindam Mandal, Siddhartha Reddy Jonnalagadda, Nikko Strom, Ariya Rastrow, Ying Shi, David Chi-Wai Tang, Nishtha Gupta, Aaron Challenner, Bonan Zheng, Angeliki Metallinou, Vincent Auvray, Minmin Shen
-
Patent number: 11887583Abstract: Some devices may perform processing using machine learning models trained at a centralized system and distributed to the device. The centralized system may update the machine learning model and distribute the update to the device (or devices). To reduce the size of an update, the centralized system may train a model update object, which may be smaller in size than the model itself and thus more suitable for sending to the device(s). A device may receive the model update object and use it to update the on-device machine learning model; for example, by changing some parameters of the model. Parameters left unchanged during the update may retain their previous value. Thus, using the model update object to update the on-device model may result in a more accurate updated model when compared to sending an updated model compressed to a size similar to that of the model update object.Type: GrantFiled: June 9, 2021Date of Patent: January 30, 2024Assignee: Amazon Technologies, Inc.Inventors: Grant Strimel, Jonathan Jenner Macoskey, Ariya Rastrow
-
Publication number: 20240029743Abstract: Some speech processing systems may handle some commands on-device rather than sending the audio data to a second device or system for processing. The first device may have limited speech processing capabilities sufficient for handling common language and/or commands, while the second device (e.g., an edge device and/or a remote system) may call on additional language models, entity libraries, skill components, etc. to perform additional tasks. An intermediate data generator may facilitate dividing speech processing operations between devices by generating a stream of data that includes a first-pass ASR output (e.g., a word or sub-word lattice) and other characteristics of the audio data such as whisper detection, speaker identification, media signatures, etc. The second device can perform the additional processing using the data stream; e.g., without using the audio data. Thus, privacy may be enhanced by processing the audio data locally without sending it to other devices/systems.Type: ApplicationFiled: June 6, 2023Publication date: January 25, 2024Inventors: Stanislaw Ignacy Pasko, Pawel Zelazko, Cagdas Bak, Eli Joshua Fidler, Michal Kowalczuk, Andrew Oberlin, Ariya Rastrow
-
Publication number: 20230360633Abstract: Techniques for an interactive turn-based reading experience are described. A system may take turns reading content, such as a book, with a user. The system may process audio data representing a user reading a portion of the content, determine reading evaluation data, and determine how to proceed for the next turn based on the reading evaluation data. For example, based on the reading evaluation data, the system may read a portion of the content by outputting synthesized speech representing the content, may ask the user re-read a portion of the content, or may ask the user to read a different, smaller portion of the content.Type: ApplicationFiled: March 13, 2023Publication date: November 9, 2023Inventors: Kevin Crews, Prasanna H. Sridhar, Ariya Rastrow, Nicholas Matthew Jutila, Andrew Oberlin, Samarth Batra, Paul Anthony Bernhardt, Veerdhawal Pande, Roland Maximilian Rolf Maas
-
Patent number: 11721347Abstract: Some speech processing systems may handle some commands on-device rather than sending the audio data to a second device or system for processing. The first device may have limited speech processing capabilities sufficient for handling common language and/or commands, while the second device (e.g., an edge device and/or a remote system) may call on additional language models, entity libraries, skill components, etc. to perform additional tasks. An intermediate data generator may facilitate dividing speech processing operations between devices by generating a stream of data that includes a first-pass ASR output (e.g., a word or sub-word lattice) and other characteristics of the audio data such as whisper detection, speaker identification, media signatures, etc. The second device can perform the additional processing using the data stream; e.g., without using the audio data. Thus, privacy may be enhanced by processing the audio data locally without sending it to other devices/systems.Type: GrantFiled: June 29, 2021Date of Patent: August 8, 2023Assignee: Amazon Technologies, Inc.Inventors: Stanislaw Ignacy Pasko, Pawel Zelazko, Cagdas Bak, Eli Joshua Fidler, Michal Kowalczuk, Andrew Oberlin, Ariya Rastrow
-
Patent number: 11705116Abstract: Systems and methods described herein relate to adapting a language model for automatic speech recognition (ASR) for a new set of words. Instead of retraining the ASR models, language models and grammar models, the system only modifies one grammar model and ensures its compatibility with the existing models in the ASR system.Type: GrantFiled: August 18, 2021Date of Patent: July 18, 2023Assignee: Amazon Technologies, Inc.Inventors: Ankur Gandhe, Ariya Rastrow, Gautam Tiwari, Ashish Vishwanath Shenoy, Chun Chen