Patents by Inventor Ke Hu
Ke Hu has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Publication number: 20260170236Abstract: Disclosed are apparatuses, systems, and techniques implementing a language error correction system that leverages shared expertise of an ensemble of editor models. The techniques include processing, using a router network, an input representation of a text generated using text generation tool(s) to obtain an ensemble of scores, an individual score characterizing a degree of correspondence of the text to a field of expertise of a respective editor model of the ensemble. The techniques further include processing, using a plurality of editor models of the ensemble, the input representation of the text to generate a plurality of output representations of the text. The techniques further include generating a final representation of the text using a weighted combination of the plurality of output representations of the text, an individual output representation weighted using a score, of the ensemble of scores, obtained for a respective editor model of the plurality of editor models.Type: ApplicationFiled: December 18, 2024Publication date: June 18, 2026Inventors: Chao-Han Huck Yang, Yen-Ting Lin, Zhehuai Chen, Piotr Zelasko, Xuesong Yang, Venkata Naga Krishna Chaitanya Puvvada, Szu-Wei Fu, Ke Hu, Jagadeesh Balam, Boris Ginsburg, Yu-Chiang Frank Wang
-
Publication number: 20260162656Abstract: A method of a multilingual ASR model includes receiving a sequence of acoustic frames characterizing an utterance of speech. At a plurality of output steps, the method further includes generating a first higher order feature representation for an acoustic frame by a first encoder that includes a first plurality of multi-head attention layers; generating a second higher order feature representation for a corresponding first higher order feature representation by a second encoder that includes a second plurality of multi-head attention layers; and generating, by a first decoder, a first probability distribution over possible speech recognition hypotheses based on the second higher order feature representation and a sequence of N previous non-blank symbols. A gating layer of each respective MoE layer configured to dynamically route an output from a previous multi-head attention layer at each of the plurality of output steps to a respective pair of feed-forward expert networks.Type: ApplicationFiled: January 28, 2026Publication date: June 11, 2026Applicant: Google LLCInventors: Ke Hu, Bo Li, Tara N. Sainath, Yu Zhang, Françoise Beaufays
-
Publication number: 20260120690Abstract: Disclosed are apparatuses, systems, and techniques that use one or more artificial intelligence models for time-aligned automatic speech recognition (ASR) of speech. The techniques include processing, an ASR model, one or more audio frames representative of a speech to generate, for a transcription unit (TU) of the speech a first set of likelihood values and a second set of likelihood values. An individual likelihood value of the first set characterizes a probability that the TU corresponds to a vocabulary token. An individual likelihood value of the second set characterizes a probability that the TU corresponds to a timestamp token. The techniques further include generating, using the first set of likelihood values and the second set of likelihood values, a timed transcription of the speech.Type: ApplicationFiled: October 29, 2024Publication date: April 30, 2026Inventors: Ke Hu, Venkata Naga Krishna Chaitanya Puvvada, Jagadeesh Balam, Elena Sergeevna Rastorgueva, Boris Ginsburg
-
Publication number: 20260080191Abstract: Disclosed are apparatuses, systems, and techniques that implement training and deployment of automatic transcription-assisted translation systems that use language models. The techniques include processing, using a first speech-to-text (S2T) model, a first input that includes a speech in a first language to generate a transcription of the speech. The techniques further include processing, using a second S2T model, a second input to generate a translation of the speech to a second language. The second input includes at least a representation of the speech, and the transcription of the speech.Type: ApplicationFiled: March 17, 2025Publication date: March 19, 2026Inventors: Ke Hu, Zhehuai Chen, Chao-Han Huck Yang, Piotr Zelasko, Oleskii Hrinchuk, Vitaly Lavrukhin, Jagadeesh Balam, Boris Ginsburg
-
Patent number: 12555573Abstract: A method of a multilingual ASR model includes receiving a sequence of acoustic frames characterizing an utterance of speech. At a plurality of output steps, the method further includes generating a first higher order feature representation for an acoustic frame by a first encoder that includes a first plurality of multi-head attention layers; generating a second higher order feature representation for a corresponding first higher order feature representation by a second encoder that includes a second plurality of multi-head attention layers; and generating, by a first decoder, a first probability distribution over possible speech recognition hypotheses based on the second higher order feature representation and a sequence of N previous non-blank symbols. A gating layer of each respective MoE layer configured to dynamically route an output from a previous multi-head attention layer at each of the plurality of output steps to a respective pair of feed-forward expert networks.Type: GrantFiled: March 7, 2024Date of Patent: February 17, 2026Assignee: Google LLCInventors: Ke Hu, Bo Li, Tara N. Sainath, Yu Zhang, Francoise Beaufays
-
Publication number: 20250308512Abstract: A method of text-only and semi-supervised training for deliberation includes receiving training data including unspoken textual utterances that are each not paired with any corresponding spoken utterance of non-synthetic speech, and training a deliberation model that includes a text encoder and a deliberation decoder on the unspoken textual utterances. The method also includes receiving, at the trained deliberation model, first-pass hypotheses and non-causal acoustic embeddings. The first-pass hypotheses is generated by a recurrent neural network-transducer (RNN-T) decoder for the non-causal acoustic embeddings encoded by a non-causal encoder. The method also includes encoding, using the text encoder, the first-pass hypotheses generated by the RNN-T decoder, and generating, using the deliberation decoder attending to both the first-pass hypotheses and the non-causal acoustic embeddings, second-pass hypotheses.Type: ApplicationFiled: June 13, 2025Publication date: October 2, 2025Applicant: Google LLCInventors: Ke Hu, Yanzhang He, Weiran Wang, Tara N. Sainath, Trevor Strohman, Rohit Prabhavalkar, Sepand Mavandadi
-
Patent number: 12354595Abstract: A method of text-only and semi-supervised training for deliberation includes receiving training data including unspoken textual utterances that are each not paired with any corresponding spoken utterance of non-synthetic speech, and training a deliberation model that includes a text encoder and a deliberation decoder on the unspoken textual utterances. The method also includes receiving, at the trained deliberation model, first-pass hypotheses and non-causal acoustic embeddings. The first-pass hypotheses is generated by a recurrent neural network-transducer (RNN-T) decoder for the non-causal acoustic embeddings encoded by a non-causal encoder. The method also includes encoding, using the text encoder, the first-pass hypotheses generated by the RNN-T decoder, and generating, using the deliberation decoder attending to both the first-pass hypotheses and the non-causal acoustic embeddings, second-pass hypotheses.Type: GrantFiled: March 18, 2023Date of Patent: July 8, 2025Assignee: Google LLCInventors: Ke Hu, Tara N. Sainath, Yanzhang He, Rohit Prabhavalkar, Sepand Mavandadi, Weiran Wang, Trevor Strohman
-
Publication number: 20250095637Abstract: A method includes receiving a textual prompt in a first language and obtaining a fine-tuned prompt embedding configured to guide a large language model (LLM) to generate text in a target language from textual prompts in the first language. The method also includes processing, using the LLM, the textual prompt conditioned on the fine-tuned prompt embedding to generate output text in the target language and concatenating the textual prompt and the generated output text to provide an unspoken textual utterance. The method also includes training a multilingual automatic speech recognition (ASR) model to learn how to recognize speech in the target language by injecting the unspoken textual utterance into a text encoder associated with the multilingual ASR model.Type: ApplicationFiled: September 16, 2024Publication date: March 20, 2025Applicant: Google LLCInventors: Ke Hu, Tara N. Sainath, Bo Li, Yu Zhang, Yong Cheng, Tao Wang, Yujing Zhang, Frederick Liu
-
Publication number: 20240428786Abstract: A method includes receiving a sequence of acoustic frames and generating, by a first encoder, a first higher order feature representation for a corresponding acoustic frame in the sequence of acoustic frames. The method also includes generating, by a first pass transducer decoder, a first pass speech recognition hypothesis for a corresponding first higher order feature representation and generating, by a text encoder, a text encoding for a corresponding first pass speech recognition hypothesis. The method also includes generating, by a second encoder, a second higher order feature representation for a corresponding first higher order feature representation. The method also includes generating, by a second pass transducer decoder, a second pass speech recognition hypothesis using a corresponding second higher order feature representation and a corresponding text encoding.Type: ApplicationFiled: September 6, 2024Publication date: December 26, 2024Applicant: Google LLCInventors: Ke Hu, Tara N. Sainath, Arun Narayanan, Ruoming Pang, Trevor Strohman
-
Patent number: 12118988Abstract: A method includes receiving a sequence of acoustic frames and generating, by a first encoder, a first higher order feature representation for a corresponding acoustic frame in the sequence of acoustic frames. The method also includes generating, by a first pass transducer decoder, a first pass speech recognition hypothesis for a corresponding first higher order feature representation and generating, by a text encoder, a text encoding for a corresponding first pass speech recognition hypothesis. The method also includes generating, by a second encoder, a second higher order feature representation for a corresponding first higher order feature representation. The method also includes generating, by a second pass transducer decoder, a second pass speech recognition hypothesis using a corresponding second higher order feature representation and a corresponding text encoding.Type: GrantFiled: September 19, 2022Date of Patent: October 15, 2024Assignee: Google LLCInventors: Ke Hu, Tara N. Sainath, Arun Narayanan, Ruoming Pang, Trevor Strohman
-
Publication number: 20240304185Abstract: A method of a multilingual ASR model includes receiving a sequence of acoustic frames characterizing an utterance of speech. At a plurality of output steps, the method further includes generating a first higher order feature representation for an acoustic frame by a first encoder that includes a first plurality of multi-head attention layers; generating a second higher order feature representation for a corresponding first higher order feature representation by a second encoder that includes a second plurality of multi-head attention layers; and generating, by a first decoder, a first probability distribution over possible speech recognition hypotheses based on the second higher order feature representation and a sequence of N previous non-blank symbols. A gating layer of each respective MoE layer configured to dynamically route an output from a previous multi-head attention layer at each of the plurality of output steps to a respective pair of feed-forward expert networks.Type: ApplicationFiled: March 7, 2024Publication date: September 12, 2024Applicant: Google LLCInventors: Ke Hu, Bo Li, Tara N. Sainath, Yu Zhang, Francoise Beaufays
-
Publication number: 20240250500Abstract: A monolithic integrated multi-segment cascade optical frequency comb and its chip are disclosed, which belongs to the technical field of sensing detection, quantum information and optical communication technology. The optical frequency comb includes a first semiconductor passive mode-locking laser, a semiconductor optical amplifier, and a second semiconductor passive mode-locking laser sequentially integrated and connected; the first semiconductor passive mode-locking laser includes a first reverse bias absorption area integrated with an optogalvanic distribution grating and a first gain cavity length extender (coupled multi-ring or multi-disk); the second semiconductor passive mode-locking laser includes a second reverse bias absorption area and a second gain cavity length extender (coupled multi-ring or multi-disk); each structure is connected to each other by electrical isolation grooves.Type: ApplicationFiled: April 2, 2024Publication date: July 25, 2024Inventors: Zhongliang Qiao, Yi Qu, Haoran Wan, Wenjun Yu, Dengqun Weng, Xiaohu Hou, Menghao Wu, Ke Hu, Zhibin Zhao, Hao Chen, Zaijin Li, Lina Zeng, Lin Li, Guojun Liu
-
Patent number: 12027158Abstract: A method of performing speech recognition using a two-pass deliberation architecture includes receiving a first-pass hypothesis and an encoded acoustic frame and encoding the first-pass hypothesis at a hypothesis encoder. The first-pass hypothesis is generated by a recurrent neural network (RNN) decoder model for the encoded acoustic frame. The method also includes generating, using a first attention mechanism attending to the encoded acoustic frame, a first context vector, and generating, using a second attention mechanism attending to the encoded first-pass hypothesis, a second context vector. The method also includes decoding the first context vector and the second context vector at a context vector decoder to form a second-pass hypothesis.Type: GrantFiled: February 6, 2023Date of Patent: July 2, 2024Assignee: Google LLCInventors: Ke Hu, Tara N. Sainath, Ruoming Pang, Rohit Prakash Prabhavalkar
-
Patent number: 11942076Abstract: A method includes receiving audio data encoding an utterance spoken by a native speaker of a first language, and receiving a biasing term list including one or more terms in a second language different than the first language. The method also includes processing, using a speech recognition model, acoustic features derived from the audio data to generate speech recognition scores for both wordpieces and corresponding phoneme sequences in the first language. The method also includes rescoring the speech recognition scores for the phoneme sequences based on the one or more terms in the biasing term list, and executing, using the speech recognition scores for the wordpieces and the rescored speech recognition scores for the phoneme sequences, a decoding graph to generate a transcription for the utterance.Type: GrantFiled: February 16, 2022Date of Patent: March 26, 2024Assignee: Google LLCInventors: Ke Hu, Golan Pundak, Rohit Prakash Prabhavalkar, Antoine Jean Bruguier, Tara N. Sainath
-
Patent number: 11908461Abstract: A method of performing speech recognition using a two-pass deliberation architecture includes receiving a first-pass hypothesis and an encoded acoustic frame and encoding the first-pass hypothesis at a hypothesis encoder. The first-pass hypothesis is generated by a recurrent neural network (RNN) decoder model for the encoded acoustic frame. The method also includes generating, using a first attention mechanism attending to the encoded acoustic frame, a first context vector, and generating, using a second attention mechanism attending to the encoded first-pass hypothesis, a second context vector. The method also includes decoding the first context vector and the second context vector at a context vector decoder to form a second-pass hypothesis.Type: GrantFiled: January 14, 2021Date of Patent: February 20, 2024Assignee: Google LLCInventors: Ke Hu, Tara N. Sainath, Ruoming Pang, Rohit Prakash Prabhavalkar
-
Publication number: 20230298563Abstract: A method of text-only and semi-supervised training for deliberation includes receiving training data including unspoken textual utterances that are each not paired with any corresponding spoken utterance of non-synthetic speech, and training a deliberation model that includes a text encoder and a deliberation decoder on the unspoken textual utterances. The method also includes receiving, at the trained deliberation model, first-pass hypotheses and non-causal acoustic embeddings. The first-pass hypotheses is generated by a recurrent neural network-transducer (RNN-T) decoder for the non-causal acoustic embeddings encoded by a non-causal encoder. The method also includes encoding, using the text encoder, the first-pass hypotheses generated by the RNN-T decoder, and generating, using the deliberation decoder attending to both the first-pass hypotheses and the non-causal acoustic embeddings, second-pass hypotheses.Type: ApplicationFiled: March 18, 2023Publication date: September 21, 2023Applicant: Google LLCInventors: Ke Hu, Tara N. Sainath, Yanzhang He, Rohit Prabhavalkar, Sepand Mavandadi, Weiran Wang, Trevor Strohman
-
Publication number: 20230186907Abstract: A method of performing speech recognition using a two-pass deliberation architecture includes receiving a first-pass hypothesis and an encoded acoustic frame and encoding the first-pass hypothesis at a hypothesis encoder. The first-pass hypothesis is generated by a recurrent neural network (RNN) decoder model for the encoded acoustic frame. The method also includes generating, using a first attention mechanism attending to the encoded acoustic frame, a first context vector, and generating, using a second attention mechanism attending to the encoded first-pass hypothesis, a second context vector.Type: ApplicationFiled: February 6, 2023Publication date: June 15, 2023Applicant: Google LLCInventors: Ke Hu, Tara N. Sainath, Ruoming Pang, Rohit Prakash Prabhavalkar
-
Publication number: 20230107248Abstract: A method includes receiving an initial alignment for a candidate hypothesis generated by a transducer decoder model during a first pass. Here, the candidate hypothesis corresponds to a candidate transcription for an utterance and the initial alignment for the candidate hypothesis includes a sequence of output labels. Each output label corresponds to a blank symbol or a hypothesized sub-word unit. The method also include receiving a subsequent sequence of audio encodings characterizing the utterance. During an initial refinement step, the method also includes generating a new alignment for a rescored sequence of output labels using a non-autoregressive decoder. The non-autoregressive decoder is configured to receive the initial alignment for the candidate hypothesis and the subsequent sequence of audio encodings.Type: ApplicationFiled: September 16, 2022Publication date: April 6, 2023Applicant: Google LLCInventors: Weiran Wang, Ke Hu, Tara N. Sainath
-
Publication number: 20230109407Abstract: A method includes receiving a sequence of acoustic frames and generating, by a first encoder, a first higher order feature representation for a corresponding acoustic frame in the sequence of acoustic frames. The method also includes generating, by a first pass transducer decoder, a first pass speech recognition hypothesis for a corresponding first higher order feature representation and generating, by a text encoder, a text encoding for a corresponding first pass speech recognition hypothesis. The method also includes generating, by a second encoder, a second higher order feature representation for a corresponding first higher order feature representation. The method also includes generating, by a second pass transducer decoder, a second pass speech recognition hypothesis using a corresponding second higher order feature representation and a corresponding text encoding.Type: ApplicationFiled: September 19, 2022Publication date: April 6, 2023Applicant: Google LLCInventors: Ke Hu, Tara N. Sainath, Arun Narayanan, Ruoming Pang, Trevor Strohman
-
Patent number: 11610586Abstract: A method includes receiving a speech recognition result, and using a confidence estimation module (CEM), for each sub-word unit in a sequence of hypothesized sub-word units for the speech recognition result: obtaining a respective confidence embedding that represents a set of confidence features; generating, using a first attention mechanism, a confidence feature vector; generating, using a second attention mechanism, an acoustic context vector; and generating, as output from an output layer of the CEM, a respective confidence output score for each corresponding sub-word unit based on the confidence feature vector and the acoustic feature vector received as input by the output layer of the CEM. For each of the one or more words formed by the sequence of hypothesized sub-word units, the method also includes determining a respective word-level confidence score for the word. The method also includes determining an utterance-level confidence score by aggregating the word-level confidence scores.Type: GrantFiled: February 23, 2021Date of Patent: March 21, 2023Assignee: Google LLCInventors: David Qiu, Qiujia Li, Yanzhang He, Yu Zhang, Bo Li, Liangliang Cao, Rohit Prabhavalkar, Deepti Bhatia, Wei Li, Ke Hu, Tara Sainath, Ian Mcgraw