Patents by Inventor Quan Wang
Quan Wang has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Patent number: 12682906Abstract: A method of generating an accurate speaker representation for an audio sample includes receiving a first audio sample from a first speaker and a second audio sample from a second speaker. The method includes dividing a respective audio sample into a plurality of audio slices. The method also includes, based on the plurality of slices, generating a set of candidate acoustic embeddings where each candidate acoustic embedding includes a vector representation of acoustic features. The method further includes removing a subset of the candidate acoustic embeddings from the set of candidate acoustic embeddings. The method additionally includes generating an aggregate acoustic embedding from the remaining candidate acoustic embeddings in the set of candidate acoustic embeddings after removing the subset of the candidate acoustic embeddings.Type: GrantFiled: September 19, 2022Date of Patent: July 14, 2026Assignee: Google LLCInventors: Yeming Fang, Quan Wang, Pedro Jose Moreno Mengibar, Ignacio Lopez Moreno, Gang Feng, Fang Chu, Jin Shi, Jason William Pelecanos
-
Patent number: 12670912Abstract: A speaker verification method includes receiving audio data corresponding to an utterance, processing the audio data to generate a reference attentive d-vector representing voice characteristics of the utterance, the evaluation ad-vector includes ne style classes each including a respective value vector concatenated with a corresponding routing vector. The method also includes generating using a self-attention mechanism, at least one multi-condition attention score that indicates a likelihood that the evaluation ad-vector matches a respective reference ad-vector associated with a respective user. The method also includes identifying the speaker of the utterance as the respective user associated with the respective reference ad-vector based on the multi-condition attention score.Type: GrantFiled: October 2, 2023Date of Patent: June 30, 2026Assignee: Google LLCInventors: Ignacio Lopez Moreno, Quan Wang, Jason Pelecanos, Yiling Huang, Mert Saglam
-
Patent number: 12639464Abstract: Provided are a file leak detection method. The method includes: acquiring a file operation event on a terminal device, wherein the file operation event is an event in which a specified operation is executed on a target file; extracting, from the file operation event, a file path of the target file involved in the file operation event; searching for file content of the target file according to the file path, and performing mapping processing on the file content of the target file, so as to obtain a file fingerprint of the target file; determining, according to the file path and the file fingerprint, whether the target file belongs to a specified file library which is used for dynamically maintaining a service file that needs to be protected; and if the target file belongs to the specified file library, determining that a file in the specified file library is leaking.Type: GrantFiled: October 27, 2022Date of Patent: May 26, 2026Assignee: Beijing Bytedance Network Technology Co., Ltd.Inventors: Quan Wang, Xinci Liu, Jian Yang, Xianglong Ma
-
Publication number: 20260134193Abstract: A method includes receiving audio data characterizing an utterance spoken by a user. The method also includes processing the audio data to generate a transcription of the utterance using a multimodal large language model (LLM). The transcription includes a sequence of terms. The method also includes processing, using the multimodal LLM, the audio data and the transcription in parallel to identify one or more revision terms in the sequence of terms. The one or more revision terms specify a revision action to perform on at least on other term in the sequence of terms. The method also includes modifying the transcription based on the one or more revision terms.Type: ApplicationFiled: November 11, 2024Publication date: May 14, 2026Applicant: Google LLCInventors: Quan Wang, Francoise Beaufays, Bhuvana Ramabhadran, Zhong Meng, Neng Chen, Antoine Bruguier, Yanzhang He, Golan Pundak, Guanlong Zhao
-
Publication number: 20260105919Abstract: A method (500) includes receiving an input audio signal (122) that corresponds to utterances (120) spoken by multiple speakers. The method also includes processing the input audio to generate a transcription (200) of the utterances and a sequence of speaker turn tokens (224) each indicating a location of a respective speaker turn. The method also includes segmenting the input audio signal into a plurality of speaker segments (225) based on the sequence of speaker tokens. The method also includes extracting a speaker-discriminative embedding from each speaker segment and performing spectral clustering on the speaker-discriminative embeddings to cluster the plurality of speaker segments into k classes. The method also includes assigning a respective speaker label (250) to each speaker segment clustered into the respective class that is different than the respective speaker label assigned to the speaker segments clustered into each other class of the k classes.Type: ApplicationFiled: October 5, 2022Publication date: April 16, 2026Applicant: Google LLCInventors: Quan Wang, Yiling Huang, Han Lu, Guanlong Zhao
-
Patent number: 12603938Abstract: Disclosed is a method for controlling data transmission for smart gas, comprising: obtaining gas data of a gas user, and storing the gas data into a storage unit; predicting a probability of full storage of the storage unit within a preset future time period based on a historical data increment of the storage unit; in response to the probability of full storage meeting a preset probability condition: obtaining the gas data of the gas user within the preset time period; determining the upload demand degree based on the gas data; and in response to the upload demand degree meeting a preset condition, transmitting the gas data within the preset time period to the smart gas management platform based on the smart gas object platform at a target upload time point of at least one smart gas meter of the gas user.Type: GrantFiled: August 27, 2024Date of Patent: April 14, 2026Assignee: CHENGDU QINCHUAN IOT TECHNOLOGY CO., LTD.Inventors: Zehua Shao, Junyan Zhou, Xiaojun Wei, Lei He, Quan Wang
-
Publication number: 20260094600Abstract: A method includes receiving a prompt directed towards an assistant large language model (LLM) and generating, using the assistant LLM, a sequence of output tokens based on the prompt. The sequence of output tokens includes a sequence of textual tokens including one or more correct textual tokens and one or more incorrect textual tokens, and one or more revision tokens each indicating a corresponding N number of incorrect textual tokens generated prior to the respective revision token and corresponding replacement textual tokens generated after the respective revision token for replacement of the corresponding N number of incorrect textual tokens. The method also includes generating a revised sequence of output tokens for the prompt based on the sequence of output tokens.Type: ApplicationFiled: September 22, 2025Publication date: April 2, 2026Applicant: Google LLCInventors: Quan Wang, Fadi Biadsy, Yonghui Xiao, Youzheng Chen
-
Publication number: 20260085619Abstract: A variable geometry turbine device includes: a turbine casing provided with an inlet duct, and defining an exhaust port and a bypass chamber therein; a gear rotating sleeve disposed inside the bypass chamber; a rack actuator; combined nozzle guide vanes including: a fixed nozzle and an elastic nozzle mechanism; and a turbine. A first annular chamber and a second annular chamber are defined between the exhaust port and the bypass chamber. A first radial channel and a second radial channel are defined under the bypass chamber, and are connected to the first annular chamber and the second annular chamber respectively. A side of the gear rotating sleeve defines a first notch and a second notch configured to be selectively connected to the first radial channel and the second radial channel respectively; and opening and closing of the first radial channel and the second radial channel have a time difference.Type: ApplicationFiled: May 24, 2025Publication date: March 26, 2026Inventors: Quan LIU, Chao MA, Jianjun HUANG, Xinyuan XIONG, Zhanhao CHEN, Quan WANG, Yangli SUN
-
Publication number: 20260073923Abstract: A method includes receiving an input audio signal that corresponds to utterances spoken by multiple speakers. The method also includes processing the input audio to generate a transcription of the utterances and a sequence of speaker turn tokens each indicating a location of a respective speaker turn. The method also includes segmenting the input audio signal into a plurality of speaker segments based on the sequence of speaker tokens. The method also includes extracting a speaker-discriminative embedding from each speaker segment and performing spectral clustering on the speaker-discriminative embeddings to cluster the plurality of speaker segments into k classes. The method also includes assigning a respective speaker label to each speaker segment clustered into the respective class that is different than the respective speaker label assigned to the speaker segments clustered into each other class of the k classes.Type: ApplicationFiled: November 13, 2025Publication date: March 12, 2026Applicant: Google LLCInventors: Quan Wang, Han Lu, Evan Clark, Ignacio Lopez Moreno, Hasim Sak, Wei Xia, Taral Joglekar, Anshuman Tripathi
-
Publication number: 20260051318Abstract: A method includes receiving training utterances that include non-synthetic speech training utterances and synthetic speech utterances. For each training utterance, the method includes processing, using a memorized neural network, a corresponding sequence of input audio frames to generate a hotword detection output indicating a likelihood the training utterance includes a hotword, determining a first loss based on the hotword detection output, obtaining a hidden layer feature vector for each corresponding input audio frame; processing, using a speech classification model, the hidden layer feature vectors to predict a classification output for the training utterance; and determining an adversarial loss based on the classification output predicted for the training utterance.Type: ApplicationFiled: August 7, 2025Publication date: February 19, 2026Applicant: GDM Holding LLCInventors: Hyun Jin Park, Kurt Edward Partridge, Quan Wang
-
Publication number: 20260050790Abstract: During a first prompt session, a method includes receiving a first prompt specifying a task for a language model (LM). For each biased attention layer of the LM, the method also includes: computing, based on the first prompt, a set of attention weights; and computing bias parameters for biasing a subsequent computation of the set of attention weights during a second prompt session. During the second prompt session, the method also includes receiving a second prompt specifying another task for the LM. For each biased attention layer, the method also includes: computing, based on the second prompt, the set of attention weights; and biasing, using the bias parameters computed during the first prompt session, the set of attention weights. The method also includes generating a corresponding response based on the biased sets of attention weights.Type: ApplicationFiled: August 4, 2025Publication date: February 19, 2026Applicant: GDM Holding LLCInventor: Quan Wang
-
Publication number: 20260045268Abstract: A method includes receiving audio data characterizing a conversation between two or more speakers. The method also includes generating a sequence of audio features based on the audio data. For each output step of a plurality of output steps, the method includes generating a corresponding set of embeddings for the corresponding audio features, selecting a subset of the embeddings for the corresponding output step from the corresponding set of embeddings, and predicting a respective voice activity indicator for each respective speaker of the two or more speakers based on the subset of the embeddings selected for corresponding output step. The respective voice activity indicator indicates whether a voice of the respective speaker is active or inactive at the corresponding output step.Type: ApplicationFiled: August 7, 2024Publication date: February 12, 2026Applicant: Google LLCInventors: Jason Pelecanos, Yiling Huang, Quan Wang
-
Publication number: 20260038507Abstract: Techniques are disclosed that enable processing of audio data to generate one or more refined versions of audio data, where each of the refined versions of audio data isolate one or more utterances of a single respective human speaker. Various implementations generate a refined version of audio data that isolates utterance(s) of a single human speaker by processing a spectrogram representation of the audio data (generated by processing the audio data with a frequency transformation) using a mask generated by processing the spectrogram of the audio data and a speaker embedding for the single human speaker using a trained voice filter model. Output generated over the trained voice filter model is processed using an inverse of the frequency transformation to generate the refined audio data.Type: ApplicationFiled: October 8, 2025Publication date: February 5, 2026Inventors: Quan Wang, Prashant Sridhar, Ignacio Lopez Moreno, Hannah Muckenhirn
-
Patent number: 12518762Abstract: A method includes obtaining a multi-utterance training sample that includes audio data characterizing utterances spoken by two or more different speakers and obtaining ground-truth speaker change intervals indicating time intervals in the audio data where speaker changes among the two or more different speakers occur. The method also includes processing the audio data to generate a sequence of predicted speaker change tokens using a sequence transduction model. For each corresponding predicted speaker change token, the method includes labeling the corresponding predicted speaker change token as correct when the predicted speaker change token overlaps with one of the ground-truth speaker change intervals. The method also includes determining a precision metric of the sequence transduction model based on a number of the predicted speaker change tokens labeled as correct and a total number of the predicted speaker change tokens in the sequence of predicted speaker change tokens.Type: GrantFiled: October 9, 2023Date of Patent: January 6, 2026Assignee: Google LLCInventors: Guanlong Zhao, Quan Wang, Han Lu, Yiling Huang, Jason Pelecanos
-
Publication number: 20250378286Abstract: A method (500) includes receiving, from an application (50) executing on a client device (110), at a speech service interface (200), configuration parameters (211) for integrating a speech service (250) into the application. The configuration parameters include a language pack directory (225) that maps a primary language code (235) to an on-device path of a primary language pack (110) of the speech service for use in recognizing speech in a primary language and each of one or more codeswitch language codes to an on-device path. The method also includes receiving audio data (102) characterizing an utterance (106) and processing, using a language ID predictor model (230), the audio data to determine that the audio data is associated with the primary language code. The method also includes processing, using the primary language pack, the audio data to determine a transcription (120) that includes one or more words in the primary language.Type: ApplicationFiled: November 23, 2022Publication date: December 11, 2025Applicant: Google LLCInventors: Quan Wang, Evan Clark, Yang Yu, Han Lu, Taral Pradeep Joglekar, Qi Cao, Dharmeshkumar Mokani, Diego Melendo Casado, Ignacio Lopez Moreno, Hasim Sak
-
Publication number: 20250378206Abstract: The present disclosure relates to the technical field of data security, and discloses a security protection method for a drawing file, an apparatus, a device, a medium, and a program product. The security protection method for a drawing file comprises: acquiring a first drawing file; acquiring a matching result of matching first feature information corresponding to the first drawing file with each second feature information in a target feature information set, where the first feature information is obtained by parsing the first drawing file to obtain description information of a target element in the first drawing file, and performing feature extraction on the obtained description information of the target element, and the target feature information set includes second feature information of at least one protected drawing file; and performing security protection management on the first drawing file based on the matching result.Type: ApplicationFiled: March 24, 2025Publication date: December 11, 2025Inventors: Quan WANG, Xianglong MA, Huanxun CHEN, Chenhui CAI, Haoxiang CHENG
-
Patent number: 12494206Abstract: A method includes obtaining a speaker identification (SID) model trained to predict speaker embeddings from utterances spoken by different speakers, the SID model includes a trained audio encoder and a trained SID head. The method also includes receiving a plurality of synthetic speech detection (SSD) training utterances that include a set of human-originated speech samples and a set of synthetic speech samples. The method also includes training, using the trained audio encoder, a SSD head on the SSD training utterances to learn to detect the presence of synthetic speech in audio encodings encoded by the trained audio encoder. The operations also include providing, for execution on a computing device, a multitask neural network model for performing both SID tasks and SSD tasks on input audio data in parallel.Type: GrantFiled: February 10, 2023Date of Patent: December 9, 2025Assignee: Google LLCInventors: Alanna Foster Slocum, Yiling Huang, Shelly Bensal, Quan Wang
-
Patent number: 12482470Abstract: A method includes receiving an input audio signal that corresponds to utterances spoken by multiple speakers. The method also includes processing the input audio to generate a transcription of the utterances and a sequence of speaker turn tokens each indicating a location of a respective speaker turn. The method also includes segmenting the input audio signal into a plurality of speaker segments based on the sequence of speaker tokens. The method also includes extracting a speaker-discriminative embedding from each speaker segment and performing spectral clustering on the speaker-discriminative embeddings to cluster the plurality of speaker segments into k classes. The method also includes assigning a respective speaker label to each speaker segment clustered into the respective class that is different than the respective speaker label assigned to the speaker segments clustered into each other class of the k classes.Type: GrantFiled: December 14, 2021Date of Patent: November 25, 2025Assignee: Google LLCInventors: Quan Wang, Han Lu, Evan Clark, Ignacio Lopez Moreno, Hasim Sak, Wei Xia, Taral Joglekar, Anshuman Tripathi
-
Patent number: 12470399Abstract: A blockchain system for ownership verification may include one or more issuer network nodes and one or more verification network nodes. An issuer network node may be configured to receive a request including a public key to issue a credential, provision the credential to the communication device, generate a payload derived from hashing the credential and the public key, store the payload in a record of a blockchain, and synchronize the record to other network nodes on the blockchain. A verification network node may be configured to receive the credential, the public key, and a signature generated by the communication device to request access to a resource, verify the signature using the public key, generate a hash value based on the credential and the public key, determine that the hash value is stored in the blockchain, and authenticate the communication device for access to the requested resource.Type: GrantFiled: June 17, 2022Date of Patent: November 11, 2025Assignee: Visa International Service AssociationInventors: Mohamed Nosseir, Anil Somani, Quan Wang
-
Patent number: 12469498Abstract: Techniques are disclosed that enable processing of audio data to generate one or more refined versions of audio data, where each of the refined versions of audio data isolate one or more utterances of a single respective human speaker. Various implementations generate a refined version of audio data that isolates utterance(s) of a single human speaker by processing a spectrogram representation of the audio data (generated by processing the audio data with a frequency transformation) using a mask generated by processing the spectrogram of the audio data and a speaker embedding for the single human speaker using a trained voice filter model. Output generated over the trained voice filter model is processed using an inverse of the frequency transformation to generate the refined audio data.Type: GrantFiled: March 4, 2024Date of Patent: November 11, 2025Assignee: GOOGLE LLCInventors: Quan Wang, Prashant Sridhar, Ignacio Lopez Moreno, Hannah Muckenhim