Patents by Inventor Yuhui Chen
Yuhui Chen has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Patent number: 12684087Abstract: Systems and methods for active speaker detection for videoconferencing are provided. For example, an audio recording system can record a set of raw audio signals. The set of raw audio signals include audio signals recorded by a set of microphones at different distances from an audio source with different orientations. An audio dataset generation system can access a virtual meeting room setup which specifies microphones used in a virtual meeting room, locations of speakers and the microphones, and orientations of the speakers. The audio dataset generation system generates a synthetic audio signal for each of the speakers specified in the virtual meeting room setup by combining audio signals selected from the set of raw audio signals according to the virtual meeting room setup.Type: GrantFiled: January 8, 2024Date of Patent: July 14, 2026Assignee: Zoom Communications, Inc.Inventors: Yuhui Chen, Qiang Gao, Zhaofeng Jia, Xian Tong, Ye Wang
-
Patent number: 12671958Abstract: Systems and methods for synthetic audio datasets generation for videoconferencing based on realistic room acoustic simulation are provided. For example, a room-independent recording for a source audio signal and a particular type of microphones can be generated and a target room setup for a target room can be obtained. The target room setup specifying one or more of a size of the target room, respective locations of a speaker and a microphone in the target room. Room characteristics of the target room can be generated based on the target room setup via a room acoustic model. A synthetic audio signal for the target room and the particular type of microphones can be generated by applying the room acoustic characteristics onto the room-independent recording.Type: GrantFiled: January 29, 2024Date of Patent: June 30, 2026Assignee: Zoom Communications, Inc.Inventors: Yuhui Chen, Yifeng Fan, Zhaofeng Jia
-
Patent number: 12645422Abstract: Techniques for adaptive audio processing during video conferencing are provided. In an example method, a client device joins a video conference hosted by a video conference provider, the video conference including a plurality of client devices. The client device receives, from an audio input device, an audio stream. The client device then processes, using an audio processing component, the audio stream. The client device determines, using a trained machine learning model, one or more characteristics of the audio stream. The client device then determines, based on the one or more characteristics of the audio stream, an audio configuration operation comprising one or more instructions. The client device executes the audio configuration operation. The client device then outputs the audio stream to the video conference provider.Type: GrantFiled: October 3, 2023Date of Patent: June 2, 2026Assignee: Zoom Communications, Inc.Inventors: Yuhui Chen, Qiang Gao, Zhaofeng Jia, Shiwei Wang
-
Patent number: 12626709Abstract: Methods, systems, and apparatus, including computer programs encoded on computer storage media relate to a method for audio super resolution. The system receives an audio signal. When the sampling rate of the audio signal is below a sampling rate threshold or the frequency range of the audio signal is below a frequency range threshold, the audio signal is input to an audio super resolution model comprising a machine learning model. The audio signal is processed by the audio super resolution model to generate a synthetic audio signal with a wider frequency range than the frequency range of the audio signal.Type: GrantFiled: October 31, 2021Date of Patent: May 12, 2026Assignee: Zoom Communications, Inc.Inventors: Yuhui Chen, Zhaofeng Jia, Qiyong Liu, Zhengwei Wei
-
Publication number: 20260099974Abstract: Systems and methods for personalized realistic video generation. In one example, a client device joins a video conference. The client device accesses a source video clip including a set of source video frames related to a user associated with the client device. The client device receives source audio data related to the user. The client device generates target video data based on the set of source video frames and the source audio data using a trained video generator model. The client device streams the target video data during the video conference.Type: ApplicationFiled: December 4, 2024Publication date: April 9, 2026Applicant: Zoom Video Communications, Inc.Inventors: Yuanqi Chen, Yuhui Chen, Dewang Hou, Bo Ling, Gengdai Liu, Liubin Liu, Matthieu Tardivel, Rui Zhang, Yian Zhu
-
Publication number: 20260046362Abstract: Example methods and systems provide machine-learning assisted acoustic echo cancellation (AEC). The AEC can be used, as an example, to improve the audio quality for online audio and video conferences. A system according to this disclosure includes a pre-trained, machine-learning, AI model designed to detect, in real time, a unitary voice signal, or a signal representing the speech of a single speaker as opposed to that of multiple speakers. A digital signal processing (DSP) algorithm can then detect the echo state, for example, whether distortion results primarily from an echo. Based on these characteristics, the system can, alternatively and automatically apply either a default mode of AEC to the audio signal, or apply a more aggressive mode of AEC.Type: ApplicationFiled: October 21, 2025Publication date: February 12, 2026Applicant: Zoom Communications, Inc.Inventors: Yuhui Chen, Zhaofeng Jia, Wei Wang
-
Publication number: 20260031097Abstract: Techniques for low-quality audio detection are provided. In an example method, a computing system joins a first client device to a first video conference including a number of connected client devices. The computing system receives, from the first client device, a first audio stream. The computing system determines, using a machine learning (“ML”) model, at least one first audio quality measurement based on the first audio stream. The computing system computes a first metric for the first audio stream based on the at least one first audio quality measurement. In response to the first metric satisfying a predetermined threshold, the computing system outputs a message including first information about a first low-quality audio status associated with the first audio stream.Type: ApplicationFiled: July 25, 2024Publication date: January 29, 2026Applicant: Zoom Video Communications, Inc.Inventors: Yuhui Chen, Zhiye Chen, Zhaofeng Jia, Shiwei Wang
-
Patent number: 12477070Abstract: Example methods and systems provide machine-learning assisted acoustic echo cancellation (AEC). The AEC can be used, as an example, to improve the audio quality for online audio and video conferences. A system according to this disclosure includes a pre-trained, machine-learning, AI model designed to detect, in real time, a unitary voice signal, or a signal representing the speech of a single speaker as opposed to that of multiple speakers. A digital signal processing (DSP) algorithm can then detect the echo state, for example, whether distortion results primarily from an echo. Based on these characteristics, the system can, alternatively and automatically apply either a default mode of AEC to the audio signal, or apply a more aggressive mode of AEC.Type: GrantFiled: November 2, 2023Date of Patent: November 18, 2025Assignee: Zoom Communications, Inc.Inventors: Yuhui Chen, Zhaofeng Jia, Wei Wang
-
Publication number: 20250201231Abstract: Systems and methods for generating speaker video and audio in multiple languages for videoconferencing are provided. For example, a computing device can access a speaker speech audio signal that includes a speaker speech in a first language, a video of the speaker and a translated speech audio signal of the speaker speech in a second language. The computing device generates, based on the translated speech audio signal, a converted translated speech audio signal that includes a speech in the second language having voice characteristics in the speaker speech. The computing device further generates a lip-synched speaker video based on the video of the speaker and the converted translated speech audio signal. Lip movements in the lip-synched speaker video correspond to the converted translated speech audio signal. The converted translated speech audio signal and the lip-synched speaker video are transmitted to a video conference provider configured to host the video conference.Type: ApplicationFiled: December 18, 2023Publication date: June 19, 2025Inventors: Yuhui Chen, Qiang Gao, Zhaofeng Jia
-
Publication number: 20250193130Abstract: Embodiments of the present disclosure provide a network traffic flow control method and apparatus based on a distributed storage system, an electrical device, and a storage medium. The method includes: acquiring a transmission link for transmitting target data in response to a first data server in the distributed storage system receiving a transmission request for the target data; configuring a first priority parameter for the transmission link by the first data server according to the priority information of the business type to which the target data belongs; and in response to transmitting the target data through the transmission link, configuring a second priority parameter for the target data by the first data server according to the first priority parameter of the transmission link, so as to enable a data transmission device to allocate a network bandwidth to the target data according to the second priority parameter.Type: ApplicationFiled: August 1, 2024Publication date: June 12, 2025Inventors: Shaoyu ZHANG, Yuhui CHEN, Yihang MIAO, Ruidong CAO, Tao WANG, Liyang ZHAO
-
Publication number: 20250182765Abstract: Methods, systems, and apparatus, including computer programs encoded on computer storage media relate to a method for target speaker extraction. A target speaker extraction system receives an audio frame of an audio signal. A multi-speaker detection model analyzes the audio frame to determine whether the audio frame includes only a single-speaker or multiple speakers. When the audio frame includes only a single-speaker, the system inputs the audio frame to a target speaker VAD model to suppress speech in the audio frame from a non-target speaker based on comparing the audio frame to a voiceprint of a target speaker. When the audio frame includes multiple speakers, the system inputs the audio frame to a speech separation model to separate the voice of the target speaker from a voice mixture in the audio frame.Type: ApplicationFiled: February 3, 2025Publication date: June 5, 2025Applicant: Zoom Communications, Inc.Inventors: Yuhui Chen, Qiyong Liu, Zhengwei Wei, Yangbin Zeng
-
Publication number: 20250086869Abstract: Various embodiments of an apparatus, method(s), system(s) and computer program product(s) described herein are directed to a Viseme Engine. The Viseme Engine receives audio data associated with a user account. The Viseme Engine predicts at least one viseme that corresponds with a portion of phoneme audio data and identifies one or more facial expression parameters associated with the predicted viseme. The facial expression parameters being applicable to a face model. The Viseme Engine renders the predicted viseme according to the one or more facial expression parameters.Type: ApplicationFiled: November 25, 2024Publication date: March 13, 2025Inventors: Yuhui Chen, Cheng Lun Hu, Bo Ling, Gengdai Liu, Huixi Zhao, Yian Zhu
-
Publication number: 20250088800Abstract: Example methods and systems provide automatic audio equalization for online conferences. A target, ideal, audio frequency response for use in online conferencing audio can be designed and preset for reference, and the equalizer can operate in real time to reach the target frequency response for any input. The equalizer can selectively apply per-band equalization to the audio signal as needed to adjust an energy value for each of multiple frequency bands to produce an output audio signal that that can be, as an example, routed through meeting servers or other infrastructure to other conference participants. The equalization can compensate for conditions local to a speaker that would otherwise adversely affect audio quality.Type: ApplicationFiled: November 26, 2024Publication date: March 13, 2025Applicant: Zoom Video Communications, Inc.Inventors: Yuhui Chen, Zhaofeng Jia
-
Publication number: 20250078852Abstract: Audio enhancement of musical content is performed by a device coupled to a network. The device receives an audio signal to be transmitted over the network, and detects when musical content is present in the audio signal based on a content probability threshold. The device disables noise suppression for the audio signal and applies a linear filter to cancel echo for the audio signal. The device disables gain control for the audio signal and encodes the audio signal using a codec designed for music.Type: ApplicationFiled: November 18, 2024Publication date: March 6, 2025Inventors: Qiyong Liu, Jiachuan Deng, Yuhui Chen, Oded Gal
-
Patent number: 12217761Abstract: Methods, systems, and apparatus, including computer programs encoded on computer storage media relate to a method for target speaker extraction. A target speaker extraction system receives an audio frame of an audio signal. A multi-speaker detection model analyzes the audio frame to determine whether the audio frame includes only a single-speaker or multiple speakers. When the audio frame includes only a single-speaker, the system inputs the audio frame to a target speaker VAD model to suppress speech in the audio frame from a non-target speaker based on comparing the audio frame to a voiceprint of a target speaker. When the audio frame includes multiple speakers, the system inputs the audio frame to a speech separation model to separate the voice of the target speaker from a voice mixture in the audio frame.Type: GrantFiled: October 31, 2021Date of Patent: February 4, 2025Assignee: Zoom Video Communications, Inc.Inventors: Yuhui Chen, Qiyong Liu, Zhengwei Wei, Yangbin Zeng
-
Patent number: 12207063Abstract: Example methods and systems provide automatic audio equalization for online conferences. A target, ideal, audio frequency response for use in online conferencing audio can be designed and preset for reference, and the equalizer can operate in real time to reach the target frequency response for any input. The equalizer can selectively apply per-band equalization to the audio signal as needed to adjust an energy value for each of multiple frequency bands to produce an output audio signal that that can be, as an example, routed through meeting servers or other infrastructure to other conference participants. The equalization can compensate for conditions local to a speaker that would otherwise adversely affect audio quality.Type: GrantFiled: September 26, 2022Date of Patent: January 21, 2025Assignee: Zoom Video Communications, Inc.Inventors: Yuhui Chen, Zhaofeng Jia
-
Patent number: 12190424Abstract: Various embodiments of an apparatus, method(s), system(s) and computer program product(s) described herein are directed to a Viseme Engine. The Viseme Engine receives audio data associated with a user account. The Viseme Engine predicts at least one viseme that corresponds with a portion of phoneme audio data and identifies one or more facial expression parameters associated with the predicted viseme. The facial expression parameters being applicable to a face model. The Viseme Engine renders the predicted viseme according to the one or more facial expression parameters.Type: GrantFiled: May 24, 2023Date of Patent: January 7, 2025Assignee: Zoom Video Communications, Inc.Inventors: Yuhui Chen, Cheng Lun Hu, Bo Ling, Gengdai Liu, Huixi Zhao, Yian Zhu
-
Patent number: 12183357Abstract: Dynamic adjustment of audio characteristics for enhancing musical sound during a networked conference is disclosed. In an embodiment, a method is provided for sound enhancement performed by a device coupled to a network. The method includes receiving an audio signal to be transmitted over the network, detecting when musical content is present in the audio signal, processing the audio signal to enhance voice characteristics to generate an enhanced audio signal when the musical content is not detected, processing the audio signal to enhance music characteristic to generate the enhanced audio signal when the musical content is detected, and transmitting the enhanced audio signal over the network.Type: GrantFiled: December 16, 2022Date of Patent: December 31, 2024Assignee: Zoom Video Communications, Inc.Inventors: Qiyong Liu, Jiachuan Deng, Yuhui Chen, Oded Gal
-
Publication number: 20240419391Abstract: Techniques for adaptive audio processing during video conferencing are provided. In an example method, a client device joins a video conference hosted by a video conference provider, the video conference including a plurality of client devices. The client device receives, from an audio input device, an audio stream. The client device then processes, using an audio processing component, the audio stream. The client device determines, using a trained machine learning model, one or more characteristics of the audio stream. The client device then determines, based on the one or more characteristics of the audio stream, an audio configuration operation comprising one or more instructions. The client device executes the audio configuration operation. The client device then outputs the audio stream to the video conference provider.Type: ApplicationFiled: October 3, 2023Publication date: December 19, 2024Inventors: Yuhui CHEN, Qiang GAO, Zhaofeng JIA, Shiwei WANG
-
Patent number: D1113422Type: GrantFiled: April 29, 2024Date of Patent: February 17, 2026Inventor: Yuhui Chen