Patents by Inventor Yuhui Chen

Yuhui Chen has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Patent number: 12684087
    Abstract: Systems and methods for active speaker detection for videoconferencing are provided. For example, an audio recording system can record a set of raw audio signals. The set of raw audio signals include audio signals recorded by a set of microphones at different distances from an audio source with different orientations. An audio dataset generation system can access a virtual meeting room setup which specifies microphones used in a virtual meeting room, locations of speakers and the microphones, and orientations of the speakers. The audio dataset generation system generates a synthetic audio signal for each of the speakers specified in the virtual meeting room setup by combining audio signals selected from the set of raw audio signals according to the virtual meeting room setup.
    Type: Grant
    Filed: January 8, 2024
    Date of Patent: July 14, 2026
    Assignee: Zoom Communications, Inc.
    Inventors: Yuhui Chen, Qiang Gao, Zhaofeng Jia, Xian Tong, Ye Wang
  • Patent number: 12671958
    Abstract: Systems and methods for synthetic audio datasets generation for videoconferencing based on realistic room acoustic simulation are provided. For example, a room-independent recording for a source audio signal and a particular type of microphones can be generated and a target room setup for a target room can be obtained. The target room setup specifying one or more of a size of the target room, respective locations of a speaker and a microphone in the target room. Room characteristics of the target room can be generated based on the target room setup via a room acoustic model. A synthetic audio signal for the target room and the particular type of microphones can be generated by applying the room acoustic characteristics onto the room-independent recording.
    Type: Grant
    Filed: January 29, 2024
    Date of Patent: June 30, 2026
    Assignee: Zoom Communications, Inc.
    Inventors: Yuhui Chen, Yifeng Fan, Zhaofeng Jia
  • Patent number: 12645422
    Abstract: Techniques for adaptive audio processing during video conferencing are provided. In an example method, a client device joins a video conference hosted by a video conference provider, the video conference including a plurality of client devices. The client device receives, from an audio input device, an audio stream. The client device then processes, using an audio processing component, the audio stream. The client device determines, using a trained machine learning model, one or more characteristics of the audio stream. The client device then determines, based on the one or more characteristics of the audio stream, an audio configuration operation comprising one or more instructions. The client device executes the audio configuration operation. The client device then outputs the audio stream to the video conference provider.
    Type: Grant
    Filed: October 3, 2023
    Date of Patent: June 2, 2026
    Assignee: Zoom Communications, Inc.
    Inventors: Yuhui Chen, Qiang Gao, Zhaofeng Jia, Shiwei Wang
  • Patent number: 12626709
    Abstract: Methods, systems, and apparatus, including computer programs encoded on computer storage media relate to a method for audio super resolution. The system receives an audio signal. When the sampling rate of the audio signal is below a sampling rate threshold or the frequency range of the audio signal is below a frequency range threshold, the audio signal is input to an audio super resolution model comprising a machine learning model. The audio signal is processed by the audio super resolution model to generate a synthetic audio signal with a wider frequency range than the frequency range of the audio signal.
    Type: Grant
    Filed: October 31, 2021
    Date of Patent: May 12, 2026
    Assignee: Zoom Communications, Inc.
    Inventors: Yuhui Chen, Zhaofeng Jia, Qiyong Liu, Zhengwei Wei
  • Publication number: 20260099974
    Abstract: Systems and methods for personalized realistic video generation. In one example, a client device joins a video conference. The client device accesses a source video clip including a set of source video frames related to a user associated with the client device. The client device receives source audio data related to the user. The client device generates target video data based on the set of source video frames and the source audio data using a trained video generator model. The client device streams the target video data during the video conference.
    Type: Application
    Filed: December 4, 2024
    Publication date: April 9, 2026
    Applicant: Zoom Video Communications, Inc.
    Inventors: Yuanqi Chen, Yuhui Chen, Dewang Hou, Bo Ling, Gengdai Liu, Liubin Liu, Matthieu Tardivel, Rui Zhang, Yian Zhu
  • Publication number: 20260046362
    Abstract: Example methods and systems provide machine-learning assisted acoustic echo cancellation (AEC). The AEC can be used, as an example, to improve the audio quality for online audio and video conferences. A system according to this disclosure includes a pre-trained, machine-learning, AI model designed to detect, in real time, a unitary voice signal, or a signal representing the speech of a single speaker as opposed to that of multiple speakers. A digital signal processing (DSP) algorithm can then detect the echo state, for example, whether distortion results primarily from an echo. Based on these characteristics, the system can, alternatively and automatically apply either a default mode of AEC to the audio signal, or apply a more aggressive mode of AEC.
    Type: Application
    Filed: October 21, 2025
    Publication date: February 12, 2026
    Applicant: Zoom Communications, Inc.
    Inventors: Yuhui Chen, Zhaofeng Jia, Wei Wang
  • Publication number: 20260031097
    Abstract: Techniques for low-quality audio detection are provided. In an example method, a computing system joins a first client device to a first video conference including a number of connected client devices. The computing system receives, from the first client device, a first audio stream. The computing system determines, using a machine learning (“ML”) model, at least one first audio quality measurement based on the first audio stream. The computing system computes a first metric for the first audio stream based on the at least one first audio quality measurement. In response to the first metric satisfying a predetermined threshold, the computing system outputs a message including first information about a first low-quality audio status associated with the first audio stream.
    Type: Application
    Filed: July 25, 2024
    Publication date: January 29, 2026
    Applicant: Zoom Video Communications, Inc.
    Inventors: Yuhui Chen, Zhiye Chen, Zhaofeng Jia, Shiwei Wang
  • Patent number: 12477070
    Abstract: Example methods and systems provide machine-learning assisted acoustic echo cancellation (AEC). The AEC can be used, as an example, to improve the audio quality for online audio and video conferences. A system according to this disclosure includes a pre-trained, machine-learning, AI model designed to detect, in real time, a unitary voice signal, or a signal representing the speech of a single speaker as opposed to that of multiple speakers. A digital signal processing (DSP) algorithm can then detect the echo state, for example, whether distortion results primarily from an echo. Based on these characteristics, the system can, alternatively and automatically apply either a default mode of AEC to the audio signal, or apply a more aggressive mode of AEC.
    Type: Grant
    Filed: November 2, 2023
    Date of Patent: November 18, 2025
    Assignee: Zoom Communications, Inc.
    Inventors: Yuhui Chen, Zhaofeng Jia, Wei Wang
  • Publication number: 20250201231
    Abstract: Systems and methods for generating speaker video and audio in multiple languages for videoconferencing are provided. For example, a computing device can access a speaker speech audio signal that includes a speaker speech in a first language, a video of the speaker and a translated speech audio signal of the speaker speech in a second language. The computing device generates, based on the translated speech audio signal, a converted translated speech audio signal that includes a speech in the second language having voice characteristics in the speaker speech. The computing device further generates a lip-synched speaker video based on the video of the speaker and the converted translated speech audio signal. Lip movements in the lip-synched speaker video correspond to the converted translated speech audio signal. The converted translated speech audio signal and the lip-synched speaker video are transmitted to a video conference provider configured to host the video conference.
    Type: Application
    Filed: December 18, 2023
    Publication date: June 19, 2025
    Inventors: Yuhui Chen, Qiang Gao, Zhaofeng Jia
  • Publication number: 20250193130
    Abstract: Embodiments of the present disclosure provide a network traffic flow control method and apparatus based on a distributed storage system, an electrical device, and a storage medium. The method includes: acquiring a transmission link for transmitting target data in response to a first data server in the distributed storage system receiving a transmission request for the target data; configuring a first priority parameter for the transmission link by the first data server according to the priority information of the business type to which the target data belongs; and in response to transmitting the target data through the transmission link, configuring a second priority parameter for the target data by the first data server according to the first priority parameter of the transmission link, so as to enable a data transmission device to allocate a network bandwidth to the target data according to the second priority parameter.
    Type: Application
    Filed: August 1, 2024
    Publication date: June 12, 2025
    Inventors: Shaoyu ZHANG, Yuhui CHEN, Yihang MIAO, Ruidong CAO, Tao WANG, Liyang ZHAO
  • Publication number: 20250182765
    Abstract: Methods, systems, and apparatus, including computer programs encoded on computer storage media relate to a method for target speaker extraction. A target speaker extraction system receives an audio frame of an audio signal. A multi-speaker detection model analyzes the audio frame to determine whether the audio frame includes only a single-speaker or multiple speakers. When the audio frame includes only a single-speaker, the system inputs the audio frame to a target speaker VAD model to suppress speech in the audio frame from a non-target speaker based on comparing the audio frame to a voiceprint of a target speaker. When the audio frame includes multiple speakers, the system inputs the audio frame to a speech separation model to separate the voice of the target speaker from a voice mixture in the audio frame.
    Type: Application
    Filed: February 3, 2025
    Publication date: June 5, 2025
    Applicant: Zoom Communications, Inc.
    Inventors: Yuhui Chen, Qiyong Liu, Zhengwei Wei, Yangbin Zeng
  • Publication number: 20250086869
    Abstract: Various embodiments of an apparatus, method(s), system(s) and computer program product(s) described herein are directed to a Viseme Engine. The Viseme Engine receives audio data associated with a user account. The Viseme Engine predicts at least one viseme that corresponds with a portion of phoneme audio data and identifies one or more facial expression parameters associated with the predicted viseme. The facial expression parameters being applicable to a face model. The Viseme Engine renders the predicted viseme according to the one or more facial expression parameters.
    Type: Application
    Filed: November 25, 2024
    Publication date: March 13, 2025
    Inventors: Yuhui Chen, Cheng Lun Hu, Bo Ling, Gengdai Liu, Huixi Zhao, Yian Zhu
  • Publication number: 20250088800
    Abstract: Example methods and systems provide automatic audio equalization for online conferences. A target, ideal, audio frequency response for use in online conferencing audio can be designed and preset for reference, and the equalizer can operate in real time to reach the target frequency response for any input. The equalizer can selectively apply per-band equalization to the audio signal as needed to adjust an energy value for each of multiple frequency bands to produce an output audio signal that that can be, as an example, routed through meeting servers or other infrastructure to other conference participants. The equalization can compensate for conditions local to a speaker that would otherwise adversely affect audio quality.
    Type: Application
    Filed: November 26, 2024
    Publication date: March 13, 2025
    Applicant: Zoom Video Communications, Inc.
    Inventors: Yuhui Chen, Zhaofeng Jia
  • Publication number: 20250078852
    Abstract: Audio enhancement of musical content is performed by a device coupled to a network. The device receives an audio signal to be transmitted over the network, and detects when musical content is present in the audio signal based on a content probability threshold. The device disables noise suppression for the audio signal and applies a linear filter to cancel echo for the audio signal. The device disables gain control for the audio signal and encodes the audio signal using a codec designed for music.
    Type: Application
    Filed: November 18, 2024
    Publication date: March 6, 2025
    Inventors: Qiyong Liu, Jiachuan Deng, Yuhui Chen, Oded Gal
  • Patent number: 12217761
    Abstract: Methods, systems, and apparatus, including computer programs encoded on computer storage media relate to a method for target speaker extraction. A target speaker extraction system receives an audio frame of an audio signal. A multi-speaker detection model analyzes the audio frame to determine whether the audio frame includes only a single-speaker or multiple speakers. When the audio frame includes only a single-speaker, the system inputs the audio frame to a target speaker VAD model to suppress speech in the audio frame from a non-target speaker based on comparing the audio frame to a voiceprint of a target speaker. When the audio frame includes multiple speakers, the system inputs the audio frame to a speech separation model to separate the voice of the target speaker from a voice mixture in the audio frame.
    Type: Grant
    Filed: October 31, 2021
    Date of Patent: February 4, 2025
    Assignee: Zoom Video Communications, Inc.
    Inventors: Yuhui Chen, Qiyong Liu, Zhengwei Wei, Yangbin Zeng
  • Patent number: 12207063
    Abstract: Example methods and systems provide automatic audio equalization for online conferences. A target, ideal, audio frequency response for use in online conferencing audio can be designed and preset for reference, and the equalizer can operate in real time to reach the target frequency response for any input. The equalizer can selectively apply per-band equalization to the audio signal as needed to adjust an energy value for each of multiple frequency bands to produce an output audio signal that that can be, as an example, routed through meeting servers or other infrastructure to other conference participants. The equalization can compensate for conditions local to a speaker that would otherwise adversely affect audio quality.
    Type: Grant
    Filed: September 26, 2022
    Date of Patent: January 21, 2025
    Assignee: Zoom Video Communications, Inc.
    Inventors: Yuhui Chen, Zhaofeng Jia
  • Patent number: 12190424
    Abstract: Various embodiments of an apparatus, method(s), system(s) and computer program product(s) described herein are directed to a Viseme Engine. The Viseme Engine receives audio data associated with a user account. The Viseme Engine predicts at least one viseme that corresponds with a portion of phoneme audio data and identifies one or more facial expression parameters associated with the predicted viseme. The facial expression parameters being applicable to a face model. The Viseme Engine renders the predicted viseme according to the one or more facial expression parameters.
    Type: Grant
    Filed: May 24, 2023
    Date of Patent: January 7, 2025
    Assignee: Zoom Video Communications, Inc.
    Inventors: Yuhui Chen, Cheng Lun Hu, Bo Ling, Gengdai Liu, Huixi Zhao, Yian Zhu
  • Patent number: 12183357
    Abstract: Dynamic adjustment of audio characteristics for enhancing musical sound during a networked conference is disclosed. In an embodiment, a method is provided for sound enhancement performed by a device coupled to a network. The method includes receiving an audio signal to be transmitted over the network, detecting when musical content is present in the audio signal, processing the audio signal to enhance voice characteristics to generate an enhanced audio signal when the musical content is not detected, processing the audio signal to enhance music characteristic to generate the enhanced audio signal when the musical content is detected, and transmitting the enhanced audio signal over the network.
    Type: Grant
    Filed: December 16, 2022
    Date of Patent: December 31, 2024
    Assignee: Zoom Video Communications, Inc.
    Inventors: Qiyong Liu, Jiachuan Deng, Yuhui Chen, Oded Gal
  • Publication number: 20240419391
    Abstract: Techniques for adaptive audio processing during video conferencing are provided. In an example method, a client device joins a video conference hosted by a video conference provider, the video conference including a plurality of client devices. The client device receives, from an audio input device, an audio stream. The client device then processes, using an audio processing component, the audio stream. The client device determines, using a trained machine learning model, one or more characteristics of the audio stream. The client device then determines, based on the one or more characteristics of the audio stream, an audio configuration operation comprising one or more instructions. The client device executes the audio configuration operation. The client device then outputs the audio stream to the video conference provider.
    Type: Application
    Filed: October 3, 2023
    Publication date: December 19, 2024
    Inventors: Yuhui CHEN, Qiang GAO, Zhaofeng JIA, Shiwei WANG
  • Patent number: D1113422
    Type: Grant
    Filed: April 29, 2024
    Date of Patent: February 17, 2026
    Inventor: Yuhui Chen