Patents by Inventor Shaofan YANG
Shaofan YANG has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Patent number: 12688000Abstract: A system for managing user-generated content (UGC) and professionally generated content (PGC) is disclosed. The system is programmed to receive digital audio data having two channels from a social media platform. The system is programmed to extract spatial features that capture differences in the two channels from the digital audio data. The system is programmed to also extract temporal features, spectral features, and background features from the digital audio data. The system is programmed to then use the extracted features to determine whether to process the digital audio data as UGC or PGC before playback.Type: GrantFiled: August 11, 2022Date of Patent: July 21, 2026Assignee: Dolby Laboratories Licensing CorporationInventors: Shaofan Yang, Kai Li
-
Publication number: 20250131941Abstract: A system for detecting speech from reverberant signals is disclosed. The system is programmed to receive spectral temporal amplitude data in the modulation frequency domain. The system is programmed to then enhance the spectral temporal amplitude data by reducing reverberation and other noise as well as smoothing based on certain properties of the spectral temporal spectrogram associated with the spectral temporal amplitude data. Next, the system is programmed to compute various features related to the presence of speech based on the enhanced spectral temporal amplitude data and other data in the modulation frequency domain or in the (acoustic) frequency domain. The system is programmed to then determine an extent of speech present in the audio data corresponding to the received spectral temporal amplitude data based on the various features. The system can be programmed to transmit the extent of speech present to an output device.Type: ApplicationFiled: August 11, 2022Publication date: April 24, 2025Applicant: Dolby Laboratories Licensing CorporationInventors: Shaofan YANG, Kai LI
-
Publication number: 20250130756Abstract: A system for managing user-generated content (UGC) and professionally generated content (PGC) is disclosed. The system is programmed to receive digital audio data having two channels from a social media platform. The system is programmed to extract spatial features that capture differences in the two channels from the digital audio data. The system is programmed to also extract temporal features, spectral features, and background features from the digital audio data. The system is programmed to then use the extracted features to determine whether to process the digital audio data as UGC or PGC before playback.Type: ApplicationFiled: August 11, 2022Publication date: April 24, 2025Applicant: Dolby Laboratories Licensing CorporationInventors: Shaofan YANG, Kai LI
-
Publication number: 20250038726Abstract: Described herein is a method of performing content-aware audio processing for an audio signal comprising a plurality of audio components of different types. The method includes source separating the audio signal into at least a voice-related audio component and a residual audio component. The method further includes determining a dynamic audio gain based on the voice-related audio component and the residual audio component. The method also includes performing audio level adjustment for the audio signal based on the determined audio gain. Further described are corresponding apparatus, programs, and computer-readable storage media.Type: ApplicationFiled: November 3, 2022Publication date: January 30, 2025Applicant: Dolby Laboratories Licensing CorporationInventors: Shaofan YANG, Kai LI, Qianqian FANG
-
Publication number: 20240363131Abstract: A method for dereverberating audio signals is provided. In some implementations, the method involves obtaining a real acoustic impulse response (AIR); identifying a first portion of the real AIR corresponding to early reflections of a direct sound and a second portion of the real AIR that corresponding to late reflections of the direct sound; generating one or more synthesized AIRs by modifying the first portion of the real AIR and/or the second portion of the real AIR; and using the real AIR and the one or more synthesized AIRs to generate a plurality of training samples, each training sample comprising an input audio signal and a reverberated audio signal, wherein the reverberated audio signal is generated based on the input audio signal and one of the real AIR or one of the one or more synthesized AIRs, which plurality of training samples are used to train a machine learning model.Type: ApplicationFiled: July 12, 2022Publication date: October 31, 2024Applicant: Dolby Laboratories Licensing CorporationInventors: Jia DAI, Kai LI, Xiaoyu LIU, Richard J. CARTWRIGHT, Shaofan YANG
-
Patent number: 12073828Abstract: Described herein is a method for Convolutional Neural Network (CNN) based speech source separation, wherein the method includes the steps of: (a) providing multiple frames of a time-frequency transform of an original noisy speech signal; (b) inputting the time-frequency transform of said multiple frames into an aggregated multi-scale CNN having a plurality of parallel convolution paths; (c) extracting and outputting, by each parallel convolution path, features from the input time-frequency transform of said multiple frames; (d) obtaining an aggregated output of the outputs of the parallel convolution paths; and (e) generating an output mask for extracting speech from the original noisy speech signal based on the aggregated output. Described herein are further an apparatus for CNN based speech source separation as well as a respective computer program product comprising a computer-readable storage medium with instructions adapted to carry out said method when executed by a device having processing capability.Type: GrantFiled: May 13, 2020Date of Patent: August 27, 2024Assignee: Dolby Laboratories Licensing CorporationInventors: Jundai Sun, Zhiwei Shuang, Lie Lu, Shaofan Yang, Jia Dai
-
Publication number: 20240170002Abstract: A method for reverberation suppression may involve receiving an input audio signal. The method may involve classifying a media type of the input audio signal as one of a group comprising at least: 1) speech; 2) music; or 3) speech over music. The method may involve determining whether to perform dereverberation on the input audio signal based at least on a determination that the media type of the input audio signal has been classified as speech. The method may involve generating an output audio signal by performing dereverberation on the input audio signal in response to determining that dereverberation is to be performed on the input audio signal.Type: ApplicationFiled: March 10, 2022Publication date: May 23, 2024Applicant: DOLBY LABORATORIES LICENSING CORPORATIONInventors: Kai LI, Shaofan YANG, Yuanxing MA
-
Publication number: 20240071411Abstract: Disclosed is a method for determining one or more dialog quality metrics of a mixed audio signal comprising a dialog component and a noise component, the method comprising separating an estimated dialog component from the mixed audio signal by means of a dialog separator using a dialog separating model determined by training the dialog separator based on the one or more quality metrics; providing the estimated dialog component from the dialog separator to a quality metrics estimator; and determining the one or more quality metrics by means of the quality metrics estimator based on the mixed signal and the estimated dialog component. Further disclosed is a method for training a dialog separator, a system comprising circuitry configured to perform the method, and a non-transitory computer-readable storage medium.Type: ApplicationFiled: January 4, 2022Publication date: February 29, 2024Applicant: Dolby Laboratories Licensing CorporationInventors: Jundai SUN, Lie LU, Shaofan YANG, Rhonda J. WILSON, Dirk Jeroen BREEBAART
-
Publication number: 20220223144Abstract: Described herein is a method for Convolutional Neural Network (CNN) based speech source separation, wherein the method includes the steps of: (a) providing multiple frames of a time-frequency transform of an original noisy speech signal; (b) inputting the time-frequency transform of said multiple frames into an aggregated multi-scale CNN having a plurality of parallel convolution paths; (c) extracting and outputting, by each parallel convolution path, features from the input time-frequency transform of said multiple frames; (d) obtaining an aggregated output of the outputs of the parallel convolution paths; and (e) generating an output mask for extracting speech from the original noisy speech signal based on the aggregated output. Described herein are further an apparatus for CNN based speech source separation as well as a respective computer program product comprising a computer-readable storage medium with instructions adapted to carry out said method when executed by a device having processing capability.Type: ApplicationFiled: May 13, 2020Publication date: July 14, 2022Applicant: Dolby Laboratories Licensing CorporationInventors: Jundai SUN, Zhiwei SHUANG, Lie LU, Shaofan YANG, Jia DAI