Patents by Inventor Xianjun XIA

Xianjun XIA has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Publication number: 20260171101
    Abstract: Embodiments of the present disclosure provide an audio processing method and device, the method including: acquiring audio data of a video file, converting each frame data in the audio data into frequency-domain data through a clipping detection model, and detecting whether clipping data exists in each frame data based on the frequency-domain data; performing, if it is detected that clipping data exists in the frame data, clipping restoration processing on the frame data through a clipping restoration model to obtain restored frame data, the clipping restoration model being configured to filter out clipping data and spectral mirroring from the frame data; and performing normalization processing on audio loudness of the restored frame data based on a maximum amplitude of the restored frame data in a time domain, to obtain restored audio data of the video file.
    Type: Application
    Filed: December 8, 2025
    Publication date: June 18, 2026
    Inventors: Zhuangqi CHEN, Xianjun XIA, Chuanzeng HUANG
  • Publication number: 20260171106
    Abstract: Provided in the present disclosure are a method for processing speech signal, an electronic device and a non-transitory computer-readable medium. The method includes: inputting a first superimposition result into an encoding module of a first convolutional recurrent network (CRN) for convolution processing to obtain a local spectral feature; inputting the local spectral feature into a feature processing module of the first CRN to obtain a global spectral feature; inputting the local spectral feature and the global spectral feature into a decoding module of the first CRN for decoding; inputting a decoded result into an activation module of the first CRN to obtain a first mask; performing, based on the first mask, first masking on a complex spectrum obtained through transformation of the speech signal acquired by the microphone of the first terminal device; and obtaining a target speech signal based on the complex spectrum obtained after the first masking.
    Type: Application
    Filed: February 1, 2024
    Publication date: June 18, 2026
    Inventors: Xianjun XIA, Zhuangqi CHEN, Cheng CHEN, Yijian XIAO
  • Publication number: 20260141909
    Abstract: Embodiments of this application provide an audio restoration method and apparatus, and relates to the technical field of audio restoration. The method includes: performing pop detection on audio to be restored to obtain a pop proportion of the audio to be restored; performing pop restoration on the audio to be restored to obtain first audio in a case where the pop proportion is greater than a first threshold; performing speech detection on the first audio to obtain a speech proportion of the first audio; converting, in a case where the speech proportion is greater than a second threshold, the first audio into a first time-frequency domain signal, segmenting the first time-frequency domain signal into a first number of sub-band signals with non-overlapping frequency bands according to a resolution of the first audio, respectively obtaining spectrum features of the first number of sub-band signals.
    Type: Application
    Filed: November 7, 2025
    Publication date: May 21, 2026
    Inventors: Xiaohuai LE, Zhuangqi Chen, Siyu Sun, Xianjun Xia, Chuanzeng Huang
  • Publication number: 20260134878
    Abstract: Embodiments of this application provide a sound source separation method and apparatus, and relates to the technical field of data processing. The method includes: transforming an audio signal to be separated from a time-domain signal into a time-frequency domain signal; performing frequency band segmentation on the time-frequency domain signal to segment the time-frequency domain signal into a plurality of sub band signals, where frequency bands of the plurality of sub band signals do not overlap; acquiring spectrum features of the plurality of sub band signals respectively; acquiring a spectral mask of at least one sound source of the audio signal to be separated according to the spectrum features of the plurality of sub band signals; and acquiring an audio signal of the at least one sound source according to the spectral mask of the at least one sound source and the time-frequency domain signal.
    Type: Application
    Filed: November 6, 2025
    Publication date: May 14, 2026
    Inventors: Xianjun XIA, Zihan ZHANG, Chuanzeng HUANG
  • Patent number: 12567429
    Abstract: Embodiments of this application provide a real-time voice call control method performed by an electronic device. The method includes: obtaining a mixed call voice in real time during a cloud conference call, where the mixed call voice includes at least one branch voice; determining energy information corresponding to each frequency point of the call voice in a frequency domain; determining an energy proportion of each branch voice at each frequency point in total energy of the frequency point based on the energy information at the frequency point; determining a quantity of branch voices comprised in the call voice based on the energy proportion of each branch voice at each frequency point; and controlling the voice call by setting a call voice control manner based on the quantity of branch voices.
    Type: Grant
    Filed: October 26, 2022
    Date of Patent: March 3, 2026
    Assignee: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
    Inventors: Juanjuan Li, Xianjun Xia
  • Publication number: 20250210057
    Abstract: The present disclosure relates to a field of computer technology and discloses a method, a system, a device, and storage medium for speech enhancement. The method for speech enhancement comprises acquiring audio data and, when speech data is detected in the audio data, extracting an embedding vector of the speech data; searching the embedding vector for a target embedding vector extracted from target speech data, and generating a registration embedding vector based on the target embedding vector; performing correlation calculation between the registration embedding vector and an audio feature vector of the audio data, to determine a masking value required for enhancing the target speech data; and enhancing, according to the masking value, the target speech data in the audio data.
    Type: Application
    Filed: December 16, 2024
    Publication date: June 26, 2025
    Inventors: Xiaohuai LE, Xianjun XIA, Yijian XIAO
  • Publication number: 20250184663
    Abstract: Embodiments of the present disclosure provide an audio processing method and apparatus, a storage medium, and an electronic device. The method includes: acquiring audio to be processed, and obtaining first restored audio by restoring, based on a first processing model, a first type of distortion in the audio to be processed; and obtain second restored audio by restoring, based on a second processing model, a second type of distortion in the first restored audio.
    Type: Application
    Filed: November 27, 2024
    Publication date: June 5, 2025
    Inventors: Xianjun Xia, Mingshuai Liu, Yijian Xiao
  • Publication number: 20230051413
    Abstract: Embodiments of this application provide a real-time voice call control method performed by an electronic device. The method includes: obtaining a mixed call voice in real time during a cloud conference call, where the mixed call voice includes at least one branch voice; determining energy information corresponding to each frequency point of the call voice in a frequency domain; determining an energy proportion of each branch voice at each frequency point in total energy of the frequency point based on the energy information at the frequency point; determining a quantity of branch voices comprised in the call voice based on the energy proportion of each branch voice at each frequency point; and controlling the voice call by setting a call voice control manner based on the quantity of branch voices.
    Type: Application
    Filed: October 26, 2022
    Publication date: February 16, 2023
    Inventors: Juanjuan LI, Xianjun XIA
  • Publication number: 20230041256
    Abstract: An artificial intelligence-based audio processing method includes: obtaining an audio clip of an audio scene, the audio clip including noise; performing audio scene classification processing based on the audio clip to obtain an audio scene type corresponding to the noise in the audio clip; and determining a target audio processing mode corresponding to the audio scene type, and applying the target audio processing mode to the audio clip of the audio scene according to a degree of interference caused by the noise in the audio clip.
    Type: Application
    Filed: October 20, 2022
    Publication date: February 9, 2023
    Inventors: Wen WU, Xianjun XIA