Patents by Inventor Wan Ding
Wan Ding has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Patent number: 12694595Abstract: A method for synthesizing a talking head video includes: obtaining speech data to be synthesized and observation data, wherein the observation data is data obtained through observation other than the speech data; performing feature extraction on the speech data to obtain speech features corresponding to the speech data, and performing feature extraction on the observation data to obtain non-speech features corresponding to the observation data; performing temporal modeling on the speech features and first non-speech features to obtain low-dimensional representations, wherein the first non-speech features are non-speech features that are sensitive to temporal changes; and performing video synthesis based on the low-dimensional representations and second non-speech features, wherein the second non-speech features are non-speech features insensitive to temporal changes.Type: GrantFiled: June 7, 2024Date of Patent: July 28, 2026Assignee: UBTECH ROBOTICS CORP LTDInventors: Wan Ding, Dongyan Huang, Xianjie Yang, Zehong Zheng, Penghui Li
-
Publication number: 20260204272Abstract: An echo cancellation method includes: acquiring a reference-aligned signal and a microphone-aligned signal; filtering the reference-aligned signal and the microphone-aligned signal by using a preset adaptive filter to obtain a filter output signal; performing speech compensation processing on the filter output signal according to the reference-aligned signal and the microphone-aligned signal to obtain a compensated speech signal; and performing nonlinear processing on the compensated speech signal to obtain an echo-cancelled speech signal.Type: ApplicationFiled: December 23, 2025Publication date: July 16, 2026Inventors: ZEHONG ZHENG, Dongyan Huang, Xianjie Yang, Wan Ding
-
Patent number: 12400635Abstract: A text-to-speech synthesis method, an electronic device, and a computer-readable storage medium are provided. The method includes: obtaining prosodic pause features of an input text by performing a prosodic pause prediction processing on the input text, and dividing the input text into a plurality of prosodic phrases according to the prosodic pause features; synthesizing short sentence audios according to the prosodic phrases by performing a streamed speech synthesis processing on each of the prosodic phrases in the input text in a manner of asynchronous processing of a thread pool; and performing an audio playback operation of the input text according to the short sentence audios corresponding to the first prosodic phrase of the input text, in response to synthesizing the short sentence audio corresponding to the first prosodic phrase of the input text.Type: GrantFiled: June 20, 2023Date of Patent: August 26, 2025Assignee: UBTECH ROBOTICS CORP LTDInventors: Wan Ding, Dongyan Huang, Zehong Zheng, Linhuang Yan, Zhiyong Yang
-
Patent number: 12374319Abstract: A speech synthesis method includes: obtaining an acoustic feature sequence of a text to be processed; processing the acoustic feature sequence by using a non-autoregressive computing model in parallel to obtain first audio information of the text to be processed, wherein the first audio information comprises audio corresponding to each segment; processing the acoustic feature sequence and the first audio information by using an autoregressive computing model to obtain a residual value corresponding to each segment; and obtaining second audio information corresponding to an i-th segment based on the first audio information corresponding to the i-th segment and the residual values corresponding to a first to an (i?1)-th segment, wherein a synthesized audio of the text to be processed comprises each of the second audio information, i=1, 2 . . . n, n is a total number of the segments.Type: GrantFiled: December 28, 2022Date of Patent: July 29, 2025Assignee: UBTECH ROBOTICS CORP LTDInventors: Wan Ding, Dongyan Huang, Zhiyuan Zhao, Zhiyong Yang
-
Patent number: 12315059Abstract: A method for generating a talking head video includes: obtaining a text and an image containing a face of a user; determining a phoneme sequence that corresponds to the text and includes one or more phonemes; determining acoustic features corresponding to the text according to the phoneme sequence, and obtaining synthesized speech corresponding to the text according to the acoustic features; determining a first mouth movement sequence corresponding to the text according to the phoneme sequence, and determining a second mouth movement sequence corresponding to the text according to the acoustic features; creating a facial action video corresponding to the user according to the first mouth movement sequence, the second mouth movement sequence and the image; and processing the synthesized speech and the facial action video synchronously to obtain a talking head video corresponding to the user.Type: GrantFiled: May 26, 2023Date of Patent: May 27, 2025Assignee: UBTECH ROBOTICS CORP LTDInventors: Wan Ding, Dongyan Huang, Linhuang Yan, Zhiyong Yang
-
Publication number: 20250133337Abstract: A sound source localization method includes: obtaining a first audio frame and at least two second audio frames, wherein the first audio frame and the at least two second audio frames are synchronously sampled, the first audio frame is obtained by processing sound signals collected by the first microphone, the at least two second audio frames are obtained by processing sound signals collected by the second microphones; calculating a time delay estimation between the first audio frame and each of the at least two second audio frames; and determining a sound source orientation corresponding to the first audio frame and the at least two second audio frames through a preset time delay-orientation lookup table according to the time delay estimation between the first audio frame and each of the at least two second audio frames.Type: ApplicationFiled: October 9, 2024Publication date: April 24, 2025Inventors: ZEHONG ZHENG, Dongyan Huang, Xianjie Yang, Wan Ding
-
Patent number: 12263600Abstract: A calibration method for an industrial robot includes receiving a first model of the industrial robot, the first model is synchronized with a pose state of the industrial robot located at a specific position in an actual environment; receiving an environment model around the industrial robot, the environment model including a second model of the industrial robot; obtaining registration information of the second model, at least by selecting at least three corresponding non-collinear point pairs in the first model and the second model to perform registration; and based on the registration information, calibrating a coordinate system of the environment model to a base coordinate system of the industrial robot.Type: GrantFiled: September 6, 2019Date of Patent: April 1, 2025Assignee: Robert Bosch GmbHInventors: Wan Ding, Liupeng Yan, William Wang
-
Publication number: 20240428493Abstract: A method for synthesizing a talking head video includes: obtaining speech data to be synthesized and observation data, wherein the observation data is data obtained through observation other than the speech data; performing feature extraction on the speech data to obtain speech features corresponding to the speech data, and performing feature extraction on the observation data to obtain non-speech features corresponding to the observation data; performing temporal modeling on the speech features and first non-speech features to obtain low-dimensional representations, wherein the first non-speech features are non-speech features that are sensitive to temporal changes; and performing video synthesis based on the low-dimensional representations and second non-speech features, wherein the second non-speech features are non-speech features insensitive to temporal changes.Type: ApplicationFiled: June 7, 2024Publication date: December 26, 2024Inventors: WAN DING, Dongyan Huang, Xianjie Yang, Zehong Zheng, Penghul Li
-
Patent number: 11941366Abstract: The present disclosure discloses a context-based multi-turn dialogue method.Type: GrantFiled: November 23, 2020Date of Patent: March 26, 2024Assignee: UBTECH ROBOTICS CORP LTDInventors: Chi Shao, Dongyan Huang, Wan Ding, Youjun Xiong
-
Publication number: 20230410791Abstract: A text-to-speech synthesis method, an electronic device, and a computer-readable storage medium are provided. The method includes: obtaining prosodic pause features of an input text by performing a prosodic pause prediction processing on the input text, and dividing the input text into a plurality of prosodic phrases according to the prosodic pause features; synthesizing short sentence audios according to the prosodic phrases by performing a streamed speech synthesis processing on each of the prosodic phrases in the input text in a manner of asynchronous processing of a thread pool; and performing an audio playback operation of the input text according to the short sentence audios corresponding to the first prosodic phrase of the input text, in response to synthesizing the short sentence audio corresponding to the first prosodic phrase of the input text.Type: ApplicationFiled: June 20, 2023Publication date: December 21, 2023Inventors: Wan Ding, Dongyuan Huang, Zehong Zheng, Linhuang Yan, Zhiyong Yang
-
Publication number: 20230386116Abstract: A method for generating a talking head video includes: obtaining a text and an image containing a face of a user; determining a phoneme sequence that corresponds to the text and includes one or more phonemes; determining acoustic features corresponding to the text according to the phoneme sequence, and obtaining synthesized speech corresponding to the text according to the acoustic features; determining a first mouth movement sequence corresponding to the text according to the phoneme sequence, and determining a second mouth movement sequence corresponding to the text according to the acoustic features; creating a facial action video corresponding to the user according to the first mouth movement sequence, the second mouth movement sequence and the image; and processing the synthesized speech and the facial action video synchronously to obtain a talking head video corresponding to the user.Type: ApplicationFiled: May 26, 2023Publication date: November 30, 2023Inventors: WAN DING, Dongyan Huang, Linhuang Yan, Zhiyong Yang
-
Patent number: 11813754Abstract: A grabbing method for an industrial robot is disclosed. The method includes obtaining an object information file. The object information file includes numbers and/or positions of detected objects. The method further includes determining collision boundary lines and collision representative objects according to the object information file. In addition, the method includes determining a collision-free grabbing path of a gripper of the industrial robot based on the determined collision boundary lines and the collision representative objects. The collision-free grabbing path is a linear path that satisfies joint limits of the industrial robot. The disclosure further relates to a grabbing device for an industrial robot, a computer storage medium, and an industrial robot.Type: GrantFiled: June 24, 2021Date of Patent: November 14, 2023Assignee: Robert Bosch GmbHInventor: Wan Ding
-
Publication number: 20230206895Abstract: A speech synthesis method includes: obtaining an acoustic feature sequence of a text to be processed; processing the acoustic feature sequence by using a non-autoregressive computing model in parallel to obtain first audio information of the text, to be processed, wherein the first audio information comprises audio corresponding to each segment; processing the acoustic feature sequence and the first audio information by using an autoregressive computing model to obtain a residual value corresponding to each segment; and obtaining second audio information corresponding to an i-th segment based on the first audio information corresponding to the i-th segment and the residual values corresponding to a first to an (i-1)-th segment, wherein a synthesized audio of the text to be processed comprises each of the second audio information, i=1, 2 . . . n, n is a total number of the segments.Type: ApplicationFiled: December 28, 2022Publication date: June 29, 2023Inventors: Wan Ding, Dongyan Huang, Zhiyuan Zhao, Zhiyong Yang
-
Publication number: 20220331969Abstract: A calibration method for an industrial robot includes receiving a first model of the industrial robot, the first model is synchronized with an attitude state of the industrial robot located at a specific position in an actual environment; receiving an environment model around the industrial robot, the environment model including a second model of the industrial robot; obtaining registration information of the second model, at least by selecting at least three corresponding non-collinear point pairs in the first model and the second model to perform registration; and based on the registration information, calibrating a coordinate system of the environment model to a base coordinate system of the industrial robot.Type: ApplicationFiled: September 6, 2019Publication date: October 20, 2022Inventors: Wan Ding, Liupeng Yan, William Wang
-
Patent number: 11282503Abstract: The present disclosure discloses a voice conversion training method. The method includes: forming a first training data set including a plurality of training voice data groups; selecting two of the training voice data groups from the first training data set to input into a voice conversion neural network for training; forming a second training data set including the first training data set and a first source speaker voice data group; inputting one of the training voice data groups selected from the first training data set and the first source speaker voice data group into the network for training; forming the third training data set including the second source speaker voice data group and the personalized voice data group that are parallel corpus with respect to each other; and inputting the second source speaker voice data group and the personalized voice data group into the network for training.Type: GrantFiled: November 12, 2020Date of Patent: March 22, 2022Assignee: UBTECH ROBOTICS CORP LTDInventors: Ruotong Wang, Dongyan Huang, Xian Li, Jiebin Xie, Zhichao Tang, Wan Ding, Yang Liu, Bai Li, Youjun Xiong
-
Publication number: 20210402604Abstract: A grabbing method for an industrial robot is disclosed. The method includes obtaining an object information file. The object information file includes numbers and/or positions of detected objects. The method further includes determining collision boundary lines and collision representative objects according to the object information file. In addition, the method includes determining a collision-free grabbing path of a gripper of the industrial robot based on the determined collision boundary lines and the collision representative objects. The collision-free grabbing path is a linear path that satisfies joint limits of the industrial robot. The disclosure further relates to a grabbing device for an industrial robot, a computer storage medium, and an industrial robot.Type: ApplicationFiled: June 24, 2021Publication date: December 30, 2021Inventor: Wan Ding
-
Publication number: 20210200961Abstract: The present disclosure discloses a context-based multi-turn dialogue method.Type: ApplicationFiled: November 23, 2020Publication date: July 1, 2021Inventors: Chi Shao, Dongyan Huang, Wan Ding, Youjun Xiong
-
Publication number: 20210201890Abstract: The present disclosure discloses a voice conversion training method. The method includes: forming a first training data set including a plurality of training voice data groups; selecting two of the training voice data groups from the first training data set to input into a voice conversion neural network for training; forming a second training data set including the first training data set and a first source speaker voice data group; inputting one of the training voice data groups selected from the first training data set and the first source speaker voice data group into the network for training; forming the third training data set including the second source speaker voice data group and the personalized voice data group that are parallel corpus with respect to each other; and inputting the second source speaker voice data group and the personalized voice data group into the network for training.Type: ApplicationFiled: November 12, 2020Publication date: July 1, 2021Inventors: Ruotong Wang, Dongyan Huang, Xian Li, Jiebin Xie, Zhichao Tang, Wan Ding, Yang Liu, Bai Li, Youjun Xiong