Patents by Inventor Dmitriy Obukhov

Dmitriy Obukhov has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Publication number: 20260120378
    Abstract: A system trains a voice synthesis model to convert text to speech, wherein the training is based on an audio training dataset comprising audio samples of one or more persons. The system receives at least one audio sample of the target person. The system trains a voice custom synthesis model to identify person-specific speech characteristics, wherein the training is based on the at least one audio sample. The system receives an input text. The system generates, using both the voice synthesis model and the voice custom synthesis model, an audio avatar that recites the input text in a voice of the target person. The system processes the audio avatar to be formatted by phrases and expressions.
    Type: Application
    Filed: December 24, 2025
    Publication date: April 30, 2026
    Inventors: Dmitriy Obukhov, Marcel de Korte, Denis Parkhomenko, Ivan Kirillov, Alexey Rybak, Laurent Dedenis, Serg Bell, Stanislav Protasov
  • Publication number: 20260120379
    Abstract: A system trains a video synthesis model using a video training dataset comprising video samples of one or more persons. The system receives a video sample of the target person. The system trains a video custom synthesis model based on the video sample. The system generates, using both the video synthesis model and the video custom synthesis model, a video avatar that mimics visuals of the target person, wherein generating the video avatar further comprises: generating a preliminary video of a head of the target person with controlled gestures based on recorded gestures from the video sample and a target gesture script; and adding lip synchronization to the preliminary video by matching a voice recording of the target person to a plurality of lip movements based on words spoken by the target person in the preliminary video.
    Type: Application
    Filed: December 24, 2025
    Publication date: April 30, 2026
    Inventors: Dmitriy Obukhov, Marcel de Korte, Denis Parkhomenko, Ivan Kirillov, Alexey Rybak, Laurent Dedenis, Serg Bell, Stanislav Protasov
  • Patent number: 12561876
    Abstract: The present disclosure relates to an avatar generator to generate an audio-visual avatar specific to an application, such as tutoring. The avatar generator includes a general synthesizer to receive a training dataset. The general synthesizer includes a voice synthesis module and a video synthesis module trained by the training dataset. The avatar generator includes a customized synthesizer consisting of a voice custom synthesis module and a video custom synthesis module trained on the audio-video samples of the target person. The avatar generator further includes a video generator to create an audio-visual avatar and is configured to synthesize a voice clone using an input text, process the voice clone, synthesize a video clone based on the video synthesis module and the video custom synthesis module, and apply the voice clone to the video clone.
    Type: Grant
    Filed: November 28, 2022
    Date of Patent: February 24, 2026
    Assignees: Constructor Technology AG, Constructor Education and Research Genossenschaft
    Inventors: Dmitriy Obukhov, Marcel de Korte, Denis Parkhomenko, Ivan Kirillov, Alexey Rybak, Laurent Dedenis, Serg Bell, Stanislav Protasov
  • Publication number: 20250391005
    Abstract: A system obtains, by a video evaluator, a video clip generated by a video generator of the avatar generator. The system obtains, by the video evaluator, video features of a target person that the avatar is representing. The system compares the video clip with the video features of the target person using a set of video metrics. The system generates a video evaluation score for the video clip based on a comparison of the video clip and the video features.
    Type: Application
    Filed: September 3, 2025
    Publication date: December 25, 2025
    Inventors: Ilya Baimetov, Denis Parkhomenko, Marcel de Korte, Ivan Kirillov, Dmitriy Obukhov, Alexey Rybak, Laurent Dedenis, Serg Bell, Stanislav Protasov
  • Publication number: 20250384537
    Abstract: A system obtains, by an audio evaluator, a speech generated by a text-to-speech module of the avatar generator. The system obtains, by the audio evaluator, audio features of a target person that the avatar is representing. The system compares the speech with the audio features of the target person using a set of audio metrics, and generating an audio evaluation score for the speech based on a comparison of the speech and the audio features, wherein generating the audio evaluation score comprises evaluating one or more of: m speech intelligibility using automatic-speech-recognition (ASR) based evaluation metrics, audio noise level using voice-activity-detection (VAD) based evaluation metrics, naturalness of speech intonation using pitch-based metrics, voice similarities using equal-error-rate (EER) and cosine (COS) metrics, and speech pronunciation statistics.
    Type: Application
    Filed: September 3, 2025
    Publication date: December 18, 2025
    Inventors: Ilya Baimetov, Denis Parkhomenko, Marcel de Korte, Ivan Kirillov, Dmitriy Obukhov, Alexey Rybak, Laurent Dedenis, Serg Bell, Stanislav Protasov
  • Patent number: 12456180
    Abstract: The present disclosure relates to a system to evaluate an avatar generated by an avatar generator. The system comprises an evaluation module including an audio evaluation module for evaluating audio features and a video evaluation module for evaluating video features. Evaluation of the avatar includes extracting audio and video features from the avatar and applying a set of evaluation metrics for generating audio and video evaluation scores. The scores are combined to generate a final score. For avatar generator evaluation, audio clip and video clip are provided to the audio evaluation module and video evaluation module, respectively. A set of evaluation metrics is applied for evaluation. Each metric can generate a score. All scores are combined to generate a final evaluation score.
    Type: Grant
    Filed: November 28, 2022
    Date of Patent: October 28, 2025
    Assignee: Constructor Technology AG
    Inventors: Ilya Baimetov, Denis Parkhomenko, Marcel de Korte, Ivan Kirillov, Dmitriy Obukhov, Alexey Rybak, Laurent Dedenis, Serg Bell, Stanislav Protasov
  • Publication number: 20240177283
    Abstract: The present disclosure relates to a system to evaluate an avatar generated by an avatar generator. The system comprises an evaluation module including an audio evaluation module for evaluating audio features and a video evaluation module for evaluating video features. Evaluation of the avatar includes extracting audio and video features from the avatar and applying a set of evaluation metrics for generating audio and video evaluation scores. The scores are combined to generate a final score. For avatar generator evaluation, audio clip and video clip are provided to the audio evaluation module and video evaluation module, respectively. A set of evaluation metrics is applied for evaluation. Each metric can generate a score. All scores are combined to generate a final evaluation score.
    Type: Application
    Filed: November 28, 2022
    Publication date: May 30, 2024
    Inventors: Ilya Baimetov, Denis Parkhomenko, Marcel de Korte, Ivan Kirillov, Dmitriy Obukhov, Alexey Rybak, Laurent Dedenis, Serg Bell, Stanislav Protasov
  • Publication number: 20240177386
    Abstract: The present disclosure relates to an avatar generator to generate an audio-visual avatar specific to an application, such as tutoring. The avatar generator includes a general synthesizer to receive a training dataset. The general synthesizer includes a voice synthesis module and a video synthesis module trained by the training dataset. The avatar generator includes a customized synthesizer consisting of a voice custom synthesis module and a video custom synthesis module trained on the audio-video samples of the target person. The avatar generator further includes a video generator to create an audio-visual avatar and is configured to synthesize a voice clone using an input text, process the voice clone, synthesize a video clone based on the video synthesis module and the video custom synthesis module, and apply the voice clone to the video clone.
    Type: Application
    Filed: November 28, 2022
    Publication date: May 30, 2024
    Inventors: Dmitriy Obukhov, Marcel de Korte, Denis Parkhomenko, Ivan Kirillov, Alexey Rybak, Laurent Dedenis, Serg Bell, Stanislav Protasov