Patents by Inventor Dmitry Korobchenko
Dmitry Korobchenko has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Publication number: 20260212622Abstract: Various examples, systems, and methods are disclosed relating to a body generation pipeline. A first computing system can generate, according to a plurality of characteristics of a body of a subject, an initial model of the body. The first computing system further can determine a plurality of measurements of a plurality of structures of the initial model. The first computing system further can determine, by at least one neural network, based at least on the plurality of measurements, a plurality of modifications to the initial model, the at least one neural network updated according to training data including a featurized representation of example body shapes and measurements of samples of the featurized representation. The first computing system further can update the initial model according to the plurality of modifications.Type: ApplicationFiled: January 21, 2025Publication date: July 23, 2026Applicant: NVIDIA CorporationInventors: Aleksey KARMANOV, Dmitry KOROBCHENKO, Adeline AUBAME, Qiao WANG, Miguel GUERRERO
-
Patent number: 12682531Abstract: In various examples, a technique for audio-driven facial animation with adaptive speech includes determining that a rate of speech associated with an audio segment exceeds a threshold. The technique also includes based at least on the rate of speech exceeding the threshold, upsampling a first set of features associated with the audio segment into a second set of features that is different in size than the first set of features. The technique further includes generating, using one or more machine learning models and based at least on at least a subset of the second set of features, a facial animation output corresponding to the audio segment.Type: GrantFiled: August 1, 2023Date of Patent: July 14, 2026Assignee: NVIDIA CORPORATIONInventors: Zhengyu Huang, Dmitry Korobchenko, Junjie Lai, Tao Li, Yeongho Seol, Rui Zhang, Weihua Zhang, Yingying Zhong
-
Patent number: 12626689Abstract: In various examples, determining emotion sequences for speech in conversational AI systems and applications is described herein. Systems and methods are disclosed that use one or more first machine learning models to determine a sequence of emotional states associated with audio data representing speech. To use the first machine learning model(s), the systems and methods may train the first machine learning model(s) using one or more second machine learning models, where the second machine learning model(s) is trained to determine scores indicating accuracies associated with sequences of emotional states. For instance, the second machine learning model(s) may be trained to determine the scores using audio data representing speech, sequences of emotional states associated with the speech, and indications of which sequences of emotional states better represent the speech as compared to other sequences of emotional states.Type: GrantFiled: August 1, 2023Date of Patent: May 12, 2026Assignee: NVIDIA CorporationInventors: Ilia Fedorov, Dmitry Korobchenko
-
Publication number: 20260105672Abstract: In various examples, systems and methods are disclosed relating to animating virtual or digital actors or avatars using audio-driven animation. A system can identify an animation for a mesh corresponding to audio data and an indication of a speaking style. The system can generate a plurality of vertex deltas using the animation and a neutral pose for the mesh. The system can update, using the plurality of vertex deltas, the audio data, and the indication of the speaking style, a machine-learning model to generate output vertex deltas for the mesh given an input speaking style and input audio data.Type: ApplicationFiled: October 16, 2024Publication date: April 16, 2026Applicant: NVIDIA CorporationInventors: Yeongho SEOL, Zhengyu HUANG, Roger BLANCO RIBERA, Dmitry KOROBCHENKO
-
Publication number: 20260105671Abstract: In various examples, systems and methods are disclosed relating to animating virtual or digital actors or avatars using audio-driven animation. A system can identify an animation for a mesh corresponding to audio data and an indication of a speaking style. The system can generate a plurality of vertex deltas using the animation and a neutral pose for the mesh. The system can update, using the plurality of vertex deltas, the audio data, and the indication of the speaking style, a machine-learning model to generate output vertex deltas for the mesh given an input speaking style and input audio data.Type: ApplicationFiled: October 16, 2024Publication date: April 16, 2026Applicant: NVIDIA CorporationInventors: Yeongho SEOL, Zhengyu HUANG, Roger BLANCO RIBERA, Dmitry KOROBCHENKO
-
Publication number: 20250272901Abstract: In various examples, determining emotional states for speech in conversational artificial intelligence (AI) and/or digital avatar systems and applications is descried herein. Systems and methods are disclosed that use one or more machine learning models to determine one or more emotional states associated with speech, where the machine learning model(s) may be trained using various processes. For instance, in some examples, the machine learning model(s) may be trained during a first training process to determine probabilities for distributions of values, where the distributions model different emotional states. For example, a distribution may include a first value for angry, a second value for happy, a third value for sad, and/or so forth. Additionally, or alternatively, in some examples, the machine learning model(s) may be trained during a second training process to more precisely determine the actual emotional states (and/or the probabilities) based on training data representing human feedback.Type: ApplicationFiled: February 26, 2024Publication date: August 28, 2025Inventors: Ilia Fedorov, Dmitry Korobchenko
-
Publication number: 20250272903Abstract: In various examples, using spatial relationships for animation retargeting in digital avatar systems and applications is described herein. Systems and methods are disclosed that determine constraints using first points (e.g., first vertices) associated with joints and/or a mesh of a source character and second points (e.g., second vertices) associated with joints and/or a mesh of a target character. As described herein, the constraints may include, but are not limited to, one or more of deformation constrains, interaction constraints, feet constraints, and angle constraints. Systems and methods are further disclosed that then use the constraints when performing optimization for animation retargeting of the target character. In some examples, the optimization is performed in vertex space, performed with respect to joints (e.g., rotations and/or transformations of the joints), and/or performed using one or more techniques, such as gradient descent.Type: ApplicationFiled: February 27, 2024Publication date: August 28, 2025Inventors: Evgenii Tumanov, Ivan Afanasev, Dmitry Korobchenko, Lina Halper, David Sebastian Minor
-
Publication number: 20250061634Abstract: Systems and methods of the present disclosure include animating virtual avatars or agents according to input audio and one or more selected or determined emotions and/or styles. For example, a deep neural network can be trained to output motion or deformation information for a character that is representative of the character uttering speech contained in audio input. The character can have different facial components or regions (e.g., head, skin, eyes, tongue) modeled separately, such that the network can output motion or deformation information for each of these different facial components. During training, the network can use a transformer-based audio encoder with locked parameters to train an associated decoder using a weighted feature vector. The network output can be provided to a renderer to generate audio-driven facial animation that is emotion-accurate.Type: ApplicationFiled: August 28, 2023Publication date: February 20, 2025Inventors: Zhengyu Huang, Rui Zhang, Tao Li, Yingying Zhong, Weihua Zhang, Junjie Lai, Yeongho Seol, Dmitry Korobchenko, Simon Yuen
-
Publication number: 20250046298Abstract: In various examples, determining emotion sequences for speech in conversational AI systems and applications is described herein. Systems and methods are disclosed that use one or more first machine learning models to determine a sequence of emotional states associated with audio data representing speech. To use the first machine learning model(s), the systems and methods may train the first machine learning model(s) using one or more second machine learning models, where the second machine learning model(s) is trained to determine scores indicating accuracies associated with sequences of emotional states. For instance, the second machine learning model(s) may be trained to determine the scores using audio data representing speech, sequences of emotional states associated with the speech, and indications of which sequences of emotional states better represent the speech as compared to other sequences of emotional states.Type: ApplicationFiled: August 1, 2023Publication date: February 6, 2025Inventors: Ilia Fedorov, Dmitry Korobchenko
-
Publication number: 20250029307Abstract: In various examples, a technique for audio-driven facial animation with adaptive speech includes determining that a rate of speech associated with an audio segment exceeds a threshold. The technique also includes based at least on the rate of speech exceeding the threshold, upsampling a first set of features associated with the audio segment into a second set of features that is different in size than the first set of features. The technique further includes generating, using one or more machine learning models and based at least on at least a subset of the second set of features, a facial animation output corresponding to the audio segment.Type: ApplicationFiled: August 1, 2023Publication date: January 23, 2025Inventors: Zhengyu HUANG, Dmitry KOROBCHENKO, Junjie LAI, Tao LI, Yeongho SEOL, Rui ZHANG, Weihua ZHANG, Yingying ZHONG
-
Publication number: 20240412440Abstract: In various examples, techniques are described for animating characters by decoupling portions of a face from other portions of the face. Systems and methods are disclosed that use one or more neural networks to generate high-fidelity facial animation using inputted audio data. In order to generate the high-fidelity facial animations, the systems and methods may decouple effects of implicit emotional states from effects of audio on the facial animations during training of the neural network(s). For instance, the training may cause the audio to drive the lower face animations while the implicit emotional states drive the upper face animations. In some examples, in order to encourage more expressive expressions, adversarial training is further used to learn a discriminator that predicts if generated emotional states are from real distribution.Type: ApplicationFiled: June 6, 2023Publication date: December 12, 2024Inventors: Rui Zhang, Zhengyu Huang, Lance Li, Weihua Zhang, Yingying Zhong, Junjie Lai, Yeongho Seol, Dmitry Korobchenko
-
Patent number: 11954791Abstract: Approaches in accordance with various embodiments provide for fluid simulation with substantially reduced time and memory requirements with respect to conventional approaches. In particular, various embodiments can perform time and energy efficient, large scale fluid simulation on processing hardware using a method that does not solve for the Navier-Stokes equations to enforce incompressibility. Instead, various embodiments generate a density tensor and rigid body map tensor for a large number of particles contained in a sub-domain. Collectively, the density tensor and rigid body map may represent input channels of a network with three spatial-dimensions. The network may apply a series of operations to the input channels to predict an updated position and updated velocity for each particle at the end of a frame. Such approaches can handle tens of millions of particles within a virtually unbounded simulation domain, as compared to classical approaches that solve for the Navier-Stokes equations.Type: GrantFiled: May 23, 2022Date of Patent: April 9, 2024Assignee: Nvidia CorporationInventors: Evgenii Tumanov, Dmitry Korobchenko, Alexey Solovey
-
Publication number: 20220358710Abstract: Approaches in accordance with various embodiments provide for fluid simulation with substantially reduced time and memory requirements with respect to conventional approaches. In particular, various embodiments can perform time and energy efficient, large scale fluid simulation on processing hardware using a method that does not solve for the Navier-Stokes equations to enforce incompressibility. Instead, various embodiments generate a density tensor and rigid body map tensor for a large number of particles contained in a sub-domain. Collectively, the density tensor and rigid body map may represent input channels of a network with three spatial-dimensions. The network may apply a series of operations to the input channels to predict an updated position and updated velocity for each particle at the end of a frame. Such approaches can handle tens of millions of particles within a virtually unbounded simulation domain, as compared to classical approaches that solve for the Navier-Stokes equations.Type: ApplicationFiled: May 23, 2022Publication date: November 10, 2022Inventors: Evgenii Tumanov, Dmitry Korobchenko, Alexey Solovey
-
Patent number: 11341710Abstract: Approaches in accordance with various embodiments provide for fluid simulation with substantially reduced time and memory requirements with respect to conventional approaches. In particular, various embodiments can perform time and energy efficient, large scale fluid simulation on processing hardware using a method that does not solve for the Navier-Stokes equations to enforce incompressibility. Instead, various embodiments generate a density tensor and rigid body map tensor for a large number of particles contained in a sub-domain. Collectively, the density tensor and rigid body map may represent input channels of a network with three spatial-dimensions. The network may apply a series of operations to the input channels to predict an updated position and updated velocity for each particle at the end of a frame. Such approaches can handle tens of millions of particles within a virtually unbounded simulation domain, as compared to classical approaches that solve for the Navier-Stokes equations.Type: GrantFiled: November 21, 2019Date of Patent: May 24, 2022Assignee: NVIDIA CORPORATIONInventors: Evgenii Tumanov, Dmitry Korobchenko, Alexey Solovey
-
Publication number: 20210158603Abstract: Approaches in accordance with various embodiments provide for fluid simulation with substantially reduced time and memory requirements with respect to conventional approaches. In particular, various embodiments can perform time and energy efficient, large scale fluid simulation on processing hardware using a method that does not solve for the Navier-Stokes equations to enforce incompressibility. Instead, various embodiments generate a density tensor and rigid body map tensor for a large number of particles contained in a sub-domain. Collectively, the density tensor and rigid body map may represent input channels of a network with three spatial-dimensions. The network may apply a series of operations to the input channels to predict an updated position and updated velocity for each particle at the end of a frame. Such approaches can handle tens of millions of particles within a virtually unbounded simulation domain, as compared to classical approaches that solve for the Navier-Stokes equations.Type: ApplicationFiled: November 21, 2019Publication date: May 27, 2021Inventors: Evegny Tumanov, Dmitry Korobchenko, Alexey Solovey