Patents by Inventor Sergey Tulyakov
Sergey Tulyakov has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Publication number: 20260205619Abstract: A method, computer readable medium, and system are disclosed for action video generation. The method includes the steps of generating, by a recurrent neural network, a sequence of motion vectors from a first set of random variables and receiving, by a generator neural network, the sequence of motion vectors and a content vector sample. The sequence of motion vectors and the content vector sample are sampled by the generator neural network to produce a video clip.Type: ApplicationFiled: December 22, 2025Publication date: July 16, 2026Inventors: Ming-Yu Liu, Xiaodong Yang, Jan Kautz, Sergey Tulyakov
-
Patent number: 12670362Abstract: Aspects of the present disclosure involve a system comprising a computer-readable storage medium storing a program and method for video synthesis. The program and method provide for accessing a primary generative adversarial network (GAN) comprising a pre-trained image generator, a motion generator comprising a plurality of neural networks, and a video discriminator; generating an updated GAN based on the primary GAN, by performing operations comprising identifying input data of the updated GAN, the input data comprising an initial latent code and a motion domain dataset, training the motion generator based on the input data, and adjusting weights of the plurality of neural networks of the primary GAN based on an output of the video discriminator; and generating a synthesized video based on the primary GAN and the input data.Type: GrantFiled: September 30, 2021Date of Patent: June 30, 2026Assignee: Snap Inc,.Inventors: Menglei Chai, Kyle Olszewski, Jian Ren, Yu Tian, Sergey Tulyakov
-
Patent number: 12663914Abstract: A system of machine learning schemes can be configured to efficiently perform image processing tasks on a user device, such as a mobile phone. The system can selectively detect and transform individual regions within each frame of a live streaming video. The system can selectively partition and toggle image effects within the live streaming video.Type: GrantFiled: August 9, 2023Date of Patent: June 23, 2026Inventors: Theresa Barton, Yanping Chen, Jaewook Chung, Christopher Yale Crutchfield, Aymeric Damien, Sergei Kotcur, Igor Kudriashov, Sergey Tulyakov, Andrew Wan, Emre Yamangil
-
Publication number: 20260162367Abstract: A system and method are described for generating 3D garments from two-dimensional (2D) scribble images drawn by users. The system includes a conditional 2D generator, a conditional 3D generator, and two intermediate media including dimension-coupling color-density pairs and flat point clouds that bridge the gap between dimensions. Given a scribble image, the 2D generator synthesizes dimension-coupling color-density pairs including the RGB projection and density map from the front and rear views of the scribble image. A density-aware sampling algorithm converts the 2D dimension-coupling color-density pairs into a 3D flat point cloud representation, where the depth information is ignored. The 3D generator predicts the depth information from the flat point cloud. Dynamic variations per garment due to deformations resulting from a wearer's pose as well as irregular wrinkles and folds may be bypassed by taking advantage of 2D generative models to bridge the dimension gap in a non-parametric way.Type: ApplicationFiled: April 15, 2025Publication date: June 11, 2026Inventors: Panagiotis Achlioptas, Menglei Chai, Hsin-Ying Lee, Kyle Olszewski, Jian Ren, Sergey Tulyakov
-
Patent number: 12645936Abstract: Techniques for training a neural network having a plurality of computational layers with associated weights and activations for computational layers in fixed-point formats include determining an optimal fractional length for weights and activations for the computational layers; training a learned clipping-level with fixed-point quantization using a PACT process for the computational layers; and quantizing on effective weights that fuses a weight of a convolution layer with a weight and running variance from a batch normalization layer. A fractional length for weights of the computational layers is determined from current values of weights using the determined optimal fractional length for the weights of the computational layers. A fixed-point activation between adjacent computational layers is related using PACT quantization of the clipping-level and an activation fractional length from a node in a following computational layer.Type: GrantFiled: December 31, 2021Date of Patent: June 2, 2026Assignee: Snap Inc.Inventors: Sumant Milind Hanumante, Qing Jin, Sergei Korolev, Denys Makoviichuk, Jian Ren, Dhritiman Sagar, Patrick Timothy McSweeney Simons, Sergey Tulyakov, Yang Wen, Richard Zhuang
-
Publication number: 20260148438Abstract: Hierarchical patch-wise diffusion models (HPDMs) use a diffusion paradigm that learns a hierarchical distribution of patches instead of whole videos for efficient patch-wise training of diffusion models. To enforce consistency between the patches, deep context fusion may be used to propagate the context information from low-scale to high-scale patches in a hierarchical manner. To accelerate patch-wise training and inference, adaptive computation also may be used to allocate more computational resources and network capacity towards coarse image details and to cheapen synthesis of high-frequency texture details. All the processing stages are jointly trained to provide spatially aligned global context to the higher levels of the cascade. As a result, the model does not operate on the full-resolution inputs, which allows the model to be trained on high-resolution video datasets in an end-to-end fashion.Type: ApplicationFiled: December 11, 2025Publication date: May 28, 2026Inventors: Willi Menapace, Aliaksandr Siarohin, Ivan Skorokhodov, Sergey Tulyakov
-
Publication number: 20260141714Abstract: A mobile vision transformer network for use on mobile devices, such as smart eyewear devices and other augmented reality (AR) and virtual reality (VR) devices. The mobile vision transformer network considers factors including number of parameters, latency, and model performance, as they reflect disk storage, mobile frames per second (FPS), and application quality, respectively. The mobile vision transformer network processes images, e.g., for image classification, segmentation, and detection. The mobile vision transformer network has a fine-grained architecture including a search algorithm performing latency-driven slimming that jointly improves model size and speed.Type: ApplicationFiled: January 19, 2026Publication date: May 21, 2026Inventors: Jian Ren, Yanyu Li, Ju Hu, Yang Wen, Georgios Evangelidis, Sergey Tulyakov, Kamyar Salahi
-
Patent number: 12620216Abstract: Described is a system for improving machine learning models. In some cases, the system improves such models by identifying an autoencoder for a latent diffusion machine learning model, the latent diffusion machine learning model is trained to receive text as input and output an image based on the received text. The system identifies a number of channels in a decoder of the autoencoder, the decoder being configured to receive latent features as input and output images. The system further identifies a performance characteristic of the decoder and changes the node topology of the decoder based on the performance characteristic to generate an updated decoder. The system retrains the latent diffusion machine learning model using the updated decoder by inputting latent features to the updated decoder, receiving an outputted image from the updated decoder, and updating one or more weights of the decoder based on an assessment of the outputted image.Type: GrantFiled: December 29, 2023Date of Patent: May 5, 2026Assignee: SNAP INC.Inventors: Pavlo Chemerys, Colin Eles, Ju Hu, Qing Jin, Yanyu Li, Ergeta Muca, Jian Ren, Dhritiman Sagar, Aleksei Stoliar, Sergey Tulyakov, Huan Wang
-
Publication number: 20260112124Abstract: Examples relate to systems and methods for generating digital effects experiences. The system performs operations including accessing a set of instructions that defines a digital effects experience. The system processes the set of instructions by a generative machine learning model to generate one or more digital effects comprising the digital effects experience. The system continuously processes one or more inputs, received by a user device, while the one or more digital effects comprising the digital effects experience are presented on the user device, along with the set of instructions in real time by the generative machine learning model to update presentation of the one or more digital effects comprising the digital effects experience.Type: ApplicationFiled: October 18, 2024Publication date: April 23, 2026Inventors: Willi Menapace, Robert Cornelius Murphy, Aliasksandr Siarohin, Aleksei Stoliar, Sergey Tulyakov
-
Patent number: 12608593Abstract: Systems and methods herein describe an image compression system. The image compression system generates a first generative adversarial network (GAN), identifies a threshold, based on the threshold, generates a second GAN by pruning channels of the first GAN, trains the second GAN using similarity-based knowledge distillation from the first GAN, and stores the trained second GAN.Type: GrantFiled: December 21, 2021Date of Patent: April 21, 2026Assignee: Snap Inc.Inventors: Jian Ren, Oliver Woodford, Sergey Tulyakov, Jiazhuo Wang, Qing Jin
-
Publication number: 20260089303Abstract: A method for generating photorealistic 4D scenes from text inputs is disclosed. The method utilizes a text-to-video diffusion model to generate a reference video and a freeze-time video. A canonical 3D representation is reconstructed using deformable 3D Gaussian Splats (D-3DGS) based on the freeze-time video. Temporal deformations are learned to capture dynamic interactions in the reference video. The method employs a novel Score Distillation Sampling strategy combining multi-view and temporal aspects to enhance consistency and robustness. The resulting 4D scenes feature multiple objects interacting with detailed background environments, viewable from different angles and times. The method enables flexible camera control and integration with augmented and virtual reality applications. Some examples include features such as image-to-4D generation.Type: ApplicationFiled: September 23, 2024Publication date: March 26, 2026Inventors: Hsin-Ying Lee, Willi Menapace, Aliaksandr Siarohin, Sergey Tulyakov, Chaoyang Wang, Heng Yu, Peiye Zhuang
-
Patent number: 12579406Abstract: Systems and methods herein describe an image compression system. The image compression system generates a first generative adversarial network (GAN), identifies a threshold, based on the threshold, generates a second GAN by pruning channels of the first GAN, trains the second GAN using similarity-based knowledge distillation from the first GAN, and stores the trained second GAN.Type: GrantFiled: December 21, 2021Date of Patent: March 17, 2026Assignee: Snap Inc.Inventors: Jian Ren, Oliver Woodford, Sergey Tulyakov, Jiazhuo Wang, Qing Jin
-
Publication number: 20260073280Abstract: Described is a system performing operations comprising deriving, based on an evaluation of variations of a first machine learning model, a second machine learning model, the second machine learning model having layers with precisions assigned based on the evaluation, training the second machine learning model to reduce error between a first test output generated by the first machine learning model based on a first test input and a second test output generated by the second machine learning model based on the first test input, training the second machine learning model to reduce error between a third test output generated by the second machine learning model based on a second test input and a ground truth output associated with the second test input, providing an input for the second machine learning model, and generating an output using the second machine learning model based on the input.Type: ApplicationFiled: September 9, 2024Publication date: March 12, 2026Inventors: Junli Cao, Ju Hu, Yerlan Idelbayev, Anil Kag, Yanyu Li, Jian Ren, Dhritiman Sagar, Yang Sui, Sergey Tulyakov
-
Publication number: 20260073578Abstract: The present disclosure addresses technological challenges arising in the field of artificial intelligence (AI) with respect to inefficient use of computing resources and runtime delay. In particular, the present disclosure provides for development of a machine learning model that generates an image sample for a video in a single forward pass. The development of this machine learning model uses an adversarial training approach involving training two machine learning models, a generator model and a discriminator model. With the generator model trained in this way, the generator model can be used to generate image samples for a video in a single forward pass.Type: ApplicationFiled: September 10, 2024Publication date: March 12, 2026Inventors: Junli Cao, Anil Kag, Yanyu Li, Willi Menapace, Jian Ren, Aliaksandr Siarohin, Ivan Skorokhodov, Sergey Tulyakov, Yushu Wu, Zhixing Zhang
-
Publication number: 20260073667Abstract: Described herein are techniques for personalizing vision-language models (VLMs) to understand user-specific concepts, without modifying original model weights. Pre-trained VLMs are augmented with external concept heads that identify user-specific concepts in input images. A concept embedding vector is computed to represent the user-specific concept within an intermediate feature space of the VLM through iterative optimization. Then, when processing an input image and the concept is detected, the concept embedding is appended to image features extracted by a vision encoder of the VLM. Personalized textual outputs incorporating the user-specific concept are generated in response to input images and language instructions. Regularization techniques balance attention between the appended concept embedding and original image features, maintaining alignment between outputs and inputs.Type: ApplicationFiled: September 10, 2024Publication date: March 12, 2026Inventors: Kfir Aberman, Yuval Alaluf, Daniel Cohen-Or, Sergey Tulyakov
-
Publication number: 20260057606Abstract: Systems and methods for generating static and articulated 3D assets are provided that include a 3D autodecoder at their core. The 3D autodecoder framework embeds properties learned from the target dataset in the latent space, which can then be decoded into a volumetric representation for rendering view-consistent appearance and geometry. The appropriate intermediate volumetric latent space is then identified and robust normalization and de-normalization operations are implemented to learn a 3D diffusion from 2D images or monocular videos of rigid or articulated objects. The methods are flexible enough to use either existing camera supervision or no camera information at all—instead efficiently learning the camera information during training.Type: ApplicationFiled: October 30, 2025Publication date: February 26, 2026Inventors: Evangelos Ntavelis, Kyle Olszewski, Aliaksandr Siarohin, Sergey Tulyakov
-
Publication number: 20260059173Abstract: Automatic captioning pipelines and methods for automatically annotating video data with subtitles, which can be obtained using automatic speech recognition (ASR). An automatic captioning pipeline with inputs of multimodal data scales up the dataset of high-quality video-caption pairs. The automatic captioning pipeline generates video-caption pairs by establishing and using a large video-language dataset along with an automatic captioning approach leveraging multimodal inputs, such as textual video description, subtitles, and individual video frames.Type: ApplicationFiled: October 31, 2025Publication date: February 26, 2026Inventors: Tsai-Shien Chen, Yuwei Fang, Hsin-Ying Lee, Jian Ren, Aliaksandr Siarohin, Sergey Tulyakov
-
Publication number: 20260051121Abstract: An environment synthesis framework generates virtual environments from a synthesized two-dimensional (2D) satellite map of a geographic area, a three-dimensional (3D) voxel environment, and a voxel-based neural rendering framework. In an example implementation, the synthesized 2D satellite map is generated by a map synthesis generative adversarial network (GAN) which is trained using sample city datasets. The multi-stage framework lifts the 2D map into a set of 3D octrees, generates an octree-based 3D voxel environment, and then converts it into a texturized 3D virtual environment using a neural rendering GAN and a set of pseudo ground truth images. The resulting 3D virtual environment is texturized, lifelike, editable, traversable in virtual reality (VR) and augmented reality (AR) experiences, and very large in scale.Type: ApplicationFiled: October 23, 2025Publication date: February 19, 2026Inventors: Menglei Chai, Hsin-Ying Lee, Chieh Lin, Willi Menapace, Aliaksandr Siarohin, Sergey Tulyakov
-
Publication number: 20260051023Abstract: A neural light field (NeLF) that runs real-time on mobile devices for neural rendering of three dimensional (3D) scenes, referred to as MobileR2L. The MobileR2L architecture runs efficiently on mobile devices with low latency and small size, and it achieves high-resolution generation while maintaining real-time inference for both synthetic and real-world 3D scenes on mobile devices. The MobileR2L has a network backbone including a convolutional layer embedding an input image at a resolution, residual blocks uploading the embedded image, and super-resolution modules receiving the uploaded embedded image and rendering an output image having a higher resolution than the embedded image. The convolution layer generates a number of rays equal to a number of pixels in the input image, where a partial number of the rays is uploaded to the super-resolution modules.Type: ApplicationFiled: October 23, 2025Publication date: February 19, 2026Inventors: Jian Ren, Pavlo Chemerys, Vladislav Shakhrai, Ju Hu, Denys Makoviichuk, Sergey Tulyakov, Junli Cao
-
Patent number: 12555371Abstract: A mobile vision transformer network for use on mobile devices, such as smart eyewear devices and other augmented reality (AR) and virtual reality (VR) devices. The mobile vision transformer network considers factors including number of parameters, latency, and model performance, as they reflect disk storage, mobile frames per second (FPS), and application quality, respectively. The mobile vision transformer network processes images, e.g., for image classification, segmentation, and detection. The mobile vision transformer network has a fine-grained architecture including a search algorithm performing latency-driven slimming that jointly improves model size and speed.Type: GrantFiled: December 14, 2022Date of Patent: February 17, 2026Assignee: Snap Inc.Inventors: Jian Ren, Yanyu Li, Ju Hu, Yang Wen, Georgios Evangelidis, Sergey Tulyakov, Kamyar Salahi