Patents by Inventor Bryan Russell
Bryan Russell has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Publication number: 20260134655Abstract: Embodiments are disclosed for a video frame encoding system trained to generate a visual features from video frames of a video sequence using a lightweight visual encoder. The method may include generating, by a visual encoder, first visual features for a first video frame of a video sequence. The disclosed systems and methods further comprise generating, by a lightweight visual encoder, first residual visual features for a first residual video frame, wherein the first residual video frame is based on the first video frame of the video sequence and a second video frame of the video sequence subsequent to the first video frame. The disclosed systems and methods further comprise generating second visual features for the second video frame of the video sequence by aggregating the first visual features and the first residual visual features.Type: ApplicationFiled: November 8, 2024Publication date: May 14, 2026Applicant: Adobe Inc.Inventors: Mattia SOLDAN, Fabian David CABA HEILBRON, Bryan RUSSELL, Josef SIVIC
-
Publication number: 20260094440Abstract: Embodiments are trained to generate a timeline of digital media assets based on natural language instructions using a machine learning model. The method may include receiving an input including digital media assets, an input visual timeline, and text input, where the text input indicates a natural language instruction describing a modification to the input visual timeline using the digital media assets. The disclosed systems and methods further comprise generating a first set of tokens for the digital media assets, a second set of tokens for the input visual timeline, and a third set of tokens for the text input. The disclosed systems and methods further comprise processing, by a large language model, the first set of tokens, the second set of tokens, and the third set of tokens to generate an output set of tokens and generating a reconstructed visual timeline using the output set of tokens.Type: ApplicationFiled: September 27, 2024Publication date: April 2, 2026Applicant: Adobe Inc.Inventors: Fabian David CABA HEILBRON, Jui-Hsien WANG, Bryan RUSSELL, Luis Alejandro PARDO GONZALEZ, Josef SIVIC
-
Publication number: 20260002354Abstract: In one example, an apparatus, which may take the form of a portion of a shed, includes a first structure such as a first panel, a second structure such as a second panel, and a connector configured to releasably connect the first structure and the second structure together. The connector includes a body, an arm that is integral with the body and configured to be removably received in a receptacle cooperatively defined by the first structure and the second structure. The connector further includes a web connecting the arm to the body.Type: ApplicationFiled: June 9, 2025Publication date: January 1, 2026Inventors: Neil Ben Watson, Kenneth William Taylor, Jeremy Ross Wendt, Nathan Craig Jones, Paul Craig Branch, Bryan Russell, Dennis Norman
-
Publication number: 20250316062Abstract: Embodiments are disclosed for correlating video sequences and audio sequences by a media recommendation system using a trained encoder network.Type: ApplicationFiled: June 23, 2025Publication date: October 9, 2025Applicant: Adobe Inc.Inventors: Justin SALAMON, Bryan RUSSELL, Didac SURIS COLL-VINENT
-
Patent number: 12340563Abstract: Embodiments are disclosed for correlating video sequences and audio sequences by a media recommendation system using a trained encoder network.Type: GrantFiled: May 11, 2022Date of Patent: June 24, 2025Assignee: Adobe Inc.Inventors: Justin Salamon, Bryan Russell, Didac Suris Coll-Vinent
-
Publication number: 20240419726Abstract: Techniques for learning to personalize vision-language models through meta-personalization are described. In one embodiment, one or more processing devices lock a pre-trained vision-language model (VLM) during a training phase. The processing devices train the pre-trained VLM to augment a text encoder of the pre-trained VLM with a set of general named video instances to form a meta-personalized VLM, the meta-personalized VLM to include global category features. The processing devices test the meta-personalized VLM to adapt the text encoder with a set of personal named video instances to form a personal VLM, the personal VLM comprising the global category features personalized with a set of personal instance weights to form a personal instance token associated with the user. Other embodiments are described and claimed.Type: ApplicationFiled: June 15, 2023Publication date: December 19, 2024Applicant: Adobe Inc.Inventors: Simon Jenni, Fabian David Caba Heilbron, Chun-Hsiao Yeh, Bryan Russell, Josef Sivic
-
Publication number: 20240386048Abstract: Embodiments are disclosed for an audio recommendation system trained to recommend music audio sequences for pairing with query video sequences using neural networks. In particular, in one or more embodiments, the disclosed systems and methods comprise receiving an input including a query video sequence and natural language text. The disclosed systems and methods further comprise generating a fused visual-text embedding based on a visual embedding and a text embedding corresponding to the input. The disclosed systems and methods further comprise comparing audio embeddings for music audio sequences of a music audio sequences database with the fused visual-text embedding. The disclosed systems and methods further comprise determining a music audio sequence from the music audio sequences database as the recommended music audio sequence for pairing with the query video sequence based on a similarity metric calculated between an audio embedding for the music audio sequence and the fused visual-text embedding.Type: ApplicationFiled: May 17, 2023Publication date: November 21, 2024Applicant: Adobe Inc.Inventors: Bryan RUSSELL, Justin SALAMON, Daniel McKEE, Josef SIVIC
-
Patent number: 12118787Abstract: Methods, system, and computer storage media are provided for multi-modal localization. Input data comprising two modalities, such as image data and corresponding text or audio data, may be received. A phrase may be extracted from the text or audio data, and a neural network system may be utilized to spatially and temporally localize the phrase within the image data. The neural network system may include a plurality of cross-modal attention layers that each compare features across the first and second modalities without comparing features of the same modality. Using the cross-modal attention layers, a region or subset of pixels within one or more frames of the image data may be identified as corresponding to the phrase, and a localization indicator may be presented for display with the image data. Embodiments may also include unsupervised training of the neural network system.Type: GrantFiled: October 12, 2021Date of Patent: October 15, 2024Assignee: ADOBE INC.Inventors: Hailin Jin, Bryan Russell, Reuben Xin Hong Tan
-
Patent number: 11949964Abstract: Systems, methods, and non-transitory computer-readable media are disclosed for automatic tagging of videos. In particular, in one or more embodiments, the disclosed systems generate a set of tagged feature vectors (e.g., tagged feature vectors based on action-rich digital videos) to utilize to generate tags for an input digital video. For instance, the disclosed systems can extract a set of frames for the input digital video and generate feature vectors from the set of frames. In some embodiments, the disclosed systems generate aggregated feature vectors from the feature vectors. Furthermore, the disclosed systems can utilize the feature vectors (or aggregated feature vectors) to identify similar tagged feature vectors from the set of tagged feature vectors. Additionally, the disclosed systems can generate a set of tags for the input digital videos by aggregating one or more tags corresponding to identified similar tagged feature vectors.Type: GrantFiled: September 9, 2021Date of Patent: April 2, 2024Assignee: Adobe Inc.Inventors: Bryan Russell, Ruppesh Nalwaya, Markus Woodson, Joon-Young Lee, Hailin Jin
-
Publication number: 20230368503Abstract: Embodiments are disclosed for correlating video sequences and audio sequences by a media recommendation system using a trained encoder network.Type: ApplicationFiled: May 11, 2022Publication date: November 16, 2023Applicant: Adobe Inc.Inventors: Justin SALAMON, Bryan RUSSELL, Didac SURIS COLL-VINENT
-
Patent number: 11721056Abstract: In some embodiments, a model training system obtains a set of animation models. For each of the animation models, the model training system renders the animation model to generate a sequence of video frames containing a character using a set of rendering parameters and extracts joint points of the character from each frame of the sequence of video frames. The model training system further determines, for each frame of the sequence of video frames, whether a subset of the joint points are in contact with a ground plane in a three-dimensional space and generates contact labels for the subset of the joint points. The model training system trains a contact estimation model using training data containing the joint points extracted from the sequences of video frames and the generated contact labels. The contact estimation model can be used to refine a motion model for a character.Type: GrantFiled: January 12, 2022Date of Patent: August 8, 2023Assignee: Adobe Inc.Inventors: Jimei Yang, Davis Rempe, Bryan Russell, Aaron Hertzmann
-
Publication number: 20230115551Abstract: Methods, system, and computer storage media are provided for multi-modal localization. Input data comprising two modalities, such as image data and corresponding text or audio data, may be received. A phrase may be extracted from the text or audio data, and a neural network system may be utilized to spatially and temporally localize the phrase within the image data. The neural network system may include a plurality of cross-modal attention layers that each compare features across the first and second modalities without comparing features of the same modality. Using the cross-modal attention layers, a region or subset of pixels within one or more frames of the image data may be identified as corresponding to the phrase, and a localization indicator may be presented for display with the image data. Embodiments may also include unsupervised training of the neural network system.Type: ApplicationFiled: October 12, 2021Publication date: April 13, 2023Inventors: Hailin Jin, Bryan Russell, Reuben Xin Hong Tan
-
Patent number: 11514632Abstract: This disclosure describes methods, non-transitory computer readable storage media, and systems that utilize a contrastive perceptual loss to modify neural networks for generating synthetic digital content items. For example, the disclosed systems generate a synthetic digital content item based on a guide input to a generative neural network. The disclosed systems utilize an encoder neural network to generate encoded representations of the synthetic digital content item and a corresponding ground-truth digital content item. Additionally, the disclosed systems sample patches from the encoded representations of the encoded digital content items and then determine a contrastive loss based on the perceptual distances between the patches in the encoded representations. Furthermore, the disclosed systems jointly update the parameters of the generative neural network and the encoder neural network utilizing the contrastive loss.Type: GrantFiled: November 6, 2020Date of Patent: November 29, 2022Assignee: Adobe Inc.Inventors: Bryan Russell, Taesung Park, Richard Zhang, Junyan Zhu, Alexander Andonian
-
Publication number: 20220148242Abstract: This disclosure describes methods, non-transitory computer readable storage media, and systems that utilize a contrastive perceptual loss to modify neural networks for generating synthetic digital content items. For example, the disclosed systems generate a synthetic digital content item based on a guide input to a generative neural network. The disclosed systems utilize an encoder neural network to generate encoded representations of the synthetic digital content item and a corresponding ground-truth digital content item. Additionally, the disclosed systems sample patches from the encoded representations of the encoded digital content items and then determine a contrastive loss based on the perceptual distances between the patches in the encoded representations. Furthermore, the disclosed systems jointly update the parameters of the generative neural network and the encoder neural network utilizing the contrastive loss.Type: ApplicationFiled: November 6, 2020Publication date: May 12, 2022Inventors: Bryan Russell, Taesung Park, Richard Zhang, Junyan Zhu, Alexander Andonian
-
Publication number: 20220139019Abstract: In some embodiments, a model training system obtains a set of animation models. For each of the animation models, the model training system renders the animation model to generate a sequence of video frames containing a character using a set of rendering parameters and extracts joint points of the character from each frame of the sequence of video frames. The model training system further determines, for each frame of the sequence of video frames, whether a subset of the joint points are in contact with a ground plane in a three-dimensional space and generates contact labels for the subset of the joint points. The model training system trains a contact estimation model using training data containing the joint points extracted from the sequences of video frames and the generated contact labels. The contact estimation model can be used to refine a motion model for a character.Type: ApplicationFiled: January 12, 2022Publication date: May 5, 2022Inventors: Jimei Yang, Davis Rempe, Bryan Russell, Aaron Hertzmann
-
Patent number: 11308329Abstract: A computer system is trained to understand audio-visual spatial correspondence using audio-visual clips having multi-channel audio. The computer system includes an audio subnetwork, video subnetwork, and pretext subnetwork. The audio subnetwork receives the two channels of audio from the audio-visual clips, and the video subnetwork receives the video frames from the audio-visual clips. In a subset of the audio-visual clips the audio-visual spatial relationship is misaligned, causing the audio-visual spatial cues for the audio and video to be incorrect. The audio subnetwork outputs an audio feature vector for each audio-visual clip, and the video subnetwork outputs a video feature vector for each audio-visual clip. The audio and video feature vectors for each audio-visual clip are merged and provided to the pretext subnetwork, which is configured to classify the merged vector as either having a misaligned audio-visual spatial relationship or not.Type: GrantFiled: May 7, 2020Date of Patent: April 19, 2022Assignee: Adobe Inc.Inventors: Justin Salamon, Bryan Russell, Karren Yang
-
Patent number: 11257298Abstract: Methods, systems, and non-transitory computer readable storage media are disclosed for reconstructing three-dimensional meshes from two-dimensional images of objects with automatic coordinate system alignment. For example, the disclosed system can generate feature vectors for a plurality of images having different views of an object. The disclosed system can process the feature vectors to generate coordinate-aligned feature vectors aligned with a coordinate system associated with an image. The disclosed system can generate a combined feature vector from the feature vectors aligned to the coordinate system. Additionally, the disclosed system can then generate a three-dimensional mesh representing the object from the combined feature vector.Type: GrantFiled: March 18, 2020Date of Patent: February 22, 2022Assignee: Adobe Inc.Inventors: Vladimir Kim, Pierre-alain Langlois, Oliver Wang, Matthew Fisher, Bryan Russell
-
Patent number: 11238634Abstract: In some embodiments, a motion model refinement system receives an input video depicting a human character and an initial motion model describing motions of individual joint points of the human character in a three-dimensional space. The motion model refinement system identifies foot joint points of the human character that are in contact with a ground plane using a trained contact estimation model. The motion model refinement system determines the ground plane based on the foot joint points and the initial motion model and constructs an optimization problem for refining the initial motion model. The optimization problem minimizes the difference between the refined motion model and the initial motion model under a set of plausibility constraints including constraints on the contact foot joint points and a time-dependent inertia tensor-based constraint. The motion model refinement system obtains the refined motion model by solving the optimization problem.Type: GrantFiled: April 28, 2020Date of Patent: February 1, 2022Assignee: Adobe Inc.Inventors: Jimei Yang, Davis Rempe, Bryan Russell, Aaron Hertzmann
-
Publication number: 20210409836Abstract: Systems, methods, and non-transitory computer-readable media are disclosed for automatic tagging of videos. In particular, in one or more embodiments, the disclosed systems generate a set of tagged feature vectors (e.g., tagged feature vectors based on action-rich digital videos) to utilize to generate tags for an input digital video. For instance, the disclosed systems can extract a set of frames for the input digital video and generate feature vectors from the set of frames. In some embodiments, the disclosed systems generate aggregated feature vectors from the feature vectors. Furthermore, the disclosed systems can utilize the feature vectors (or aggregated feature vectors) to identify similar tagged feature vectors from the set of tagged feature vectors. Additionally, the disclosed systems can generate a set of tags for the input digital videos by aggregating one or more tags corresponding to identified similar tagged feature vectors.Type: ApplicationFiled: September 9, 2021Publication date: December 30, 2021Inventors: Bryan Russell, Ruppesh Nalwaya, Markus Woodson, Joon-Young Lee, Hailin Jin
-
Patent number: 11189094Abstract: Techniques are disclosed for 3D object reconstruction using photometric mesh representations. A decoder is pretrained to transform points sampled from 2D patches of representative objects into 3D polygonal meshes. An image frame of the object is fed into an encoder to get an initial latent code vector. For each frame and camera pair from the sequence, a polygonal mesh is rendered at the given viewpoints. The mesh is optimized by creating a virtual viewpoint, rasterized to obtain a depth map. The 3D mesh projections are aligned by projecting the coordinates corresponding to the polygonal face vertices of the rasterized mesh to both selected viewpoints. The photometric error is determined from RGB pixel intensities sampled from both frames. Gradients from the photometric error are backpropagated into the vertices of the assigned polygonal indices by relating the barycentric coordinates of each image to update the latent code vector.Type: GrantFiled: August 5, 2020Date of Patent: November 30, 2021Assignee: Adobe, Inc.Inventors: Oliver Wang, Vladimir Kim, Matthew Fisher, Elya Shechtman, Chen-Hsuan Lin, Bryan Russell