Patents by Inventor Sai Bi
Sai Bi has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Patent number: 12682426Abstract: In some examples, a computing system accesses a field of view (FOV) image that has a field of view less than 360 degrees and has low dynamic range (LDR) values. The computing system estimates lighting parameters from a scene depicted in the FOV image and generates a lighting image based on the lighting parameters. The computing system further generates lighting features generated the lighting image and image features generated from the FOV image. These features are aggregated into aggregated features and a machine learning model is applied to the image features and the aggregated features to generate a panorama image having high dynamic range (HDR) values.Type: GrantFiled: August 25, 2023Date of Patent: July 14, 2026Assignee: Adobe Inc.Inventors: Mohammad Reza Karimi Dastjerdi, Yannick Hold-Geoffroy, Sai Bi, Jonathan Eisenmann, Jean-François Lalonde
-
Patent number: 12682547Abstract: Embodiments are configured to render 3D models using an importance sampling method. First, embodiments obtain a 3D model including a plurality of density values corresponding to a plurality of locations in a 3D space, respectively. Embodiments then sample the color information from within a random subset of the plurality of locations using a probability distribution based on the plurality of density values. Embodiments have a higher probability to sample each location within the random subset of locations if the location has a higher density probability. Embodiments then an image depicting a view of the 3D model based on the sampling within the random subset of the plurality of locations.Type: GrantFiled: November 1, 2023Date of Patent: July 14, 2026Assignee: ADOBE INC.Inventors: Milos Hasan, Iliyan Georgiev, Sai Bi, Julien Philip, Kalyan K. Sunkavalli, Xin Sun, Fujun Luan, Kevin James Blackburn-Matzen, Zexiang Xu, Kai Zhang
-
Publication number: 20260148487Abstract: In some embodiments, a computing system accesses multiple input images of a specular object with a scene. The computing system encodes near-field interreflections of the scene on the specular object to obtain a first set of feature representations of the specular object in multiple viewing directions based on the multiple input images. The computing system encodes far-field reflections of the scene on the specular object to obtain a second set of feature representations in the multiple viewing directions based on the multiple input images. The computing system determines a set of specular color values for the specular object in the multiple viewing directions based on the first set of feature representations and the second set of feature representations using a multi-layer perceptron algorithm. The computing system renders the specular object representation at least based on the set of specular color values using a neural rendering algorithm.Type: ApplicationFiled: November 25, 2024Publication date: May 28, 2026Inventors: Sai Bi, Zexiang Xu, Liwen Wu, Kalyan Sunkavalli, Kai Zhang, Iliyan Georgiev, Fujun Luan, Ravi Ramamoorthi
-
Publication number: 20260100010Abstract: Techniques for depth guided text-based editing of 3D neural radiance fields are provided. A method includes receiving input 2D images corresponding to views of a target and generating a 3D representation from the input 2D images. The 3D representation includes points forming a point cloud, where each point has a color and density value. The method also includes accumulating the color and density values to generate a volumetric 3D scene having a geometry, extracting distance maps from the volumetric 3D scene based on the geometry, and generating a plurality of masks associated with the target for each view. The method also includes aggregating the masks into the volumetric 3D scene using the geometry, providing the input 2D images, the masks, and the distance maps to a diffusion model, and modifying an appearance of the target in the volumetric 3D scene by providing a text command to the diffusion model.Type: ApplicationFiled: October 7, 2024Publication date: April 9, 2026Inventors: Sara Rojas Martinez, Julien Philip, Kai Zhang, Sai Bi, Fujun Luan, Kalyan Sunkavalli
-
Publication number: 20260087600Abstract: In implementing per-asset denoising for real-time rendering of neural radiance fields (NeRFs), a processing device receives a three-dimensional (3D) representation of a scene as a NeRF. The processing device generates an intermediate rendering of the scene using the NeRF. The intermediate rendering is denoised using a machine-learning model to generate a final rendering. The machine-learning model is trained on another rendering of this scene, which was rendered using a non-real-time, high-quality rendering scheme. In other words, the machine-learning model is optimized for each scene and provides a lightweight denoising network to provide real-time NeRF rendering while maintaining the high-quality visuals of non-real-time rendering schemes. The final rendering is then presented via a display device.Type: ApplicationFiled: September 24, 2024Publication date: March 26, 2026Applicants: Adobe Inc., THE REGENTS OF THE UNIVERSITY OF CALIFORNIAInventors: Sai Bi, Zexiang Xu, Xin Sun, Miloš Hašan, Kunal Gupta, Kevin Blackburn-Matzen, Kalyan Krishna Sunkavalli, Kai Zhang, Julien Olivier Victor Philip, Fujun Luan, Manmohan Chandraker, Iliyan Atanasov Georgiev
-
Patent number: 12586291Abstract: Embodiments are disclosed for fast large-scale radiance field reconstruction. A method of fast large-scale radiance field reconstruction may include receiving a sequence of input images that depict views of a scene and extracting, using an image encoder, image features from the sequence of input images. A first one or more machine learning models may generate a local volume based on the image features corresponding to one or more images from the sequence of input images. A second one or more machine learning models may generate a global volume based on the local volume. A novel view of the scene is synthesized based on the global volume.Type: GrantFiled: March 16, 2023Date of Patent: March 24, 2026Assignees: Adobe Inc., The Regents of the University of CalforniaInventors: Zexiang Xu, Xiaoshuai Zhang, Sai Bi, Kalyan Sunkavalli, Hao Su
-
Publication number: 20260080609Abstract: Techniques for volumetric re-lighting of 3D objects are disclosed. In an example method, a computing system receives a first image of a three-dimensional (“3D”) object. The computing system generates a de-lighted image of the 3D object based on the first image. The computing system generates an embedded representation of the 3D object based on the de-lighted image and a first representation of the de-lighted image based on the embedded representation using a first machine learning (“ML”) model. The computing system generates a second representation of the 3D object using a second ML model based on orientation and lighting information and one or more internal states of the first ML model. The computing system generates a third representation of the 3D object by combining the first and second representations. The computing system renders a second image of the 3D object based on the third representation of the 3D object.Type: ApplicationFiled: September 18, 2024Publication date: March 19, 2026Inventors: He Zhang, Zhixin Shu, Yiqun Mei, Xuaner Zhang, Sai Bi, Jianming Zhang
-
Publication number: 20260073580Abstract: The present disclosure relates to systems, methods, and non-transitory computer-readable media that generates an image or a video from a text prompt. For example, the disclosed systems receive a text prompt and generates text tokens from the text prompt. Moreover, the disclosed systems generate combined tokens by combining the text tokens with noised tokens. Further, the disclosed systems generate denoised tokens by removing noise from noised tokens in a manner that incorporates a context indicated by the text tokens and further generates an image or video from the denoised tokens.Type: ApplicationFiled: October 29, 2024Publication date: March 12, 2026Inventors: Kai Zhang, Jianming Zhang, Sai Bi, Zexiang Xu, Hao Tan, Wei-An Lin
-
Publication number: 20260045041Abstract: In implementation of techniques for generating meshes by decoding volume representations, a computing device implements a mesh generation system to receive digital images depicting an object from different angles. The mesh generation system generates a volume representation of the object using a transformer model based on the digital images. By decoding information from the volume representation using an algorithm, the mesh generation system then generates a mesh of the object from the volume representation. The mesh generation system then presents the mesh of the object in a user interface.Type: ApplicationFiled: August 8, 2024Publication date: February 12, 2026Applicants: Adobe Inc., THE REGENTS OF THE UNIVERSITY OF CALIFORNIAInventors: Kai Zhang, Zexiang Xu, Xinyue Wei, Valentin Mathieu Deschaintre, Sai Bi, Kalyan Krishna Sunkavalli, Hao Tan, Fujun Luan, Hao Su
-
Publication number: 20260024278Abstract: A method, apparatus, non-transitory computer readable medium, and system for image generation image generation may include obtaining a first image depicting a first view of an object, generating a second image depicting a second view of the object based on the first image, and generating a third image depicting a third view of the object based on the first image, where the third view is structurally consistent with the second view.Type: ApplicationFiled: July 18, 2024Publication date: January 22, 2026Inventors: Desai Xie, Jiahao Li, Hao Tan, Xin Sun, Zhixin Shu, Yi Zhou, Sai Bi
-
Patent number: 12524954Abstract: Systems and methods for generating a 3D model from a single input image are described. Embodiments are configured to obtain an input image and camera view information corresponding to the input image; encode the input image to obtain 2D features comprising a plurality of 2D tokens corresponding to patches of the input image; decode the 2D features based on the camera view information to obtain 3D features comprising a plurality of 3D tokens corresponding to regions of a 3D representation; and generate a 3D model of the input image based on the 3D features.Type: GrantFiled: September 5, 2023Date of Patent: January 13, 2026Assignee: ADOBE INC.Inventors: Hao Tan, Yicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi, Yang Zhou, Difan Liu, Feng Liu, Kalyan K. Sunkavalli, Trung Huu Bui
-
Publication number: 20260004537Abstract: Certain aspects and features of this disclosure relate to providing a controllable, dynamic appearance for neural 3D portraits. For example, a method involves projecting a color at points in a digital video portrait based on location, surface normal, and viewing direction for each respective point in a canonical space. The method also involves projecting, using the color, dynamic face normals for the points as changing according to an articulated head pose and facial expression in the digital video portrait. The method further involves disentangling, based on the dynamic face normals, a facial appearance in the digital video portrait into intrinsic components in the canonical space. The method additionally involves storing and/or rendering at least a portion of a head pose as a controllable, neural 3D portrait based on the digital video portrait using the intrinsic components.Type: ApplicationFiled: September 4, 2025Publication date: January 1, 2026Inventors: Zhixin Shu, Zexiang Xu, Shahrukh Athar, Sai Bi, Kalyan Sunkavalli, Fujun Luan
-
Publication number: 20250390713Abstract: In some embodiments, a computing system receives an input prompt describing a 3-dimensional (3D) object. The computing system generates one or more levels of latent features based on the input prompt using a latent diffusion model. The computing system decodes the one or more levels of latent features to generate a 3D shape representation using a hierarchical autoencoder. The computing system generates an output shape based on the 3D shape representation.Type: ApplicationFiled: June 25, 2024Publication date: December 25, 2025Inventors: Yang Zhou, Yicong Hong, Sai Bi, Kai Zhang, Hao Tan, Feng Liu, Difan Liu
-
Publication number: 20250336154Abstract: In implementation of techniques for three-dimensional reconstructions based on Gaussian primitives, a computing device implements a reconstruction system to receive a first digital image depicting an object from a first angle and a second digital image depicting the object from a second angle. The reconstruction system segments the first digital image and the second digital image into patches. The reconstruction system then generates, using a machine learning model, three-dimensional Gaussian primitives that predict parameters of points of the object in a three-dimensional space that correspond on a per-pixel basis to pixels of the patches. The reconstruction system then forms a three-dimensional reconstruction of the object for display in a user interface by merging the three-dimensional Gaussian primitives.Type: ApplicationFiled: April 25, 2024Publication date: October 30, 2025Applicant: Adobe Inc.Inventors: Kai Zhang, Hao Tan, Sai Bi, Zexiang Xu, Nanxuan Zhao, Kalyan Krishna Sunkavalli
-
Patent number: 12437492Abstract: Certain aspects and features of this disclosure relate to providing a controllable, dynamic appearance for neural 3D portraits. For example, a method involves projecting a color at points in a digital video portrait based on location, surface normal, and viewing direction for each respective point in a canonical space. The method also involves projecting, using the color, dynamic face normals for the points as changing according to an articulated head pose and facial expression in the digital video portrait. The method further involves disentangling, based on the dynamic face normals, a facial appearance in the digital video portrait into intrinsic components in the canonical space. The method additionally involves storing and/or rendering at least a portion of a head pose as a controllable, neural 3D portrait based on the digital video portrait using the intrinsic components.Type: GrantFiled: April 7, 2023Date of Patent: October 7, 2025Assignee: Adobe Inc.Inventors: Zhixin Shu, Zexiang Xu, Shahrukh Athar, Sai Bi, Kalyan Sunkavalli, Fujun Luan
-
Publication number: 20250139883Abstract: Embodiments are configured to render 3D models using an importance sampling method. First, embodiments obtain a 3D model including a plurality of density values corresponding to a plurality of locations in a 3D space, respectively. Embodiments then sample the color information from within a random subset of the plurality of locations using a probability distribution based on the plurality of density values. Embodiments have a higher probability to sample each location within the random subset of locations if the location has a higher density probability. Embodiments then an image depicting a view of the 3D model based on the sampling within the random subset of the plurality of locations.Type: ApplicationFiled: November 1, 2023Publication date: May 1, 2025Inventors: Milos Hasan, Iliyan Georgiev, Sai Bi, Julien Philip, Kalyan K. Sunkavalli, Xin Sun, Fujun Luan, Kevin James Blackburn-Matzen, Zexiang Xu, Kai Zhang
-
Publication number: 20250104349Abstract: A method, apparatus, non-transitory computer readable medium, and system for 3D model generation include obtaining a plurality of input images depicting an object and a set of 3D position embeddings, where each of the plurality of input images depicts the object from a different perspective, encoding the plurality of input images to obtain a plurality of 2D features corresponding to the plurality of input images, respectively, generating 3D features based on the plurality of 2D features and the set of 3D position embeddings, and generating a 3D model of the object based on the 3D features.Type: ApplicationFiled: September 24, 2024Publication date: March 27, 2025Inventors: Sai Bi, Jiahao Li, Hao Tan, Kai Zhang, Zexiang Xu, Fujun Luan, Yinghao Xu, Yicong Hong, Kalyan K. Sunkavalli
-
Patent number: 12254570Abstract: The present disclosure relates to systems, methods, and non-transitory computer readable media that generate three-dimensional hybrid mesh-volumetric representations for digital objects. For instance, in one or more embodiments, the disclosed systems generate a mesh for a digital object from a plurality of digital images that portray the digital object using a multi-view stereo model. Additionally, the disclosed systems determine a set of sample points for a thin volume around the mesh. Using a neural network, the disclosed systems further generate a three-dimensional hybrid mesh-volumetric representation for the digital object utilizing the set of sample points for the thin volume and the mesh.Type: GrantFiled: May 3, 2022Date of Patent: March 18, 2025Assignee: Adobe Inc.Inventors: Sai Bi, Yang Liu, Zexiang Xu, Fujun Luan, Kalyan Sunkavalli
-
Publication number: 20250078393Abstract: Systems and methods for generating a 3D model from a single input image are described. Embodiments are configured to obtain an input image and camera view information corresponding to the input image; encode the input image to obtain 2D features comprising a plurality of 2D tokens corresponding to patches of the input image; decode the 2D features based on the camera view information to obtain 3D features comprising a plurality of 3D tokens corresponding to regions of a 3D representation; and generate a 3D model of the input image based on the 3D features.Type: ApplicationFiled: September 5, 2023Publication date: March 6, 2025Inventors: HAO TAN, YICONG HONG, KAI ZHANG, JIUXIANG GU, SAI BI, YANG ZHOU, DIFAN LIU, FENG LIU, KALYAN K. SUNKAVALLI, TRUNG HUU BUI
-
Patent number: 12211225Abstract: A scene reconstruction system renders images of a scene with high-quality geometry and appearance and supports view synthesis, relighting, and scene editing. Given a set of input images of a scene, the scene reconstruction system trains a network to learn a volume representation of the scene that includes separate geometry and reflectance parameters. Using the volume representation, the scene reconstruction system can render images of the scene under arbitrary viewing (view synthesis) and lighting (relighting) locations. Additionally, the scene reconstruction system can render images that change the reflectance of objects in the scene (scene editing).Type: GrantFiled: April 15, 2021Date of Patent: January 28, 2025Assignee: ADOBE INC.Inventors: Sai Bi, Zexiang Xu, Kalyan Krishna Sunkavalli, Miloš Hašan, Yannick Hold-Geoffroy, David Jay Kriegman, Ravi Ramamoorthi