Patents by Inventor Sai Bi

Sai Bi has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Patent number: 12682426
    Abstract: In some examples, a computing system accesses a field of view (FOV) image that has a field of view less than 360 degrees and has low dynamic range (LDR) values. The computing system estimates lighting parameters from a scene depicted in the FOV image and generates a lighting image based on the lighting parameters. The computing system further generates lighting features generated the lighting image and image features generated from the FOV image. These features are aggregated into aggregated features and a machine learning model is applied to the image features and the aggregated features to generate a panorama image having high dynamic range (HDR) values.
    Type: Grant
    Filed: August 25, 2023
    Date of Patent: July 14, 2026
    Assignee: Adobe Inc.
    Inventors: Mohammad Reza Karimi Dastjerdi, Yannick Hold-Geoffroy, Sai Bi, Jonathan Eisenmann, Jean-François Lalonde
  • Patent number: 12682547
    Abstract: Embodiments are configured to render 3D models using an importance sampling method. First, embodiments obtain a 3D model including a plurality of density values corresponding to a plurality of locations in a 3D space, respectively. Embodiments then sample the color information from within a random subset of the plurality of locations using a probability distribution based on the plurality of density values. Embodiments have a higher probability to sample each location within the random subset of locations if the location has a higher density probability. Embodiments then an image depicting a view of the 3D model based on the sampling within the random subset of the plurality of locations.
    Type: Grant
    Filed: November 1, 2023
    Date of Patent: July 14, 2026
    Assignee: ADOBE INC.
    Inventors: Milos Hasan, Iliyan Georgiev, Sai Bi, Julien Philip, Kalyan K. Sunkavalli, Xin Sun, Fujun Luan, Kevin James Blackburn-Matzen, Zexiang Xu, Kai Zhang
  • Publication number: 20260148487
    Abstract: In some embodiments, a computing system accesses multiple input images of a specular object with a scene. The computing system encodes near-field interreflections of the scene on the specular object to obtain a first set of feature representations of the specular object in multiple viewing directions based on the multiple input images. The computing system encodes far-field reflections of the scene on the specular object to obtain a second set of feature representations in the multiple viewing directions based on the multiple input images. The computing system determines a set of specular color values for the specular object in the multiple viewing directions based on the first set of feature representations and the second set of feature representations using a multi-layer perceptron algorithm. The computing system renders the specular object representation at least based on the set of specular color values using a neural rendering algorithm.
    Type: Application
    Filed: November 25, 2024
    Publication date: May 28, 2026
    Inventors: Sai Bi, Zexiang Xu, Liwen Wu, Kalyan Sunkavalli, Kai Zhang, Iliyan Georgiev, Fujun Luan, Ravi Ramamoorthi
  • Publication number: 20260100010
    Abstract: Techniques for depth guided text-based editing of 3D neural radiance fields are provided. A method includes receiving input 2D images corresponding to views of a target and generating a 3D representation from the input 2D images. The 3D representation includes points forming a point cloud, where each point has a color and density value. The method also includes accumulating the color and density values to generate a volumetric 3D scene having a geometry, extracting distance maps from the volumetric 3D scene based on the geometry, and generating a plurality of masks associated with the target for each view. The method also includes aggregating the masks into the volumetric 3D scene using the geometry, providing the input 2D images, the masks, and the distance maps to a diffusion model, and modifying an appearance of the target in the volumetric 3D scene by providing a text command to the diffusion model.
    Type: Application
    Filed: October 7, 2024
    Publication date: April 9, 2026
    Inventors: Sara Rojas Martinez, Julien Philip, Kai Zhang, Sai Bi, Fujun Luan, Kalyan Sunkavalli
  • Publication number: 20260087600
    Abstract: In implementing per-asset denoising for real-time rendering of neural radiance fields (NeRFs), a processing device receives a three-dimensional (3D) representation of a scene as a NeRF. The processing device generates an intermediate rendering of the scene using the NeRF. The intermediate rendering is denoised using a machine-learning model to generate a final rendering. The machine-learning model is trained on another rendering of this scene, which was rendered using a non-real-time, high-quality rendering scheme. In other words, the machine-learning model is optimized for each scene and provides a lightweight denoising network to provide real-time NeRF rendering while maintaining the high-quality visuals of non-real-time rendering schemes. The final rendering is then presented via a display device.
    Type: Application
    Filed: September 24, 2024
    Publication date: March 26, 2026
    Applicants: Adobe Inc., THE REGENTS OF THE UNIVERSITY OF CALIFORNIA
    Inventors: Sai Bi, Zexiang Xu, Xin Sun, Miloš Hašan, Kunal Gupta, Kevin Blackburn-Matzen, Kalyan Krishna Sunkavalli, Kai Zhang, Julien Olivier Victor Philip, Fujun Luan, Manmohan Chandraker, Iliyan Atanasov Georgiev
  • Patent number: 12586291
    Abstract: Embodiments are disclosed for fast large-scale radiance field reconstruction. A method of fast large-scale radiance field reconstruction may include receiving a sequence of input images that depict views of a scene and extracting, using an image encoder, image features from the sequence of input images. A first one or more machine learning models may generate a local volume based on the image features corresponding to one or more images from the sequence of input images. A second one or more machine learning models may generate a global volume based on the local volume. A novel view of the scene is synthesized based on the global volume.
    Type: Grant
    Filed: March 16, 2023
    Date of Patent: March 24, 2026
    Assignees: Adobe Inc., The Regents of the University of Calfornia
    Inventors: Zexiang Xu, Xiaoshuai Zhang, Sai Bi, Kalyan Sunkavalli, Hao Su
  • Publication number: 20260080609
    Abstract: Techniques for volumetric re-lighting of 3D objects are disclosed. In an example method, a computing system receives a first image of a three-dimensional (“3D”) object. The computing system generates a de-lighted image of the 3D object based on the first image. The computing system generates an embedded representation of the 3D object based on the de-lighted image and a first representation of the de-lighted image based on the embedded representation using a first machine learning (“ML”) model. The computing system generates a second representation of the 3D object using a second ML model based on orientation and lighting information and one or more internal states of the first ML model. The computing system generates a third representation of the 3D object by combining the first and second representations. The computing system renders a second image of the 3D object based on the third representation of the 3D object.
    Type: Application
    Filed: September 18, 2024
    Publication date: March 19, 2026
    Inventors: He Zhang, Zhixin Shu, Yiqun Mei, Xuaner Zhang, Sai Bi, Jianming Zhang
  • Publication number: 20260073580
    Abstract: The present disclosure relates to systems, methods, and non-transitory computer-readable media that generates an image or a video from a text prompt. For example, the disclosed systems receive a text prompt and generates text tokens from the text prompt. Moreover, the disclosed systems generate combined tokens by combining the text tokens with noised tokens. Further, the disclosed systems generate denoised tokens by removing noise from noised tokens in a manner that incorporates a context indicated by the text tokens and further generates an image or video from the denoised tokens.
    Type: Application
    Filed: October 29, 2024
    Publication date: March 12, 2026
    Inventors: Kai Zhang, Jianming Zhang, Sai Bi, Zexiang Xu, Hao Tan, Wei-An Lin
  • Publication number: 20260045041
    Abstract: In implementation of techniques for generating meshes by decoding volume representations, a computing device implements a mesh generation system to receive digital images depicting an object from different angles. The mesh generation system generates a volume representation of the object using a transformer model based on the digital images. By decoding information from the volume representation using an algorithm, the mesh generation system then generates a mesh of the object from the volume representation. The mesh generation system then presents the mesh of the object in a user interface.
    Type: Application
    Filed: August 8, 2024
    Publication date: February 12, 2026
    Applicants: Adobe Inc., THE REGENTS OF THE UNIVERSITY OF CALIFORNIA
    Inventors: Kai Zhang, Zexiang Xu, Xinyue Wei, Valentin Mathieu Deschaintre, Sai Bi, Kalyan Krishna Sunkavalli, Hao Tan, Fujun Luan, Hao Su
  • Publication number: 20260024278
    Abstract: A method, apparatus, non-transitory computer readable medium, and system for image generation image generation may include obtaining a first image depicting a first view of an object, generating a second image depicting a second view of the object based on the first image, and generating a third image depicting a third view of the object based on the first image, where the third view is structurally consistent with the second view.
    Type: Application
    Filed: July 18, 2024
    Publication date: January 22, 2026
    Inventors: Desai Xie, Jiahao Li, Hao Tan, Xin Sun, Zhixin Shu, Yi Zhou, Sai Bi
  • Patent number: 12524954
    Abstract: Systems and methods for generating a 3D model from a single input image are described. Embodiments are configured to obtain an input image and camera view information corresponding to the input image; encode the input image to obtain 2D features comprising a plurality of 2D tokens corresponding to patches of the input image; decode the 2D features based on the camera view information to obtain 3D features comprising a plurality of 3D tokens corresponding to regions of a 3D representation; and generate a 3D model of the input image based on the 3D features.
    Type: Grant
    Filed: September 5, 2023
    Date of Patent: January 13, 2026
    Assignee: ADOBE INC.
    Inventors: Hao Tan, Yicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi, Yang Zhou, Difan Liu, Feng Liu, Kalyan K. Sunkavalli, Trung Huu Bui
  • Publication number: 20260004537
    Abstract: Certain aspects and features of this disclosure relate to providing a controllable, dynamic appearance for neural 3D portraits. For example, a method involves projecting a color at points in a digital video portrait based on location, surface normal, and viewing direction for each respective point in a canonical space. The method also involves projecting, using the color, dynamic face normals for the points as changing according to an articulated head pose and facial expression in the digital video portrait. The method further involves disentangling, based on the dynamic face normals, a facial appearance in the digital video portrait into intrinsic components in the canonical space. The method additionally involves storing and/or rendering at least a portion of a head pose as a controllable, neural 3D portrait based on the digital video portrait using the intrinsic components.
    Type: Application
    Filed: September 4, 2025
    Publication date: January 1, 2026
    Inventors: Zhixin Shu, Zexiang Xu, Shahrukh Athar, Sai Bi, Kalyan Sunkavalli, Fujun Luan
  • Publication number: 20250390713
    Abstract: In some embodiments, a computing system receives an input prompt describing a 3-dimensional (3D) object. The computing system generates one or more levels of latent features based on the input prompt using a latent diffusion model. The computing system decodes the one or more levels of latent features to generate a 3D shape representation using a hierarchical autoencoder. The computing system generates an output shape based on the 3D shape representation.
    Type: Application
    Filed: June 25, 2024
    Publication date: December 25, 2025
    Inventors: Yang Zhou, Yicong Hong, Sai Bi, Kai Zhang, Hao Tan, Feng Liu, Difan Liu
  • Publication number: 20250336154
    Abstract: In implementation of techniques for three-dimensional reconstructions based on Gaussian primitives, a computing device implements a reconstruction system to receive a first digital image depicting an object from a first angle and a second digital image depicting the object from a second angle. The reconstruction system segments the first digital image and the second digital image into patches. The reconstruction system then generates, using a machine learning model, three-dimensional Gaussian primitives that predict parameters of points of the object in a three-dimensional space that correspond on a per-pixel basis to pixels of the patches. The reconstruction system then forms a three-dimensional reconstruction of the object for display in a user interface by merging the three-dimensional Gaussian primitives.
    Type: Application
    Filed: April 25, 2024
    Publication date: October 30, 2025
    Applicant: Adobe Inc.
    Inventors: Kai Zhang, Hao Tan, Sai Bi, Zexiang Xu, Nanxuan Zhao, Kalyan Krishna Sunkavalli
  • Patent number: 12437492
    Abstract: Certain aspects and features of this disclosure relate to providing a controllable, dynamic appearance for neural 3D portraits. For example, a method involves projecting a color at points in a digital video portrait based on location, surface normal, and viewing direction for each respective point in a canonical space. The method also involves projecting, using the color, dynamic face normals for the points as changing according to an articulated head pose and facial expression in the digital video portrait. The method further involves disentangling, based on the dynamic face normals, a facial appearance in the digital video portrait into intrinsic components in the canonical space. The method additionally involves storing and/or rendering at least a portion of a head pose as a controllable, neural 3D portrait based on the digital video portrait using the intrinsic components.
    Type: Grant
    Filed: April 7, 2023
    Date of Patent: October 7, 2025
    Assignee: Adobe Inc.
    Inventors: Zhixin Shu, Zexiang Xu, Shahrukh Athar, Sai Bi, Kalyan Sunkavalli, Fujun Luan
  • Publication number: 20250139883
    Abstract: Embodiments are configured to render 3D models using an importance sampling method. First, embodiments obtain a 3D model including a plurality of density values corresponding to a plurality of locations in a 3D space, respectively. Embodiments then sample the color information from within a random subset of the plurality of locations using a probability distribution based on the plurality of density values. Embodiments have a higher probability to sample each location within the random subset of locations if the location has a higher density probability. Embodiments then an image depicting a view of the 3D model based on the sampling within the random subset of the plurality of locations.
    Type: Application
    Filed: November 1, 2023
    Publication date: May 1, 2025
    Inventors: Milos Hasan, Iliyan Georgiev, Sai Bi, Julien Philip, Kalyan K. Sunkavalli, Xin Sun, Fujun Luan, Kevin James Blackburn-Matzen, Zexiang Xu, Kai Zhang
  • Publication number: 20250104349
    Abstract: A method, apparatus, non-transitory computer readable medium, and system for 3D model generation include obtaining a plurality of input images depicting an object and a set of 3D position embeddings, where each of the plurality of input images depicts the object from a different perspective, encoding the plurality of input images to obtain a plurality of 2D features corresponding to the plurality of input images, respectively, generating 3D features based on the plurality of 2D features and the set of 3D position embeddings, and generating a 3D model of the object based on the 3D features.
    Type: Application
    Filed: September 24, 2024
    Publication date: March 27, 2025
    Inventors: Sai Bi, Jiahao Li, Hao Tan, Kai Zhang, Zexiang Xu, Fujun Luan, Yinghao Xu, Yicong Hong, Kalyan K. Sunkavalli
  • Patent number: 12254570
    Abstract: The present disclosure relates to systems, methods, and non-transitory computer readable media that generate three-dimensional hybrid mesh-volumetric representations for digital objects. For instance, in one or more embodiments, the disclosed systems generate a mesh for a digital object from a plurality of digital images that portray the digital object using a multi-view stereo model. Additionally, the disclosed systems determine a set of sample points for a thin volume around the mesh. Using a neural network, the disclosed systems further generate a three-dimensional hybrid mesh-volumetric representation for the digital object utilizing the set of sample points for the thin volume and the mesh.
    Type: Grant
    Filed: May 3, 2022
    Date of Patent: March 18, 2025
    Assignee: Adobe Inc.
    Inventors: Sai Bi, Yang Liu, Zexiang Xu, Fujun Luan, Kalyan Sunkavalli
  • Publication number: 20250078393
    Abstract: Systems and methods for generating a 3D model from a single input image are described. Embodiments are configured to obtain an input image and camera view information corresponding to the input image; encode the input image to obtain 2D features comprising a plurality of 2D tokens corresponding to patches of the input image; decode the 2D features based on the camera view information to obtain 3D features comprising a plurality of 3D tokens corresponding to regions of a 3D representation; and generate a 3D model of the input image based on the 3D features.
    Type: Application
    Filed: September 5, 2023
    Publication date: March 6, 2025
    Inventors: HAO TAN, YICONG HONG, KAI ZHANG, JIUXIANG GU, SAI BI, YANG ZHOU, DIFAN LIU, FENG LIU, KALYAN K. SUNKAVALLI, TRUNG HUU BUI
  • Patent number: 12211225
    Abstract: A scene reconstruction system renders images of a scene with high-quality geometry and appearance and supports view synthesis, relighting, and scene editing. Given a set of input images of a scene, the scene reconstruction system trains a network to learn a volume representation of the scene that includes separate geometry and reflectance parameters. Using the volume representation, the scene reconstruction system can render images of the scene under arbitrary viewing (view synthesis) and lighting (relighting) locations. Additionally, the scene reconstruction system can render images that change the reflectance of objects in the scene (scene editing).
    Type: Grant
    Filed: April 15, 2021
    Date of Patent: January 28, 2025
    Assignee: ADOBE INC.
    Inventors: Sai Bi, Zexiang Xu, Kalyan Krishna Sunkavalli, Miloš Hašan, Yannick Hold-Geoffroy, David Jay Kriegman, Ravi Ramamoorthi