Patents by Inventor Yin Cui

Yin Cui has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Publication number: 20260148471
    Abstract: Apparatuses, systems, and techniques to use one or more neural networks to generate texture maps are described. In at least one embodiment, one or more neural networks are used to generate one or more texture maps of a second resolution based, at least in part, on one or more texture maps of a first resolution, less than the second resolution.
    Type: Application
    Filed: November 26, 2024
    Publication date: May 28, 2026
    Inventors: Chen-Hsuan Lin, Tsung-Yi Lin, Zekun Hao, Donglai Xiang, Zhaoshuo Li, Xiaohui Zeng, Jingyi Jin, Qianli Ma, Yen-Chen Lin, Yunhao Ge, Yin Cui, Ming-Yu Liu
  • Publication number: 20260141611
    Abstract: Approaches presented herein may be used to generate three-dimensional (3D) scenes using one or more 3D assets obtained from a layout image. The layout image may be produced from an input prompt, such as a text prompt, and objects at least partially depicted by the layout image may be identified and then used to generate the one or more 3D assets. One or more trained models may be used as part of a pipeline to receive an input, generate the one or more layout images, identify objects at least partially depicted by the one or more layout images, produce one or more 3D assets based on a description of the objects and/or a cropped image of the objects, and then to produce a 3D scene representing the layout image.
    Type: Application
    Filed: November 15, 2024
    Publication date: May 21, 2026
    Inventors: Yifan Ding, Yin Cui, Yunhao Ge, Tsung-Yi Lin, Ming-Yu Liu, Chen-Hsuan Lin, Zekun Hao, Zhaoshuo Li, Xiaohui Zeng, Zeqi Gu
  • Publication number: 20260120487
    Abstract: Apparatuses, systems, and techniques to obtain one or more captions for a video using machine learning. In at least one embodiment, at least one machine learning process is used to generate at least one output caption using at least one image-level caption, at least one video-level caption, and/or at least one motion caption. In at least one embodiment, the video-level caption(s) is/are generated by one or more second machine learning processes using the video, and the image-level caption(s) is/are generated by one or more third machine learning processes using one or more images sampled from the video.
    Type: Application
    Filed: October 11, 2024
    Publication date: April 30, 2026
    Inventors: Boyi Li, Ligeng Zhu, Ran Tian, Shuhan Tan, Yao Lu, Yin Cui, Yuxiao Chen, Xinshuo Weng, Sushant Veer, Jonah Philion, Max Ehrlich, Andrew Tao, Sanja Fidler, Ming-Yu Liu, Boris Ivanovic, Song Han, Marco Pavone
  • Publication number: 20260101026
    Abstract: Embodiments of the present disclosure provide systems and methods for generating a three-dimensional (3D) asset. A first multi-view diffusion model generates a multi-view image based on an input prompt. A second multi-view diffusion model generates a multi-view surface normal image based on the input prompt and the multi-view image. A 3D reconstruction engine processes the multi-view image and the multi-view surface normal image to generate an intermediate 3D representation of the object, including a polygon mesh and a low-resolution texture map. A rendering engine renders a second multi-view image based on the intermediate 3D representation. A third multi-view diffusion model processes the input prompt and the second multi-view image to generate a high-resolution multi-view image. The 3D reconstruction engine upscales the low-resolution texture map based on the high-resolution multi-view image and generates the 3D asset, which includes the polygon mesh and the upscaled texture map.
    Type: Application
    Filed: April 4, 2025
    Publication date: April 9, 2026
    Inventors: Chen-Hsuan Lin, Tsung-Yi Lin, Ming-Yu Liu, Xiaohui Zeng, Zhaoshuo Li, Zekun Hao, Donglai Xiang, Qianli Ma, Jingyi Jin, Fangyin Wei, Yin Cui
  • Publication number: 20260094239
    Abstract: The disclosed method for generating images includes performing, based on one or more inputs, one or more first denoising diffusion operations using a first trained machine learning model to generate a first image at a first resolution; and performing, based on the one or more inputs and the first image, one or more second denoising diffusion operations using a second trained machine learning model to generate a second image at a second resolution.
    Type: Application
    Filed: August 19, 2025
    Publication date: April 2, 2026
    Inventors: Yogesh BALAJI, Ting-Chun WANG, Jiaojiao FAN, Qinsheng ZHANG, Xiaohui ZENG, Maciej BALA, Yin CUI, Yuval ATZMON, Aaron LICATA, Pooya JANNATY, Siddharth GURURANI, Seungjun NAH, Yu ZENG, John LEWIS, Jacob Samuel HUFFMAN, Yunhao GE, Fitsum REDA, Ming-Yu LIU
  • Publication number: 20260094245
    Abstract: The disclosed method for generating images includes performing, based on one or more inputs, one or more first denoising diffusion operations using a first trained machine learning model to generate a first image at a first resolution; and performing, based on the one or more inputs and the first image, one or more second denoising diffusion operations using a second trained machine learning model to generate a second image at a second resolution.
    Type: Application
    Filed: August 19, 2025
    Publication date: April 2, 2026
    Inventors: Yogesh BALAJI, Ting-Chun WANG, Jiaojiao FAN, Qinsheng ZHANG, Xiaohui ZENG, Maciej BALA, Yin CUI, Yuval ATZMON, Aaron LICATA, Pooya JANNATY, Siddharth GURURANI, Seungjun NAH, Yu ZENG, John LEWIS, Jacob Samuel HUFFMAN, Yunhao GE, Fitsum REDA, Ming-Yu LIU
  • Publication number: 20260038190
    Abstract: In various examples, techniques for performing data filtering for AI systems and applications is described herein. Systems and methods described herein may use a pipeline that is configured to filter shapes, such as three-dimensional shapes, in order to identify high-quality shapes for a final dataset. In some examples, to identify the high-quality shapes, training shapes along with ground truth scores associated with the training shapes may be used to train a machine learning model. During and/or after the training, the machine learning model may then be used to process data associated with additional shapes in order to determine quality scores associated with the additional shapes. These quality scores may then be used to select the high-quality shapes, such as shapes that satisfy a threshold quality score. In some examples, additional filtering may be performed by the pipeline, such as by using one or more rules for removing low-quality shapes.
    Type: Application
    Filed: July 31, 2024
    Publication date: February 5, 2026
    Inventors: Zekun Hao, Yin Cui, Fangyin Wei, Yunhao Ge, Yifan Ding, Xiaohui Zeng, Tsung-Yi Lin, Chen-Hsuan Lin, Ming-Yu Liu
  • Publication number: 20260038191
    Abstract: In various examples, techniques for automatic annotation of shapes for AI systems and applications is described herein. Systems and methods described herein may use a pipeline that is configured to generate annotations for shapes, such as three-dimensional shapes, using various types of captions. For instance, image data representing images of the shapes, data representing description of the shapes, and/or data representing a format for the annotations may be input into one or more multimodal language models. The multimodal language model(s) may then be configured to process the data and, based at least on the processing, generate short captions and long captions associated with the shapes. These captions may then be stored in association with the shapes and/or the images. In some examples, embeddings may initially be generated for the captions, where the embeddings are then stored in association with the shapes and/or the images.
    Type: Application
    Filed: July 31, 2024
    Publication date: February 5, 2026
    Inventors: Zekun Hao, Yin Cui, Fangyin Wei, Yunhao Ge, Yifan Ding, Xiaohui Zeng, Tsung-Yi Lin, Chen-Hsuan Lin, Ming-Yu Liu
  • Publication number: 20260038213
    Abstract: In various examples, techniques for aligning shapes using pose information for AI systems and applications is described herein. Systems and methods described herein may use one or more pipelines that are configured to align shapes, such as three-dimensional shapes, using poses of the shapes. In some examples, to identify the poses, training images along with ground truth poses associated with the training images may be used to train a machine learning model. During and/or after the training, the machine learning model may then be used to process additional images in order to identify poses of additional shapes. In some examples, a pose associated with a shape may include a gravity orientation and/or an azimuth orientation of the shape as represented by an image. These poses may then be used to align the additional shapes, such as with respect to a canonical pose.
    Type: Application
    Filed: July 31, 2024
    Publication date: February 5, 2026
    Inventors: Zekun Hao, Yin Cui, Fangyin Wei, Yunhao Ge, Yifan Ding, Xiaohui Zeng, Tsung-Yi Lin, Chen-Hsuan Lin, Ming-Yu Liu
  • Publication number: 20250384650
    Abstract: An example method of training a detector head for object detection of a training object category based on a frozen vision and language model (VLM) is provided. The method includes receiving the frozen VLM pre-trained on a plurality of image-text pairs. The method includes determining, for an image embedding generated by a pre-trained image encoder of the frozen VLM and by the detector head, a detection region embedding indicative of one or more regions of interest in an image. The method includes generating, by a pre-trained text encoder of the frozen VLM, a text embedding of the training object category. The method includes predicting, by the detector head and based on the detection region embedding and the text embedding of the training object category, an object from a target object vocabulary associated with the training object category. The method includes providing the pre-trained frozen VLM and the trained detector head.
    Type: Application
    Filed: June 28, 2023
    Publication date: December 18, 2025
    Inventors: Wei-Cheng Kuo, Yin Cui, Xiuye Gu, Anthony Jacob Piergiovanni, Anelia Angelova
  • Publication number: 20250378603
    Abstract: Apparatuses, systems, and techniques to identify a location in which to place objects within a graphically rendered scene. In at least one embodiment, a location in which to place objects is identified using one or more neural networks, based, at least in part, on text or speech input to the one or more neural networks.
    Type: Application
    Filed: June 7, 2024
    Publication date: December 11, 2025
    Inventors: Yunhao Ge, Yifan Ding, Yin Cui, Tsung-Yi Lin, Chen-Hsuan Lin, Zekun Hao, Zhaoshuo Li, Xiaohui Zeng, Hanzi Mao, Ming-Yu Liu, John Peter Lewis, Donglai Xiang, Qianli Ma, Jiashu Xu Xu
  • Publication number: 20240378509
    Abstract: A computer-implemented method of generating scale-permuted models can generate models having improved accuracy and reduced evaluation computational requirements. The method can include defining, by a computing system including one or more computing devices, a search space including a plurality of candidate permutations of a plurality of candidate feature blocks, each of the plurality of candidate feature blocks having a respective scale. The method can include performing, by the computing system, a plurality of search iterations by a search algorithm to select a scale-permuted model from the search space, the scale-permuted model based at least in part on a candidate permutation of the plurality of candidate permutations.
    Type: Application
    Filed: July 25, 2024
    Publication date: November 14, 2024
    Inventors: Xianzhi Du, Yin Cui, Tsung-Yi Lin, Quoc V. Le, Pengchong Jin, Mingxing Tan, Golnaz Ghiasi, Xiaodan Song
  • Publication number: 20240362460
    Abstract: The technology relates to providing personalized neural network-based models according to user input, which can be generated upon request or otherwise as needed. This may include receiving, by one or more processors of a computing device, input corresponding to a task description. Then the input corresponding to the task description is encoded into a set of text embeddings. Based on this, the system applies mixer prediction to the set of text embeddings to generate a set of mixers and learns a set of basis models according to the set of mixers. The set of basis models are combined to form a single personalized model corresponding to the task description. This personalized model can then be used in video understanding, quality assessment, providing a recommendation, performing a classification, or performing a search.
    Type: Application
    Filed: April 4, 2024
    Publication date: October 31, 2024
    Inventors: Li Zhang, Yandong Li, Yin Cui, Hong-You Chen, Mingda Zhang
  • Patent number: 12079695
    Abstract: A computer-implemented method of generating scale-permuted models can generate models having improved accuracy and reduced evaluation computational requirements. The method can include defining, by a computing system including one or more computing devices, a search space including a plurality of candidate permutations of a plurality of candidate feature blocks, each of the plurality of candidate feature blocks having a respective scale. The method can include performing, by the computing system, a plurality of search iterations by a search algorithm to select a scale-permuted model from the search space, the scale-permuted model based at least in part on a candidate permutation of the plurality of candidate permutations.
    Type: Grant
    Filed: October 1, 2020
    Date of Patent: September 3, 2024
    Assignee: GOOGLE LLC
    Inventors: Xianzhi Du, Yin Cui, Tsung-Yi Lin, Quoc V. Le, Pengchong Jin, Mingxing Tan, Golnaz Ghiasi, Xiaodan Song
  • Publication number: 20240282131
    Abstract: Systems and methods for zero-shot prompt ensembling for zero-shot classification with text-image models can include utilizing a pre-trained text-image model to perform downstream tasks based on prompt-based weighting. The systems and methods may adjust for frequency-based bias and may automatically determine different prompt associations with a given downstream task. The systems and methods can aggregate weighted text embeddings and then determine a classification output based on similarity measures between an image embedding and the aggregated weighted text embeddings.
    Type: Application
    Filed: January 24, 2024
    Publication date: August 22, 2024
    Inventors: Jie Ren, Zhe Liu, James Urquhart Allingham, Michael Ward Dusenberry, Dustin Tran, Yin Cui, Balaji Lakshminarayanan, Xiuye Gu
  • Publication number: 20220108204
    Abstract: A computer-implemented method of generating scale-permuted models can generate models having improved accuracy and reduced evaluation computational requirements. The method can include defining, by a computing system including one or more computing devices, a search space including a plurality of candidate permutations of a plurality of candidate feature blocks, each of the plurality of candidate feature blocks having a respective scale. The method can include performing, by the computing system, a plurality of search iterations by a search algorithm to select a scale-permuted model from the search space, the scale-permuted model based at least in part on a candidate permutation of the plurality of candidate permutations.
    Type: Application
    Filed: October 1, 2020
    Publication date: April 7, 2022
    Inventors: Xianzhi Du, Yin Cui, Tsung-Yi Lin, Quoc V. Le, Pengchong Jin, Mingxing Tan, Golnaz Ghiasi, Xiaodan Song