Patents by Inventor Yin Cui
Yin Cui has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Publication number: 20260148471Abstract: Apparatuses, systems, and techniques to use one or more neural networks to generate texture maps are described. In at least one embodiment, one or more neural networks are used to generate one or more texture maps of a second resolution based, at least in part, on one or more texture maps of a first resolution, less than the second resolution.Type: ApplicationFiled: November 26, 2024Publication date: May 28, 2026Inventors: Chen-Hsuan Lin, Tsung-Yi Lin, Zekun Hao, Donglai Xiang, Zhaoshuo Li, Xiaohui Zeng, Jingyi Jin, Qianli Ma, Yen-Chen Lin, Yunhao Ge, Yin Cui, Ming-Yu Liu
-
Publication number: 20260141611Abstract: Approaches presented herein may be used to generate three-dimensional (3D) scenes using one or more 3D assets obtained from a layout image. The layout image may be produced from an input prompt, such as a text prompt, and objects at least partially depicted by the layout image may be identified and then used to generate the one or more 3D assets. One or more trained models may be used as part of a pipeline to receive an input, generate the one or more layout images, identify objects at least partially depicted by the one or more layout images, produce one or more 3D assets based on a description of the objects and/or a cropped image of the objects, and then to produce a 3D scene representing the layout image.Type: ApplicationFiled: November 15, 2024Publication date: May 21, 2026Inventors: Yifan Ding, Yin Cui, Yunhao Ge, Tsung-Yi Lin, Ming-Yu Liu, Chen-Hsuan Lin, Zekun Hao, Zhaoshuo Li, Xiaohui Zeng, Zeqi Gu
-
Publication number: 20260120487Abstract: Apparatuses, systems, and techniques to obtain one or more captions for a video using machine learning. In at least one embodiment, at least one machine learning process is used to generate at least one output caption using at least one image-level caption, at least one video-level caption, and/or at least one motion caption. In at least one embodiment, the video-level caption(s) is/are generated by one or more second machine learning processes using the video, and the image-level caption(s) is/are generated by one or more third machine learning processes using one or more images sampled from the video.Type: ApplicationFiled: October 11, 2024Publication date: April 30, 2026Inventors: Boyi Li, Ligeng Zhu, Ran Tian, Shuhan Tan, Yao Lu, Yin Cui, Yuxiao Chen, Xinshuo Weng, Sushant Veer, Jonah Philion, Max Ehrlich, Andrew Tao, Sanja Fidler, Ming-Yu Liu, Boris Ivanovic, Song Han, Marco Pavone
-
Publication number: 20260101026Abstract: Embodiments of the present disclosure provide systems and methods for generating a three-dimensional (3D) asset. A first multi-view diffusion model generates a multi-view image based on an input prompt. A second multi-view diffusion model generates a multi-view surface normal image based on the input prompt and the multi-view image. A 3D reconstruction engine processes the multi-view image and the multi-view surface normal image to generate an intermediate 3D representation of the object, including a polygon mesh and a low-resolution texture map. A rendering engine renders a second multi-view image based on the intermediate 3D representation. A third multi-view diffusion model processes the input prompt and the second multi-view image to generate a high-resolution multi-view image. The 3D reconstruction engine upscales the low-resolution texture map based on the high-resolution multi-view image and generates the 3D asset, which includes the polygon mesh and the upscaled texture map.Type: ApplicationFiled: April 4, 2025Publication date: April 9, 2026Inventors: Chen-Hsuan Lin, Tsung-Yi Lin, Ming-Yu Liu, Xiaohui Zeng, Zhaoshuo Li, Zekun Hao, Donglai Xiang, Qianli Ma, Jingyi Jin, Fangyin Wei, Yin Cui
-
Publication number: 20260094239Abstract: The disclosed method for generating images includes performing, based on one or more inputs, one or more first denoising diffusion operations using a first trained machine learning model to generate a first image at a first resolution; and performing, based on the one or more inputs and the first image, one or more second denoising diffusion operations using a second trained machine learning model to generate a second image at a second resolution.Type: ApplicationFiled: August 19, 2025Publication date: April 2, 2026Inventors: Yogesh BALAJI, Ting-Chun WANG, Jiaojiao FAN, Qinsheng ZHANG, Xiaohui ZENG, Maciej BALA, Yin CUI, Yuval ATZMON, Aaron LICATA, Pooya JANNATY, Siddharth GURURANI, Seungjun NAH, Yu ZENG, John LEWIS, Jacob Samuel HUFFMAN, Yunhao GE, Fitsum REDA, Ming-Yu LIU
-
Publication number: 20260094245Abstract: The disclosed method for generating images includes performing, based on one or more inputs, one or more first denoising diffusion operations using a first trained machine learning model to generate a first image at a first resolution; and performing, based on the one or more inputs and the first image, one or more second denoising diffusion operations using a second trained machine learning model to generate a second image at a second resolution.Type: ApplicationFiled: August 19, 2025Publication date: April 2, 2026Inventors: Yogesh BALAJI, Ting-Chun WANG, Jiaojiao FAN, Qinsheng ZHANG, Xiaohui ZENG, Maciej BALA, Yin CUI, Yuval ATZMON, Aaron LICATA, Pooya JANNATY, Siddharth GURURANI, Seungjun NAH, Yu ZENG, John LEWIS, Jacob Samuel HUFFMAN, Yunhao GE, Fitsum REDA, Ming-Yu LIU
-
Publication number: 20260038190Abstract: In various examples, techniques for performing data filtering for AI systems and applications is described herein. Systems and methods described herein may use a pipeline that is configured to filter shapes, such as three-dimensional shapes, in order to identify high-quality shapes for a final dataset. In some examples, to identify the high-quality shapes, training shapes along with ground truth scores associated with the training shapes may be used to train a machine learning model. During and/or after the training, the machine learning model may then be used to process data associated with additional shapes in order to determine quality scores associated with the additional shapes. These quality scores may then be used to select the high-quality shapes, such as shapes that satisfy a threshold quality score. In some examples, additional filtering may be performed by the pipeline, such as by using one or more rules for removing low-quality shapes.Type: ApplicationFiled: July 31, 2024Publication date: February 5, 2026Inventors: Zekun Hao, Yin Cui, Fangyin Wei, Yunhao Ge, Yifan Ding, Xiaohui Zeng, Tsung-Yi Lin, Chen-Hsuan Lin, Ming-Yu Liu
-
Publication number: 20260038191Abstract: In various examples, techniques for automatic annotation of shapes for AI systems and applications is described herein. Systems and methods described herein may use a pipeline that is configured to generate annotations for shapes, such as three-dimensional shapes, using various types of captions. For instance, image data representing images of the shapes, data representing description of the shapes, and/or data representing a format for the annotations may be input into one or more multimodal language models. The multimodal language model(s) may then be configured to process the data and, based at least on the processing, generate short captions and long captions associated with the shapes. These captions may then be stored in association with the shapes and/or the images. In some examples, embeddings may initially be generated for the captions, where the embeddings are then stored in association with the shapes and/or the images.Type: ApplicationFiled: July 31, 2024Publication date: February 5, 2026Inventors: Zekun Hao, Yin Cui, Fangyin Wei, Yunhao Ge, Yifan Ding, Xiaohui Zeng, Tsung-Yi Lin, Chen-Hsuan Lin, Ming-Yu Liu
-
Publication number: 20260038213Abstract: In various examples, techniques for aligning shapes using pose information for AI systems and applications is described herein. Systems and methods described herein may use one or more pipelines that are configured to align shapes, such as three-dimensional shapes, using poses of the shapes. In some examples, to identify the poses, training images along with ground truth poses associated with the training images may be used to train a machine learning model. During and/or after the training, the machine learning model may then be used to process additional images in order to identify poses of additional shapes. In some examples, a pose associated with a shape may include a gravity orientation and/or an azimuth orientation of the shape as represented by an image. These poses may then be used to align the additional shapes, such as with respect to a canonical pose.Type: ApplicationFiled: July 31, 2024Publication date: February 5, 2026Inventors: Zekun Hao, Yin Cui, Fangyin Wei, Yunhao Ge, Yifan Ding, Xiaohui Zeng, Tsung-Yi Lin, Chen-Hsuan Lin, Ming-Yu Liu
-
Publication number: 20250384650Abstract: An example method of training a detector head for object detection of a training object category based on a frozen vision and language model (VLM) is provided. The method includes receiving the frozen VLM pre-trained on a plurality of image-text pairs. The method includes determining, for an image embedding generated by a pre-trained image encoder of the frozen VLM and by the detector head, a detection region embedding indicative of one or more regions of interest in an image. The method includes generating, by a pre-trained text encoder of the frozen VLM, a text embedding of the training object category. The method includes predicting, by the detector head and based on the detection region embedding and the text embedding of the training object category, an object from a target object vocabulary associated with the training object category. The method includes providing the pre-trained frozen VLM and the trained detector head.Type: ApplicationFiled: June 28, 2023Publication date: December 18, 2025Inventors: Wei-Cheng Kuo, Yin Cui, Xiuye Gu, Anthony Jacob Piergiovanni, Anelia Angelova
-
Publication number: 20250378603Abstract: Apparatuses, systems, and techniques to identify a location in which to place objects within a graphically rendered scene. In at least one embodiment, a location in which to place objects is identified using one or more neural networks, based, at least in part, on text or speech input to the one or more neural networks.Type: ApplicationFiled: June 7, 2024Publication date: December 11, 2025Inventors: Yunhao Ge, Yifan Ding, Yin Cui, Tsung-Yi Lin, Chen-Hsuan Lin, Zekun Hao, Zhaoshuo Li, Xiaohui Zeng, Hanzi Mao, Ming-Yu Liu, John Peter Lewis, Donglai Xiang, Qianli Ma, Jiashu Xu Xu
-
Publication number: 20240378509Abstract: A computer-implemented method of generating scale-permuted models can generate models having improved accuracy and reduced evaluation computational requirements. The method can include defining, by a computing system including one or more computing devices, a search space including a plurality of candidate permutations of a plurality of candidate feature blocks, each of the plurality of candidate feature blocks having a respective scale. The method can include performing, by the computing system, a plurality of search iterations by a search algorithm to select a scale-permuted model from the search space, the scale-permuted model based at least in part on a candidate permutation of the plurality of candidate permutations.Type: ApplicationFiled: July 25, 2024Publication date: November 14, 2024Inventors: Xianzhi Du, Yin Cui, Tsung-Yi Lin, Quoc V. Le, Pengchong Jin, Mingxing Tan, Golnaz Ghiasi, Xiaodan Song
-
Publication number: 20240362460Abstract: The technology relates to providing personalized neural network-based models according to user input, which can be generated upon request or otherwise as needed. This may include receiving, by one or more processors of a computing device, input corresponding to a task description. Then the input corresponding to the task description is encoded into a set of text embeddings. Based on this, the system applies mixer prediction to the set of text embeddings to generate a set of mixers and learns a set of basis models according to the set of mixers. The set of basis models are combined to form a single personalized model corresponding to the task description. This personalized model can then be used in video understanding, quality assessment, providing a recommendation, performing a classification, or performing a search.Type: ApplicationFiled: April 4, 2024Publication date: October 31, 2024Inventors: Li Zhang, Yandong Li, Yin Cui, Hong-You Chen, Mingda Zhang
-
Patent number: 12079695Abstract: A computer-implemented method of generating scale-permuted models can generate models having improved accuracy and reduced evaluation computational requirements. The method can include defining, by a computing system including one or more computing devices, a search space including a plurality of candidate permutations of a plurality of candidate feature blocks, each of the plurality of candidate feature blocks having a respective scale. The method can include performing, by the computing system, a plurality of search iterations by a search algorithm to select a scale-permuted model from the search space, the scale-permuted model based at least in part on a candidate permutation of the plurality of candidate permutations.Type: GrantFiled: October 1, 2020Date of Patent: September 3, 2024Assignee: GOOGLE LLCInventors: Xianzhi Du, Yin Cui, Tsung-Yi Lin, Quoc V. Le, Pengchong Jin, Mingxing Tan, Golnaz Ghiasi, Xiaodan Song
-
Publication number: 20240282131Abstract: Systems and methods for zero-shot prompt ensembling for zero-shot classification with text-image models can include utilizing a pre-trained text-image model to perform downstream tasks based on prompt-based weighting. The systems and methods may adjust for frequency-based bias and may automatically determine different prompt associations with a given downstream task. The systems and methods can aggregate weighted text embeddings and then determine a classification output based on similarity measures between an image embedding and the aggregated weighted text embeddings.Type: ApplicationFiled: January 24, 2024Publication date: August 22, 2024Inventors: Jie Ren, Zhe Liu, James Urquhart Allingham, Michael Ward Dusenberry, Dustin Tran, Yin Cui, Balaji Lakshminarayanan, Xiuye Gu
-
Publication number: 20220108204Abstract: A computer-implemented method of generating scale-permuted models can generate models having improved accuracy and reduced evaluation computational requirements. The method can include defining, by a computing system including one or more computing devices, a search space including a plurality of candidate permutations of a plurality of candidate feature blocks, each of the plurality of candidate feature blocks having a respective scale. The method can include performing, by the computing system, a plurality of search iterations by a search algorithm to select a scale-permuted model from the search space, the scale-permuted model based at least in part on a candidate permutation of the plurality of candidate permutations.Type: ApplicationFiled: October 1, 2020Publication date: April 7, 2022Inventors: Xianzhi Du, Yin Cui, Tsung-Yi Lin, Quoc V. Le, Pengchong Jin, Mingxing Tan, Golnaz Ghiasi, Xiaodan Song