Patents by Inventor Lu Yuan
Lu Yuan has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Publication number: 20260249102Abstract: Systems and methods for predicting radiation dose distribution and planning a radiotherapy treatment are disclosed. An exemplary system includes a memory to store a computational model (such as a trained machine learning model), a dose prediction engine to predict a dose profile, and a treatment planning circuit. The dose prediction engine executes a dose simulation to determine a preliminary dose profile at a first statistical uncertainty level, applies the determined preliminary dose profile to the computational model to predict a refined dose profile at a second statistical uncertainty level lower than the first statistical uncertainty level. Based at least in part on refined dose profile, the treatment planning system can generate or update a radiotherapy treatment plan for use in a radiation treatment session.Type: ApplicationFiled: February 14, 2023Publication date: August 27, 2026Inventors: Shufei Chen, Lu Yuan
-
Patent number: 12711758Abstract: Example solutions for ranking object detection results generate or receive a plurality of segmentation masks each corresponding to one or more images. Each segmentation mask of each plurality of segmentation masks is generated using a different object detector or setting options. A quality predictor assigns a quality score to each segmentation mask, without using ground truth for the image(s). A set (one or more, but less than all) of the highest quality scores is identified for each image. In some examples, an image processing task is performed using the segmentation masks having an assigned quality score that is within the set of highest quality scores. In some examples, only the segmentation mask having the highest quality score for an image is used in the image processing task. In some examples, a quality threshold is provided, and the segmentation masks meeting the quality threshold are used in the image processing task.Type: GrantFiled: February 2, 2024Date of Patent: August 18, 2026Assignee: Microsoft Technology Licensing, LLCInventors: Dongdong Chen, Yunsheng Li, Lu Yuan, Xuelu Feng
-
Patent number: 12657940Abstract: Example solutions for image paragraph captioning use a first vision language model to generate visual information (comprising text) for an image. The visual information may include tags, an initial image caption, and information on objects within the image (e.g., further tags and captions, and object attributes and locations within the image). In some examples, the visual information further includes visual clues. A generative language model generates a plurality of image story caption candidates (e.g., descriptive paragraphs) from the visual information. A second vision language model evaluates the plurality of image story caption candidates and selects a caption as the final output caption.Type: GrantFiled: June 29, 2022Date of Patent: June 16, 2026Assignee: Microsoft Technology Licensing, LLC.Inventors: Yujia Xie, Lu Yuan, Nguyen Hung Bach
-
Publication number: 20260127865Abstract: Examples are provided for pre-training a computer vision foundation model. A representative method comprises curating a pre-training database of image-text pairs from weakly labeled data. Language is encoded of text descriptions from the image-text pairs. The images of the image-text pairs are encoded using a hierarchical vision transformer with shifted windows and convolutional embedding. Based on the encoded images and the encoded language, the computer vision foundation model is pre-trained via unified image-text contrastive learning.Type: ApplicationFiled: January 5, 2026Publication date: May 7, 2026Applicant: Microsoft Technology Licensing, LLCInventors: Lu YUAN, Chunyuan LI, Jianwei YANG, Bin XIAO
-
Patent number: 12619301Abstract: A method and an apparatus is for adjusting a display device and includes: obtaining a first facial image of a first user; determining first spatial coordinates of a facial feature point of the first user based on the first facial image; and adjusting, based on the first spatial coordinates, an orientation of a first side of an image displayed by the display device, where the first side of the image includes information to be communicated to the first user.Type: GrantFiled: November 8, 2024Date of Patent: May 5, 2026Assignee: SHENZHEN YINWANG INTELLIGENT TECHNOLOGIES CO., LTD.Inventors: Lu Yuan, Jingfang Zha
-
Patent number: 12596908Abstract: A neural architecture search (NAS) with a weak predictor comprises: receiving network architecture scoring information; iteratively sampling a search space, wherein the sampling comprises: generating a set of candidate architectures within the search space; learning a first predictor; evaluating performance of the candidate architectures; and based on at least the performance of the set of candidate architectures and the network architecture scoring information, refining the search space to a smaller search space; based on at least the network architecture scoring information, thresholding the performance of candidate architectures to determine scored output candidate architectures; and reporting the scored output candidate architectures. In some examples, the candidate architectures each comprise a machine learning (ML) model, for example a neural network (NN).Type: GrantFiled: December 15, 2020Date of Patent: April 7, 2026Assignee: Microsoft Technology Licensing, LLC.Inventors: Xiyang Dai, Dongdong Chen, Yinpeng Chen, Mengchen Liu, Ye Yu, Zicheng Liu, Mei Chen, Lu Yuan, Junru Wu
-
Publication number: 20260065434Abstract: The disclosure herein describes training an encoder network to inpaint images with masked portions. A primary encoding process is used to encode a visible portion of a masked input image into encoded token data. The encoded token data is then decoded into both pixel regression output and feature prediction output, wherein both outputs include inpainted image data associated with the masked portion of the masked input image. A pixel regression loss is determined using the pixel regression output and pixel data of an unmasked version of the masked input image. A feature prediction loss is determined using the feature prediction output and ground truth encoding output of the unmasked version of the masked input image. The primary encoding process is then trained using the pixel regression loss and the feature prediction loss, whereby the primary encoding process is trained to encode structural features of input images into encoded token data.Type: ApplicationFiled: November 10, 2025Publication date: March 5, 2026Inventors: Dongdong CHEN, Jianmin BAO, Ting ZHANG, Lu YUAN, Dong CHEN, Fang WEN, Xiaoyi DONG
-
Patent number: 12518512Abstract: Examples are provided for pre-training a computer vision foundation model. A representative method comprises curating a pre-training database of image-text pairs from weakly labeled data. Language is encoded of text descriptions from the image-text pairs. The images of the image-text pairs are encoded using a hierarchical vision transformer with shifted windows and convolutional embedding. Based on the encoded images and the encoded language, the computer vision foundation model is pre-trained via unified image-text contrastive learning.Type: GrantFiled: August 23, 2022Date of Patent: January 6, 2026Assignee: Microsoft Technology Licensing, LLCInventors: Lu Yuan, Chunyuan Li, Jianwei Yang, Bin Xiao
-
Patent number: 12488430Abstract: The disclosure herein describes training an encoder network to inpaint images with masked portions. A primary encoding process is used to encode a visible portion of a masked input image into encoded token data. The encoded token data is then decoded into both pixel regression output and feature prediction output, wherein both outputs include inpainted image data associated with the masked portion of the masked input image. A pixel regression loss is determined using the pixel regression output and pixel data of an unmasked version of the masked input image. A feature prediction loss is determined using the feature prediction output and ground truth encoding output of the unmasked version of the masked input image. The primary encoding process is then trained using the pixel regression loss and the feature prediction loss, whereby the primary encoding process is trained to encode structural features of input images into encoded token data.Type: GrantFiled: May 19, 2022Date of Patent: December 2, 2025Assignee: Microsoft Technology Licensing, LLC.Inventors: Dongdong Chen, Jianmin Bao, Ting Zhang, Lu Yuan, Dong Chen, Fang Wen, Xiaoyi Dong
-
Patent number: 12420112Abstract: Systems and methods for generating a beam model for radiotherapy treatment planning are discussed. An exemplary system includes a memory to store a trained deep learning model, and a processor circuit to generate a beam model. The deep learning model can be trained to establish a relationship between machine scanning data and values of beam model parameters, and validated for accuracy. The processor circuit can receive machine scanning data indicative of a configuration or an operation status of the radiation therapy device, apply the machine scanning data to the trained deep learning model to determine values for the beam model parameters, and generate a beam model based on the determined values of the plurality of beam model parameters. The beam model may be provided to a user, or a treatment planning system.Type: GrantFiled: September 2, 2020Date of Patent: September 23, 2025Assignee: Elekta (Shanghai) Technology Co., Ltd.Inventors: Shufei Chen, Lu Yuan
-
Publication number: 20250278843Abstract: Training a multi-object tracking model includes: generating a plurality of training images based at least on scene generation information, each training image comprising a plurality of objects to be tracked; generating, for each training image, original simulated data based at least on the scene generation information, the original simulated data comprising tag data for a first object; locating, within the original simulated data, tag data for the first object, based on at least an anomaly alert (e.g., occlusion alert, proximity alert, motion alert) associated with the first object in the first training image; based at least on locating the tag data for the first object, modifying at least a portion of the tag data for the first object from the original simulated data, thereby generating preprocessed training data from the original simulated data; and training a multi-object tracking model with the preprocessed training data to produce a trained multi-object tracker.Type: ApplicationFiled: May 20, 2025Publication date: September 4, 2025Inventors: Ishani CHAKRABORTY, Jonathan C. HANZELKA, Lu YUAN, Pedro Urbina ESCOS, Thomas M. SOEMO
-
Publication number: 20250252705Abstract: Example solutions for pluralistic salient object detection are disclosed. A received image shows multiple objects, such as a first object and a second object. A pluralistic object detector is trained to learn tokens. When provided with the image and the first token, it generates a first segmentation mask corresponding to the first object, but not the second object, and when provided with the image and the second token, it generates a second segmentation mask corresponding to at least the second object (and possibly also the first image). When the pluralistic object detector is trained on five tokens, up to five different segmentation masks, each corresponding to a different selection of up to five objects, may be generated. Additionally, a quality predictor is disclosed that assigns quality scores to each of the different segmentation masks, without requiring ground truth for the image.Type: ApplicationFiled: February 2, 2024Publication date: August 7, 2025Inventors: Dongdong CHEN, Yunsheng LI, Lu YUAN, Xuelu FENG
-
Publication number: 20250252731Abstract: Example solutions for ranking object detection results generate or receive a plurality of segmentation masks each corresponding to one or more images. Each segmentation mask of each plurality of segmentation masks is generated using a different object detector or setting options. A quality predictor assigns a quality score to each segmentation mask, without using ground truth for the image(s). A set (one or more, but less than all) of the highest quality scores is identified for each image. In some examples, an image processing task is performed using the segmentation masks having an assigned quality score that is within the set of highest quality scores. In some examples, only the segmentation mask having the highest quality score for an image is used in the image processing task. In some examples, a quality threshold is provided, and the segmentation masks meeting the quality threshold are used in the image processing task.Type: ApplicationFiled: February 2, 2024Publication date: August 7, 2025Inventors: Dongdong CHEN, Yunsheng LI, Lu YUAN, Xuelu FENG
-
Patent number: 12340521Abstract: Training a multi-object tracking model includes: generating a plurality of training images based at least on scene generation information, each training image comprising a plurality of objects to be tracked; generating, for each training image, original simulated data based at least on the scene generation information, the original simulated data comprising tag data for a first object; locating, within the original simulated data, tag data for the first object, based on at least an anomaly alert (e.g., occlusion alert, proximity alert, motion alert) associated with the first object in the first training image; based at least on locating the tag data for the first object, modifying at least a portion of the tag data for the first object from the original simulated data, thereby generating preprocessed training data from the original simulated data; and training a multi-object tracking model with the preprocessed training data to produce a trained multi-object tracker.Type: GrantFiled: November 13, 2023Date of Patent: June 24, 2025Assignee: Microsoft Technology Licensing, LLCInventors: Ishani Chakraborty, Jonathan C. Hanzelka, Lu Yuan, Pedro Urbina Escos, Thomas M. Soemo
-
Publication number: 20250148765Abstract: A method for annotating images to create a corpus for training a multi-task computer vision machine learning model is presented. The method comprises receiving, at one or more annotation specialist models, a plurality of images to be annotated. Via operation of the one or more annotation specialist models, pre-filtered annotations are generated for the plurality of images. Via operation of a data filtering and enhancement module, the pre-filtered annotations are filtered in accordance with predefined noise criteria so as to output candidate annotations for the plurality of images. The method further comprises, for each of one or more candidate annotations, selectively (1) storing the candidate annotation into the corpus as a final annotation for its associated image, or (2) adding the candidate annotation to its associated image using the one or more annotation specialist models and the data filtering and enhancement module for subsequent iterative annotation and filtering.Type: ApplicationFiled: January 30, 2024Publication date: May 8, 2025Applicant: Microsoft Technology Licensing, LLCInventors: Lu YUAN, Bin XIAO, Haiping WU, Weijian XU, Xiyang DAI, Houdong HU, Yumao LU, Nanshan ZENG, Ce Christopher LIU
-
Publication number: 20250123683Abstract: A method and an apparatus is for adjusting a display device and includes: obtaining a first facial image of a first user; determining first spatial coordinates of a facial feature point of the first user based on the first facial image; and adjusting, based on the first spatial coordinates, an orientation of a first side of an image displayed by the display device, where the first side of the image includes information to be communicated to the first user.Type: ApplicationFiled: November 8, 2024Publication date: April 17, 2025Inventors: Lu Yuan, Jingfang Zha
-
Patent number: 12223412Abstract: A computer device for automatic feature detection comprises a processor, a communication device, and a memory configured to hold instructions executable by the processor to instantiate a dynamic convolution neural network, receive input data via the communication network, and execute the dynamic convolution neural network to automatically detect features in the input data. The dynamic convolution neural network compresses the input data from an input space having a dimensionality equal to a predetermined number of channels into an intermediate space having a dimensionality less than the number of channels. The dynamic convolution neural network dynamically fuses the channels into an intermediate representation within the intermediate space and expands the intermediate representation from the intermediate space to an expanded representation in an output space having a higher dimensionality than the dimensionality of the intermediate space.Type: GrantFiled: December 16, 2020Date of Patent: February 11, 2025Assignee: Microsoft Technology Licensing, LLCInventors: Yinpeng Chen, Xiyang Dai, Mengchen Liu, Dongdong Chen, Lu Yuan, Zicheng Liu, Ye Yu, Mei Chen, Yunsheng Li
-
Publication number: 20250037252Abstract: The disclosure herein describes generating an inpainted image from a masked image using a patch-based encoder and an unquantized transformer. An image including a masked region and an unmasked region is received, and the received image is divided into a plurality of patches including masked patches. The plurality of patches is encoded into a plurality of feature vectors, wherein each patch is encoded to a feature vector. Using a transformer, a predicted token is generated for each masked patch using a feature vector encoded from the masked patch, and a quantized vector of the masked patch is determined using generated predicted token and a masked patch-specific codebook. The determined quantized vector of the masked patch is included into a set of quantized vectors associated with the plurality of patches, and an output image is generated from the set of quantized vectors using a decoder.Type: ApplicationFiled: October 11, 2024Publication date: January 30, 2025Inventors: Dongdong CHEN, Xiyang DAI, Yinpeng CHEN, Mengchen LIU, Lu YUAN
-
Patent number: 12190588Abstract: A system for tracking a target object across a plurality of image frames. The system comprises a logic machine and a storage machine. The storage machine holds instructions executable by the logic machine to calculate a trajectory for the target object over one or more previous frames occurring before a target frame. Responsive to assessing no detection of the target object in the target frame, the instructions are executable to predict an estimated region for the target object based on the trajectory, predict an occlusion center based on a set of candidate occluding locations for a set of other objects within a threshold distance of the estimated region, each location of the set of candidate occluding locations overlapping with the estimated region, and automatically estimate a bounding box for the target object in the target frame based on the occlusion center.Type: GrantFiled: June 4, 2021Date of Patent: January 7, 2025Assignee: Microsoft Technology Licensing, LLCInventors: Dongdong Chen, Qiankun Liu, Lu Yuan, Lei Zhang
-
Patent number: 12148131Abstract: The disclosure herein describes generating an inpainted image from a masked image using a patch-based encoder and an unquantized transformer. An image including a masked region and an unmasked region is received, and the received image is divided into a plurality of patches including masked patches. The plurality of patches is encoded into a plurality of feature vectors, wherein each patch is encoded to a feature vector. Using a transformer, a predicted token is generated for each masked patch using a feature vector encoded from the masked patch, and a quantized vector of the masked patch is determined using generated predicted token and a masked patch-specific codebook. The determined quantized vector of the masked patch is included into a set of quantized vectors associated with the plurality of patches, and an output image is generated from the set of quantized vectors using a decoder.Type: GrantFiled: April 29, 2022Date of Patent: November 19, 2024Assignee: Microsoft Technology Licensing, LLC.Inventors: Dongdong Chen, Xiyang Dai, Yinpeng Chen, Mengchen Liu, Lu Yuan