Patents by Inventor Xiaodan Liang
Xiaodan Liang has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Patent number: 12670359Abstract: A neural network construction method and apparatus in the field of artificial intelligence, to accurately and efficiently construct a target neural network. The constructed target neural network has high output accuracy, may be further applied to different application scenarios, and has a strong generalization capability. The method includes: obtaining a start point network, where the start point network includes a plurality of serial subnets; performing at least one time of transformation on the start point network based on a preset first search space to obtain a serial network, where the first search space includes a range of parameters used for transforming the start point network; and if the serial network meets a preset condition, training the serial network by using a preset dataset to obtain a trained serial network; and if the trained serial network meets a termination condition, obtaining a target neural network based on the trained serial network.Type: GrantFiled: November 23, 2022Date of Patent: June 30, 2026Assignee: HUAWEI TECHNOLOGIES CO., LTD.Inventors: Chenhan Jiang, Hang Xu, Zhenguo Li, Xiaodan Liang
-
Patent number: 12664710Abstract: A system for performing artificial intelligence (AI)-based customized storytelling video generation includes a story designer AI agent, a storyboard generator AI agent, a video creator AI agent, an agent manager AI agent, and an observer AI agent. Coordinated by the agent manager AI agent, the agents collaboratively process a textual prompt and a reference video provided by a user to generate a multi-shot video depicting a story of a customized subject from the reference video. The story designer, agent manager, and observer AI agents leverage Large Language Models (LLMs), while the storyboard generator AI agent employs a three-step pipeline of generation, removal, and redrawing to maintain character detail consistency across video shots. The video creator AI agent utilizes a Latent Diffusion Model (LDM) based Image-to-Video (I2V) model to ensure intra-shot character detail consistency. The system achieves high-quality, coherent storytelling videos with customizable subject fidelity.Type: GrantFiled: January 27, 2026Date of Patent: June 23, 2026Assignee: MOHAMED BIN ZAYED UNIVERSITY OF ARTIFICIAL INTELLIGENCEInventors: Panwen Hu, Xiaodan Liang
-
Publication number: 20260170302Abstract: A method is provided for data processing performed by a processing system. The method comprises determining a set of first tokens for first data and a set of second token for second data, each token comprising information associated with a segment of the respective data, determining pair-wise similarities between the set of first tokens and the set of second tokens, each pair comprising a first token in the set of first tokens and a second token in the set of second tokens, determining, for each first token in the set of first tokens, a maximum similarity based on the determined pair-wise similarities between the respective first token and the second tokens in the set of second tokens, and determining a first similarity between the first data and the second data by aggregating the maximum similarities corresponding to the first tokens in the set of first set of tokens.Type: ApplicationFiled: January 30, 2026Publication date: June 18, 2026Applicant: HUAWEI TECHNOLOGIES CO., LTD.Inventors: Hang Xu, Lu Hou, Guansong Lu, Minzhe Niu, Zhenguo Li, Runhui Huang, Lewei Yao, Chunjing Xu, Xiaodan Liang
-
Patent number: 12657896Abstract: A processing apparatus includes a collection module and a training module, the training module includes a backbone network and a region proposal network (RPN) layer, the backbone network is connected to the RPN layer, and the RPN layer includes a class activation map (CAM) unit. The collection module is configured to obtain an image, where the image includes an image with an instance-level label and an image with an image-level label. The backbone network is used to output a feature map of the image based on the image obtained by the collection module.Type: GrantFiled: August 3, 2022Date of Patent: June 16, 2026Assignee: HUAWEI TECHNOLOGIES CO., LTD.Inventors: Hang Xu, Zhili Liu, Fengwei Zhou, Jiawei Li, Xiaodan Liang, Zhenguo Li, Li Qian
-
Publication number: 20260141694Abstract: This application discloses an image processing method, a method for training a limb part image prediction model, an apparatus, a computer device, a computer-readable storage medium, and a computer program product. The method includes: obtaining a first limb image, the first limb image being an image from a first perspective; calling a feature encoding network to perform image encoding on the first limb image, to obtain a fused feature representation of each of at least two query points on a limb object, an nth fused feature representation indicating a feature representation of an image region with symmetry in a physiological structure; calling a decoding network to perform feature decoding on fused feature representations of the at least two query points, to obtain decoded features; and performing rendering based on the decoded features to obtain a second limb image, the second limb image being image information from a second perspective.Type: ApplicationFiled: January 16, 2026Publication date: May 21, 2026Applicant: Tencent Technology (Shenzhen) Company LimitedInventors: Xuan HUANG, Hanhui LI, Zejun YANG, Zhisheng WANG, Xiaodan LIANG
-
Patent number: 12633089Abstract: Provided is a three-dimensional object detection framework based on multi-source data knowledge transfer. By outputting an image feature extracted by an image feature extraction unit, an interested target selection unit outputs point cloud data of an interested target to a point cloud feature extraction unit according to the image feature; the point cloud feature extraction unit extracts a point cloud feature from the point cloud data; in a knowledge transfer unit, enable the image feature to learn the point cloud feature and update parameters of the image feature extraction unit, while a three-dimensional target parameter prediction unit updates parameters of the image feature and point cloud feature extraction units according to the image feature and the point cloud feature. Finally, the updated image feature extraction unit re-extracts the image feature to the three-dimensional target parameter prediction unit, which reckons and inputs three-dimensional parameters according to the image feature.Type: GrantFiled: January 28, 2021Date of Patent: May 19, 2026Assignee: Sun Yat-sen UniversityInventors: Xiaojun Tan, Dapeng Feng, Xiaodan Liang, Huanyu Wang, Chenrushi Yang, Mengyu Yang
-
Patent number: 12572780Abstract: A method is provided for data processing performed by a processing system. The method comprises determining a set of first tokens for first data and a set of second token for second data, each token comprising information associated with a segment of the respective data, determining pair-wise similarities between the set of first tokens and the set of second tokens, each pair comprising a first token in the set of first tokens and a second token in the set of second tokens, determining, for each first token in the set of first tokens, a maximum similarity based on the determined pair-wise similarities between the respective first token and the second tokens in the set of second tokens, and determining a first similarity between the first data and the second data by aggregating the maximum similarities corresponding to the first tokens in the set of first set of tokens.Type: GrantFiled: August 31, 2022Date of Patent: March 10, 2026Assignee: Huawei Technologies Co., Ltd.Inventors: Hang Xu, Lu Hou, Guansong Lu, Minzhe Niu, Zhenguo Li, Runhui Huang, Lewei Yao, Chunjing Xu, Xiaodan Liang
-
Publication number: 20260004501Abstract: Motion generation model-based motion generation method, device, and storage medium relate to the field of artificial intelligence technologies. The method includes: obtaining a text containing motion information; generating a text feature of the text through a text encoder; generating an intermediate motion sequence in a feature space of a first dimension based on the text feature through a first diffusion model; and performing detail enhancement processing on the intermediate motion sequence in a feature space of a second dimension through a second diffusion model, to obtain an output motion sequence matching the text, the second dimension being greater than the first dimension. In this application, the intermediate motion sequence is preliminarily generated through the first diffusion model, and detail enhancement processing is performed on the intermediate motion sequence through the second diffusion model, thereby improving the richness of details in the output motion sequence.Type: ApplicationFiled: September 5, 2025Publication date: January 1, 2026Applicant: Tencent Technology (Shenzhen) Company LimitedInventors: Yang WU, Zhenyu XIE, Zhongqian SUN, Wei YANG, Xiaodan LIANG
-
Publication number: 20240070436Abstract: A method is provided for data processing performed by a processing system. The method comprises determining a set of first tokens for first data and a set of second token for second data, each token comprising information associated with a segment of the respective data, determining pair-wise similarities between the set of first tokens and the set of second tokens, each pair comprising a first token in the set of first tokens and a second token in the set of second tokens, determining, for each first token in the set of first tokens, a maximum similarity based on the determined pair-wise similarities between the respective first token and the second tokens in the set of second tokens, and determining a first similarity between the first data and the second data by aggregating the maximum similarities corresponding to the first tokens in the set of first set of tokens.Type: ApplicationFiled: August 31, 2022Publication date: February 29, 2024Inventors: Hang XU, Lu HOU, Guansong LU, Minzhe NIU, Zhenguo LI, Runhui HUANG, Lewei YAO, Chunjing XU, Xiaodan LIANG
-
Publication number: 20230260255Abstract: Provided is a three-dimensional object detection framework based on multi-source data knowledge transfer. By outputting an image feature extracted by an image feature extraction unit, an interested target selection unit outputs point cloud data of an interested target to a point cloud feature extraction unit according to the image feature; the point cloud feature extraction unit extracts a point cloud feature from the point cloud data; in a knowledge transfer unit, enable the image feature to learn the point cloud feature and update parameters of the image feature extraction unit, while a three-dimensional target parameter prediction unit updates parameters of the image feature and point cloud feature extraction units according to the image feature and the point cloud feature. Finally, the updated image feature extraction unit re-extracts the image feature to the three-dimensional target parameter prediction unit, which reckons and inputs three-dimensional parameters according to the image feature.Type: ApplicationFiled: January 28, 2021Publication date: August 17, 2023Inventors: Xiaojun TAN, Dapeng FENG, Xiaodan LIANG, Huanyu WANG, Chenrushi YANG, Mengyu YANG
-
Publication number: 20230089380Abstract: A neural network construction method and apparatus in the field of artificial intelligence, to accurately and efficiently construct a target neural network. The constructed target neural network has high output accuracy, may be further applied to different application scenarios, and has a strong generalization capability. The method includes: obtaining a start point network, where the start point network includes a plurality of serial subnets; performing at least one time of transformation on the start point network based on a preset first search space to obtain a serial network, where the first search space includes a range of parameters used for transforming the start point network; and if the serial network meets a preset condition, training the serial network by using a preset dataset to obtain a trained serial network; and if the trained serial network meets a termination condition, obtaining a target neural network based on the trained serial network.Type: ApplicationFiled: November 23, 2022Publication date: March 23, 2023Inventors: Chenhan JIANG, Hang XU, Zhenguo LI, Xiaodan LIANG
-
Patent number: 11543830Abstract: An unsupervised real to virtual domain unification model for highway driving, or DU-drive, employs a conditional generative adversarial network to transform driving images in a real domain to their canonical representations in the virtual domain, from which vehicle control commands are predicted. In the case where there are multiple real datasets, a real-to-virtual generator may be independently trained for each real domain and a global predictor could be trained with data from multiple real domains. Qualitative experiment results show this model can effectively transform real images to the virtual domain while only keeping the minimal sufficient information, and quantitative results verify that such canonical representation can eliminate domain shift and boost the performance of control command prediction task.Type: GrantFiled: September 24, 2018Date of Patent: January 3, 2023Assignee: PETUUM, INC.Inventors: Xiaodan Liang, Eric P Xing
-
Publication number: 20220375213Abstract: A processing apparatus includes a collection module and a training module, the training module includes a backbone network and a region proposal network (RPN) layer, the backbone network is connected to the RPN layer, and the RPN layer includes a class activation map (CAM) unit. The collection module is configured to obtain an image, where the image includes an image with an instance-level label and an image with an image-level label. The backbone network is used to output a feature map of the image based on the image obtained by the collection module.Type: ApplicationFiled: August 3, 2022Publication date: November 24, 2022Inventors: Hang Xu, Zhili Liu, Fengwei Zhou, Jiawei Li, Xiaodan Liang, Zhenguo Li, Li Qian
-
Publication number: 20220261659Abstract: This application provides a method and related apparatus for determining a neural network in the field of artificial intelligence. The method includes: obtaining a plurality of initial search spaces; determining M candidate neural networks based on the plurality of initial search spaces, where the candidate neural network includes a plurality of candidate subnetworks, the plurality of candidate subnetworks belong to the plurality of initial search spaces, and any two of the plurality of candidate subnetworks belong to different initial search spaces; evaluating the M candidate neural networks to obtain M evaluation results; and determining N candidate neural networks from the M candidate neural networks based on the M evaluation results, and determining N first target neural networks based on the N candidate neural networks. According to the method and the related apparatus provided in this application, a combined neural network with relatively high performance can be obtained.Type: ApplicationFiled: May 6, 2022Publication date: August 18, 2022Inventors: Hang Xu, Zhenguo Li, Wei Zhang, Xiaodan Liang, Chenhan Jiang
-
Patent number: 11282205Abstract: Organ segmentation in chest X-rays using convolutional neural networks is disclosed. One embodiment provides a method to train a convolutional segmentation network with chest X-ray images to generate pixel-level predictions of target classes. Another embodiment will also train a critic network with an input mask, wherein the input mask is one of a segmentation network mask and a ground truth annotation, and outputting a probability that the input mask is the ground truth annotation instead of the prediction by the segmentation network, and to provide the probability output by the critic network to the segmentation network to guide the segmentation network to generate masks more consistent with learned higher-order structures.Type: GrantFiled: April 1, 2020Date of Patent: March 22, 2022Assignee: PETUUM, INC.Inventors: Wei Dai, Xiaodan Liang, Hao Zhang, Eric Xing, Joseph Doyle
-
Publication number: 20200234448Abstract: Organ segmentation in chest X-rays using convolutional neural networks is disclosed. One embodiment provides a method to train a convolutional segmentation network with chest X-ray images to generate pixel-level predictions of target classes. Another embodiment will also train a critic network with an input mask, wherein the input mask is one of a segmentation network mask and a ground truth annotation, and outputting a probability that the input mask is the ground truth annotation instead of the prediction by the segmentation network, and to provide the probability output by the critic network to the segmentation network to guide the segmentation network to generate masks more consistent with learned higher-order structures.Type: ApplicationFiled: April 1, 2020Publication date: July 23, 2020Inventors: Wei Dai, Xiaodan Liang, Hao Zhang, Eric Xing, Joseph Doyle
-
Patent number: 10699412Abstract: Organ segmentation in chest X-rays using convolutional neural networks is disclosed. One embodiment provides a method to train a convolutional segmentation network with chest X-ray images to generate pixel-level predictions of target classes. Another embodiment will also train a critic network with an input mask, wherein the input mask is one of a segmentation network mask and a ground truth annotation, and outputting a probability that the input mask is the ground truth annotation instead of the prediction by the segmentation network, and to provide the probability output by the critic network to the segmentation network to guide the segmentation network to generate masks more consistent with learned higher-order structures.Type: GrantFiled: March 20, 2018Date of Patent: June 30, 2020Assignee: PETUUM INC.Inventors: Wei Dai, Xiaodan Liang, Hao Zhang, Eric Xing, Joseph Doyle
-
Publication number: 20190171223Abstract: An unsupervised real to virtual domain unification model for highway driving, or DU-drive, employs a conditional generative adversarial network to transform driving images in a real domain to their canonical representations in the virtual domain, from which vehicle control commands are predicted. In the case where there are multiple real datasets, a real-to-virtual generator may be independently trained for each real domain and a global predictor could be trained with data from multiple real domains. Qualitative experiment results show this model can effectively transform real images to the virtual domain while only keeping the minimal sufficient information, and quantitative results verify that such canonical representation can eliminate domain shift and boost the performance of control command prediction task.Type: ApplicationFiled: September 24, 2018Publication date: June 6, 2019Inventors: Xiaodan Liang, Eric P. Xing
-
Publication number: 20180276825Abstract: Organ segmentation in chest X-rays using convolutional neural networks is disclosed. One embodiment provides a method to train a convolutional segmentation network with chest X-ray images to generate pixel-level predictions of target classes. Another embodiment will also train a critic network with an input mask, wherein the input mask is one of a segmentation network mask and a ground truth annotation, and outputting a probability that the input mask is the ground truth annotation instead of the prediction by the segmentation network, and to provide the probability output by the critic network to the segmentation network to guide the segmentation network to generate masks more consistent with learned higher-order structures.Type: ApplicationFiled: March 20, 2018Publication date: September 27, 2018Inventors: Wei Dai, Xiaodan Liang, Hao Zhang, Eric Xing, Joseph Doyle