Patents by Inventor Xiaodan Liang

Xiaodan Liang has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Patent number: 12670359
    Abstract: A neural network construction method and apparatus in the field of artificial intelligence, to accurately and efficiently construct a target neural network. The constructed target neural network has high output accuracy, may be further applied to different application scenarios, and has a strong generalization capability. The method includes: obtaining a start point network, where the start point network includes a plurality of serial subnets; performing at least one time of transformation on the start point network based on a preset first search space to obtain a serial network, where the first search space includes a range of parameters used for transforming the start point network; and if the serial network meets a preset condition, training the serial network by using a preset dataset to obtain a trained serial network; and if the trained serial network meets a termination condition, obtaining a target neural network based on the trained serial network.
    Type: Grant
    Filed: November 23, 2022
    Date of Patent: June 30, 2026
    Assignee: HUAWEI TECHNOLOGIES CO., LTD.
    Inventors: Chenhan Jiang, Hang Xu, Zhenguo Li, Xiaodan Liang
  • Patent number: 12664710
    Abstract: A system for performing artificial intelligence (AI)-based customized storytelling video generation includes a story designer AI agent, a storyboard generator AI agent, a video creator AI agent, an agent manager AI agent, and an observer AI agent. Coordinated by the agent manager AI agent, the agents collaboratively process a textual prompt and a reference video provided by a user to generate a multi-shot video depicting a story of a customized subject from the reference video. The story designer, agent manager, and observer AI agents leverage Large Language Models (LLMs), while the storyboard generator AI agent employs a three-step pipeline of generation, removal, and redrawing to maintain character detail consistency across video shots. The video creator AI agent utilizes a Latent Diffusion Model (LDM) based Image-to-Video (I2V) model to ensure intra-shot character detail consistency. The system achieves high-quality, coherent storytelling videos with customizable subject fidelity.
    Type: Grant
    Filed: January 27, 2026
    Date of Patent: June 23, 2026
    Assignee: MOHAMED BIN ZAYED UNIVERSITY OF ARTIFICIAL INTELLIGENCE
    Inventors: Panwen Hu, Xiaodan Liang
  • Publication number: 20260170302
    Abstract: A method is provided for data processing performed by a processing system. The method comprises determining a set of first tokens for first data and a set of second token for second data, each token comprising information associated with a segment of the respective data, determining pair-wise similarities between the set of first tokens and the set of second tokens, each pair comprising a first token in the set of first tokens and a second token in the set of second tokens, determining, for each first token in the set of first tokens, a maximum similarity based on the determined pair-wise similarities between the respective first token and the second tokens in the set of second tokens, and determining a first similarity between the first data and the second data by aggregating the maximum similarities corresponding to the first tokens in the set of first set of tokens.
    Type: Application
    Filed: January 30, 2026
    Publication date: June 18, 2026
    Applicant: HUAWEI TECHNOLOGIES CO., LTD.
    Inventors: Hang Xu, Lu Hou, Guansong Lu, Minzhe Niu, Zhenguo Li, Runhui Huang, Lewei Yao, Chunjing Xu, Xiaodan Liang
  • Patent number: 12657896
    Abstract: A processing apparatus includes a collection module and a training module, the training module includes a backbone network and a region proposal network (RPN) layer, the backbone network is connected to the RPN layer, and the RPN layer includes a class activation map (CAM) unit. The collection module is configured to obtain an image, where the image includes an image with an instance-level label and an image with an image-level label. The backbone network is used to output a feature map of the image based on the image obtained by the collection module.
    Type: Grant
    Filed: August 3, 2022
    Date of Patent: June 16, 2026
    Assignee: HUAWEI TECHNOLOGIES CO., LTD.
    Inventors: Hang Xu, Zhili Liu, Fengwei Zhou, Jiawei Li, Xiaodan Liang, Zhenguo Li, Li Qian
  • Publication number: 20260141694
    Abstract: This application discloses an image processing method, a method for training a limb part image prediction model, an apparatus, a computer device, a computer-readable storage medium, and a computer program product. The method includes: obtaining a first limb image, the first limb image being an image from a first perspective; calling a feature encoding network to perform image encoding on the first limb image, to obtain a fused feature representation of each of at least two query points on a limb object, an nth fused feature representation indicating a feature representation of an image region with symmetry in a physiological structure; calling a decoding network to perform feature decoding on fused feature representations of the at least two query points, to obtain decoded features; and performing rendering based on the decoded features to obtain a second limb image, the second limb image being image information from a second perspective.
    Type: Application
    Filed: January 16, 2026
    Publication date: May 21, 2026
    Applicant: Tencent Technology (Shenzhen) Company Limited
    Inventors: Xuan HUANG, Hanhui LI, Zejun YANG, Zhisheng WANG, Xiaodan LIANG
  • Patent number: 12633089
    Abstract: Provided is a three-dimensional object detection framework based on multi-source data knowledge transfer. By outputting an image feature extracted by an image feature extraction unit, an interested target selection unit outputs point cloud data of an interested target to a point cloud feature extraction unit according to the image feature; the point cloud feature extraction unit extracts a point cloud feature from the point cloud data; in a knowledge transfer unit, enable the image feature to learn the point cloud feature and update parameters of the image feature extraction unit, while a three-dimensional target parameter prediction unit updates parameters of the image feature and point cloud feature extraction units according to the image feature and the point cloud feature. Finally, the updated image feature extraction unit re-extracts the image feature to the three-dimensional target parameter prediction unit, which reckons and inputs three-dimensional parameters according to the image feature.
    Type: Grant
    Filed: January 28, 2021
    Date of Patent: May 19, 2026
    Assignee: Sun Yat-sen University
    Inventors: Xiaojun Tan, Dapeng Feng, Xiaodan Liang, Huanyu Wang, Chenrushi Yang, Mengyu Yang
  • Patent number: 12572780
    Abstract: A method is provided for data processing performed by a processing system. The method comprises determining a set of first tokens for first data and a set of second token for second data, each token comprising information associated with a segment of the respective data, determining pair-wise similarities between the set of first tokens and the set of second tokens, each pair comprising a first token in the set of first tokens and a second token in the set of second tokens, determining, for each first token in the set of first tokens, a maximum similarity based on the determined pair-wise similarities between the respective first token and the second tokens in the set of second tokens, and determining a first similarity between the first data and the second data by aggregating the maximum similarities corresponding to the first tokens in the set of first set of tokens.
    Type: Grant
    Filed: August 31, 2022
    Date of Patent: March 10, 2026
    Assignee: Huawei Technologies Co., Ltd.
    Inventors: Hang Xu, Lu Hou, Guansong Lu, Minzhe Niu, Zhenguo Li, Runhui Huang, Lewei Yao, Chunjing Xu, Xiaodan Liang
  • Publication number: 20260004501
    Abstract: Motion generation model-based motion generation method, device, and storage medium relate to the field of artificial intelligence technologies. The method includes: obtaining a text containing motion information; generating a text feature of the text through a text encoder; generating an intermediate motion sequence in a feature space of a first dimension based on the text feature through a first diffusion model; and performing detail enhancement processing on the intermediate motion sequence in a feature space of a second dimension through a second diffusion model, to obtain an output motion sequence matching the text, the second dimension being greater than the first dimension. In this application, the intermediate motion sequence is preliminarily generated through the first diffusion model, and detail enhancement processing is performed on the intermediate motion sequence through the second diffusion model, thereby improving the richness of details in the output motion sequence.
    Type: Application
    Filed: September 5, 2025
    Publication date: January 1, 2026
    Applicant: Tencent Technology (Shenzhen) Company Limited
    Inventors: Yang WU, Zhenyu XIE, Zhongqian SUN, Wei YANG, Xiaodan LIANG
  • Publication number: 20240070436
    Abstract: A method is provided for data processing performed by a processing system. The method comprises determining a set of first tokens for first data and a set of second token for second data, each token comprising information associated with a segment of the respective data, determining pair-wise similarities between the set of first tokens and the set of second tokens, each pair comprising a first token in the set of first tokens and a second token in the set of second tokens, determining, for each first token in the set of first tokens, a maximum similarity based on the determined pair-wise similarities between the respective first token and the second tokens in the set of second tokens, and determining a first similarity between the first data and the second data by aggregating the maximum similarities corresponding to the first tokens in the set of first set of tokens.
    Type: Application
    Filed: August 31, 2022
    Publication date: February 29, 2024
    Inventors: Hang XU, Lu HOU, Guansong LU, Minzhe NIU, Zhenguo LI, Runhui HUANG, Lewei YAO, Chunjing XU, Xiaodan LIANG
  • Publication number: 20230260255
    Abstract: Provided is a three-dimensional object detection framework based on multi-source data knowledge transfer. By outputting an image feature extracted by an image feature extraction unit, an interested target selection unit outputs point cloud data of an interested target to a point cloud feature extraction unit according to the image feature; the point cloud feature extraction unit extracts a point cloud feature from the point cloud data; in a knowledge transfer unit, enable the image feature to learn the point cloud feature and update parameters of the image feature extraction unit, while a three-dimensional target parameter prediction unit updates parameters of the image feature and point cloud feature extraction units according to the image feature and the point cloud feature. Finally, the updated image feature extraction unit re-extracts the image feature to the three-dimensional target parameter prediction unit, which reckons and inputs three-dimensional parameters according to the image feature.
    Type: Application
    Filed: January 28, 2021
    Publication date: August 17, 2023
    Inventors: Xiaojun TAN, Dapeng FENG, Xiaodan LIANG, Huanyu WANG, Chenrushi YANG, Mengyu YANG
  • Publication number: 20230089380
    Abstract: A neural network construction method and apparatus in the field of artificial intelligence, to accurately and efficiently construct a target neural network. The constructed target neural network has high output accuracy, may be further applied to different application scenarios, and has a strong generalization capability. The method includes: obtaining a start point network, where the start point network includes a plurality of serial subnets; performing at least one time of transformation on the start point network based on a preset first search space to obtain a serial network, where the first search space includes a range of parameters used for transforming the start point network; and if the serial network meets a preset condition, training the serial network by using a preset dataset to obtain a trained serial network; and if the trained serial network meets a termination condition, obtaining a target neural network based on the trained serial network.
    Type: Application
    Filed: November 23, 2022
    Publication date: March 23, 2023
    Inventors: Chenhan JIANG, Hang XU, Zhenguo LI, Xiaodan LIANG
  • Patent number: 11543830
    Abstract: An unsupervised real to virtual domain unification model for highway driving, or DU-drive, employs a conditional generative adversarial network to transform driving images in a real domain to their canonical representations in the virtual domain, from which vehicle control commands are predicted. In the case where there are multiple real datasets, a real-to-virtual generator may be independently trained for each real domain and a global predictor could be trained with data from multiple real domains. Qualitative experiment results show this model can effectively transform real images to the virtual domain while only keeping the minimal sufficient information, and quantitative results verify that such canonical representation can eliminate domain shift and boost the performance of control command prediction task.
    Type: Grant
    Filed: September 24, 2018
    Date of Patent: January 3, 2023
    Assignee: PETUUM, INC.
    Inventors: Xiaodan Liang, Eric P Xing
  • Publication number: 20220375213
    Abstract: A processing apparatus includes a collection module and a training module, the training module includes a backbone network and a region proposal network (RPN) layer, the backbone network is connected to the RPN layer, and the RPN layer includes a class activation map (CAM) unit. The collection module is configured to obtain an image, where the image includes an image with an instance-level label and an image with an image-level label. The backbone network is used to output a feature map of the image based on the image obtained by the collection module.
    Type: Application
    Filed: August 3, 2022
    Publication date: November 24, 2022
    Inventors: Hang Xu, Zhili Liu, Fengwei Zhou, Jiawei Li, Xiaodan Liang, Zhenguo Li, Li Qian
  • Publication number: 20220261659
    Abstract: This application provides a method and related apparatus for determining a neural network in the field of artificial intelligence. The method includes: obtaining a plurality of initial search spaces; determining M candidate neural networks based on the plurality of initial search spaces, where the candidate neural network includes a plurality of candidate subnetworks, the plurality of candidate subnetworks belong to the plurality of initial search spaces, and any two of the plurality of candidate subnetworks belong to different initial search spaces; evaluating the M candidate neural networks to obtain M evaluation results; and determining N candidate neural networks from the M candidate neural networks based on the M evaluation results, and determining N first target neural networks based on the N candidate neural networks. According to the method and the related apparatus provided in this application, a combined neural network with relatively high performance can be obtained.
    Type: Application
    Filed: May 6, 2022
    Publication date: August 18, 2022
    Inventors: Hang Xu, Zhenguo Li, Wei Zhang, Xiaodan Liang, Chenhan Jiang
  • Patent number: 11282205
    Abstract: Organ segmentation in chest X-rays using convolutional neural networks is disclosed. One embodiment provides a method to train a convolutional segmentation network with chest X-ray images to generate pixel-level predictions of target classes. Another embodiment will also train a critic network with an input mask, wherein the input mask is one of a segmentation network mask and a ground truth annotation, and outputting a probability that the input mask is the ground truth annotation instead of the prediction by the segmentation network, and to provide the probability output by the critic network to the segmentation network to guide the segmentation network to generate masks more consistent with learned higher-order structures.
    Type: Grant
    Filed: April 1, 2020
    Date of Patent: March 22, 2022
    Assignee: PETUUM, INC.
    Inventors: Wei Dai, Xiaodan Liang, Hao Zhang, Eric Xing, Joseph Doyle
  • Publication number: 20200234448
    Abstract: Organ segmentation in chest X-rays using convolutional neural networks is disclosed. One embodiment provides a method to train a convolutional segmentation network with chest X-ray images to generate pixel-level predictions of target classes. Another embodiment will also train a critic network with an input mask, wherein the input mask is one of a segmentation network mask and a ground truth annotation, and outputting a probability that the input mask is the ground truth annotation instead of the prediction by the segmentation network, and to provide the probability output by the critic network to the segmentation network to guide the segmentation network to generate masks more consistent with learned higher-order structures.
    Type: Application
    Filed: April 1, 2020
    Publication date: July 23, 2020
    Inventors: Wei Dai, Xiaodan Liang, Hao Zhang, Eric Xing, Joseph Doyle
  • Patent number: 10699412
    Abstract: Organ segmentation in chest X-rays using convolutional neural networks is disclosed. One embodiment provides a method to train a convolutional segmentation network with chest X-ray images to generate pixel-level predictions of target classes. Another embodiment will also train a critic network with an input mask, wherein the input mask is one of a segmentation network mask and a ground truth annotation, and outputting a probability that the input mask is the ground truth annotation instead of the prediction by the segmentation network, and to provide the probability output by the critic network to the segmentation network to guide the segmentation network to generate masks more consistent with learned higher-order structures.
    Type: Grant
    Filed: March 20, 2018
    Date of Patent: June 30, 2020
    Assignee: PETUUM INC.
    Inventors: Wei Dai, Xiaodan Liang, Hao Zhang, Eric Xing, Joseph Doyle
  • Publication number: 20190171223
    Abstract: An unsupervised real to virtual domain unification model for highway driving, or DU-drive, employs a conditional generative adversarial network to transform driving images in a real domain to their canonical representations in the virtual domain, from which vehicle control commands are predicted. In the case where there are multiple real datasets, a real-to-virtual generator may be independently trained for each real domain and a global predictor could be trained with data from multiple real domains. Qualitative experiment results show this model can effectively transform real images to the virtual domain while only keeping the minimal sufficient information, and quantitative results verify that such canonical representation can eliminate domain shift and boost the performance of control command prediction task.
    Type: Application
    Filed: September 24, 2018
    Publication date: June 6, 2019
    Inventors: Xiaodan Liang, Eric P. Xing
  • Publication number: 20180276825
    Abstract: Organ segmentation in chest X-rays using convolutional neural networks is disclosed. One embodiment provides a method to train a convolutional segmentation network with chest X-ray images to generate pixel-level predictions of target classes. Another embodiment will also train a critic network with an input mask, wherein the input mask is one of a segmentation network mask and a ground truth annotation, and outputting a probability that the input mask is the ground truth annotation instead of the prediction by the segmentation network, and to provide the probability output by the critic network to the segmentation network to guide the segmentation network to generate masks more consistent with learned higher-order structures.
    Type: Application
    Filed: March 20, 2018
    Publication date: September 27, 2018
    Inventors: Wei Dai, Xiaodan Liang, Hao Zhang, Eric Xing, Joseph Doyle