Patents by Inventor Haijun Shan

Haijun Shan has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Patent number: 11669741
    Abstract: Disclosed is a method for meta-knowledge fine-tuning and platform based on domain-invariant features. According to the method, highly transferable common knowledge, i.e., domain-invariant features, in different data sets of the same kind of tasks is learnt, the common domain features in different domains corresponding to different data sets of the same kind of tasks learnt in the network set are fine-tuned to be quickly adapted to any different domains. According to the present application, the parameter initialization ability and generalization ability of the universal language model of the same kind of tasks are improved, and finally a common compression framework of the universal language model of the same kind of downstream tasks is obtained through fine tuning. In the meta-knowledge fine-tuning network, a loss function of the domain-invariant features is designed in the present application, and domain-independent universal knowledge is learn.
    Type: Grant
    Filed: February 18, 2022
    Date of Patent: June 6, 2023
    Assignee: ZHEJIANG LAB
    Inventors: Hongsheng Wang, Haijun Shan, Shengjian Hu
  • Patent number: 11526774
    Abstract: Disclosed is a method for automatically compressing multi-task oriented pre-trained language model and a platform thereof. According to the method, a meta-network of a structure generator is designed, a knowledge distillation coding vector is constructed based on a knowledge distillation method of Transformer layer sampling, and a distillation structure model corresponding to a currently input coding vector is generated by using the structure generator; at the same time, a Bernoulli distribution sampling method is provided for training the structure generator; in each iteration, each encoder unit is transferred by Bernoulli distribution sampling to form a corresponding coding vector; by changing the coding vector input to the structure generator and a small batch of training data, the structure generator and the corresponding distillation structure are jointly trained, and a structure generator capable of generating weights for different distillation structures can be acquired.
    Type: Grant
    Filed: December 28, 2021
    Date of Patent: December 13, 2022
    Assignee: ZHEJIANG LAB
    Inventors: Hongsheng Wang, Haijun Shan, Jiaqing Fu
  • Publication number: 20220222529
    Abstract: Disclosed is a method for meta-knowledge fine-tuning and platform based on domain-invariant features. According to the method, highly transferable common knowledge, i.e., domain-invariant features, in different data sets of the same kind of tasks is learnt, the common domain features in different domains corresponding to different data sets of the same kind of tasks learnt in the network set are fine-tuned to be quickly adapted to any different domains. According to the present application, the parameter initialization ability and generalization ability of the universal language model of the same kind of tasks are improved, and finally a common compression framework of the universal language model of the same kind of downstream tasks is obtained through fine tuning. In the meta-knowledge fine-tuning network, a loss function of the domain-invariant features is designed in the present application, and domain-independent universal knowledge is learn.
    Type: Application
    Filed: February 18, 2022
    Publication date: July 14, 2022
    Inventors: Hongsheng WANG, Haijun SHAN, Shengjian HU
  • Publication number: 20220188658
    Abstract: Disclosed is a method for automatically compressing multi-task oriented pre-trained language model and a platform thereof. According to the method, a meta-network of a structure generator is designed, a knowledge distillation coding vector is constructed based on a knowledge distillation method of Transformer layer sampling, and a distillation structure model corresponding to a currently input coding vector is generated by using the structure generator; at the same time, a Bernoulli distribution sampling method is provided for training the structure generator; in each iteration, each encoder unit is transferred by Bernoulli distribution sampling to form a corresponding coding vector; by changing the coding vector input to the structure generator and a small batch of training data, the structure generator and the corresponding distillation structure are jointly trained, and a structure generator capable of generating weights for different distillation structures can be acquired.
    Type: Application
    Filed: December 28, 2021
    Publication date: June 16, 2022
    Inventors: Hongsheng WANG, Haijun SHAN, Jiaqing FU
  • Patent number: 11354499
    Abstract: Disclosed is a meta-knowledge fine tuning method and platform for a multi-task language model. The method is to obtain highly transferable shared knowledge, that is, meta-knowledge, on different data sets of tasks of the same category, perform interrelation and mutual reinforcement on the learning processes of the tasks of the same category that correspond to different data sets and are in different domains, so as to improve the fine tuning effect of downstream tasks of the same category on data sets of different domains in the application of the language model, and improve the parameter initialization ability and the generalization ability of a general language model for the tasks of the same category.
    Type: Grant
    Filed: November 22, 2021
    Date of Patent: June 7, 2022
    Assignee: ZHEJIANG LAB
    Inventors: Hongsheng Wang, Haijun Shan, Shengjian Hu
  • Patent number: 11341326
    Abstract: Provided is a method and a platform for compressing a pre-training language model based on knowledge distillation. According to the method, a universal knowledge distillation strategy of feature migration is firstly designed, and in the process of knowledge distillation from the teacher model to the student model, the feature mapping of each layer of the student model is approaching the teacher's features, focusing on the ability of small samples to express features in the intermediate layer of the teacher model, and guiding the student model by using these features; then, a knowledge distillation method based on self-attention cross is constructed; finally, a linear transfer strategy based on Bernoulli probability distribution is designed to gradually complete the knowledge transfer of feature mapping and self-attention distribution from teachers to students.
    Type: Grant
    Filed: September 24, 2021
    Date of Patent: May 24, 2022
    Assignee: ZHEJIANG LAB
    Inventors: Hongsheng Wang, Haijun Shan, Fei Yang
  • Publication number: 20220138414
    Abstract: Disclosed is a meta-knowledge fine tuning method and platform for a multi-task language model. The method is to obtain highly transferable shared knowledge, that is, meta-knowledge, on different data sets of tasks of the same category, perform interrelation and mutual reinforcement on the learning processes of the tasks of the same category that correspond to different data sets and are in different domains, so as to improve the fine tuning effect of downstream tasks of the same category on data sets of different domains in the application of the language model, and improve the parameter initialization ability and the generalization ability of a general language model for the tasks of the same category.
    Type: Application
    Filed: November 22, 2021
    Publication date: May 5, 2022
    Inventors: Hongsheng Wang, Haijun Shan, Shengjian Hu
  • Publication number: 20220067274
    Abstract: Provided is a method and a platform for compressing a pre-training language model based on knowledge distillation. According to the method, a universal knowledge distillation strategy of feature migration is firstly designed, and in the process of knowledge distillation from the teacher model to the student model, the feature mapping of each layer of the student model is approaching the teacher's features, focusing on the ability of small samples to express features in the intermediate layer of the teacher model, and guiding the student model by using these features; then, a knowledge distillation method based on self-attention cross is constructed; finally, a linear transfer strategy based on Bernoulli probability distribution is designed to gradually complete the knowledge transfer of feature mapping and self-attention distribution from teachers to students.
    Type: Application
    Filed: September 24, 2021
    Publication date: March 3, 2022
    Inventors: Hongsheng WANG, Haijun SHAN, Fei YANG
  • Patent number: 11216310
    Abstract: A capacity expansion method includes obtaining a measured workload of a service of an application, obtaining an application model of the application, and obtaining a measured workload of each upper-level service of the service; determining a predicted workload of the service based on the measured workload of the service, determining the measured workload of each upper-level service of the first service, and determining a first workload ratio corresponding to a first calling relationship; and determining a predicted workload of each lower-level service based on the predicted workload of the service and determining a second workload ratio corresponding to a second calling relationship.
    Type: Grant
    Filed: July 26, 2019
    Date of Patent: January 4, 2022
    Assignee: HUAWEI TECHNOLOGIES CO., LTD.
    Inventors: Donghui Zhuo, Jun Xu, Haijun Shan
  • Publication number: 20190347134
    Abstract: Embodiments of this application relate to a capacity expansion method. In this method, a measured workload of a service of an application and an application model of the application a measured workload of each upper-level service of the service is obtained. Then, a predicted workload of the service based on the measured workload of the service, the measured workload of each upper-level service of the first service, and a first workload ratio corresponding to a first calling relationship are determined. Then, a predicted workload of each lower-level service is determined based on the predicted workload of the service and a second workload ratio corresponding to a second calling relationship.
    Type: Application
    Filed: July 26, 2019
    Publication date: November 14, 2019
    Inventors: Donghui Zhuo, Jun Xu, Haijun Shan