Patents by Inventor Shuming MA

Shuming MA has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Publication number: 20260023956
    Abstract: A computer system is provided that includes processing circuitry. The computer system being configured to implement a machine learning (ML) model having a transformer architecture that, during a training operation or inference operation, is configured to receive an activation input matrix of activation input values and obtain a weight matrix of weight values. The ML model is further configured to perform ultra-low precision (ULP) quantization by quantizing each of the weight values in the weight matrix to a corresponding selected value from a predefined set of binary or ternary quantized weight values and compute a matrix arithmetic result based on at least a portion of the weight matrix with the quantized weight values and at least a portion of the activation input matrix.
    Type: Application
    Filed: July 16, 2024
    Publication date: January 22, 2026
    Applicant: Microsoft Technology Licensing, LLC
    Inventors: Shuming MA, Li DONG, Shaohan HUANG, Wenhui WANG, Furu WEI, Jilong XUE, Lingxiao MA, Hongyu WANG
  • Publication number: 20250202679
    Abstract: A server computing device is provided, including a processor configured to receive a homomorphically encrypted input embedding vector from a client computing device. At a transformer network, the processor may generate a plurality of homomorphically encrypted intermediate vectors at least in part by performing inferencing on the homomorphically encrypted input embedding vector. The processor may transmit the plurality of homomorphically encrypted intermediate output vectors to the client computing device. The processor may receive a plurality of homomorphically encrypted intermediate input vectors from the client computing device subsequently to transmitting the homomorphically encrypted intermediate output vectors to the client computing device. At the transformer network, the processor may generate a homomorphically encrypted output vector at least in part by performing additional inferencing on the homomorphically encrypted intermediate input vectors.
    Type: Application
    Filed: March 30, 2022
    Publication date: June 19, 2025
    Applicant: Microsoft Technology Licensing, LLC
    Inventors: Shaohan HUANG, Li DONG, Shuming MA, Furu WEI
  • Publication number: 20240320482
    Abstract: A computing system is provided, including a processor configured to receive a training data set. Based at least in part on the training data set, the processor is further configured to train a transformer network that includes a plurality of layers. The plurality of layers each respectively include a plurality of sub-layers including an attention sub-layer, a feed-forward sub-layer, and a plurality of normalization sub-layers. The plurality of normalization sub-layers are downstream from corresponding sub-layers of the plurality of sub-layers. Each of the plurality of normalization sub-layers is configured to apply layer normalization to a sum of: a first scaling parameter multiplied by an input vector of the sub-layer; and an output vector of the sub-layer.
    Type: Application
    Filed: February 28, 2023
    Publication date: September 26, 2024
    Applicant: Microsoft Technology Licensing, LLC
    Inventors: Shuming MA, Li DONG, Shaohan HUANG, Dongdong ZHANG, Furu WEI, Hongyu WANG