Patents by Inventor Qichen FU

Qichen FU has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Publication number: 20250037018
    Abstract: The subject technology provides memory-efficient differentiable weight clustering for large language model compression. An apparatus determines a tensor including an attention map between learned weights of a trained machine learning model and corresponding centroids. The apparatus also determines a compressed attention table and a plurality of index lists during compression of the trained machine learning model based on an uniquification of the attention map and sharding of an associated index list. The apparatus determines whether the tensor exists at a destination device during compression of the trained machine learning model using a marshaling layer. The apparatus refrains from copying the tensor to the destination device when the tensor exists at the destination device, or copies the tensor to the destination device when the tensor does not exist at the destination device. The apparatus deploys a compressed machine learning model based on the compression of the trained machine learning model.
    Type: Application
    Filed: May 8, 2024
    Publication date: January 30, 2025
    Inventors: Minsik CHO, Keivan ALIZADEH VAHID, Qichen FU, Saurabh ADYA, Carlo Eduardo Cabanero DEL MUNDO, Mohammad RASTEGARI, Devang K. NAIK, Peter ZATLOUKAL