Patents by Inventor Markus Nagel

Markus Nagel has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Publication number: 20260161998
    Abstract: Aspects described herein provide techniques for performing quantization robust federated learning of a machine learning model, comprising: receiving a model from a federated learning server; training the model using a local objective function, wherein the local objective function includes a modification configured to increase quantization robustness at a client device; and transmitting to the federated learning server an updated model.
    Type: Application
    Filed: January 4, 2023
    Publication date: June 11, 2026
    Inventors: Kartik GUPTA, Marios FOURNARAKIS, Matthias REISSER, Christos LOUIZOS, Markus NAGEL
  • Publication number: 20260093956
    Abstract: A processor-implemented method for providing parameter-free attention operations includes receiving, by an attention mechanism of a machine learning model, an input. The attention mechanism generates a set of matrices based on the input. An attention matrix is generated based on a reconstruction objective computed based on a linear combination of the input. The machine learning model computes an output based on the attention matrix.
    Type: Application
    Filed: September 25, 2025
    Publication date: April 2, 2026
    Inventors: Hanno ACKERMANN, Leyla MIRVAKHABOVA, Hong CAI, Fatih Murat PORIKLI, Farhad GHAZVINIAN ZANJANI, Markus NAGEL
  • Publication number: 20260093953
    Abstract: Certain aspects of the present disclosure provide techniques and apparatus for cache aware dynamic module selection for a computation model. An example method generally includes generating at least one output, in a first inference round, using a first subset of modules of a computational model loaded in a cache memory from another memory, evaluating modules of the computational model to use for a second inference round, using a function that biases evaluation of the first subset of modules of the computational model already in the cache, and performing the second inference round with a second subset of modules of the computational module, based on the evaluation.
    Type: Application
    Filed: September 30, 2024
    Publication date: April 2, 2026
    Inventors: Marinus Willem VAN BAALEN, Davide BELLI, Andrii SKLIAR, Bence MAJOR, Markus NAGEL, Babak EHTESHAMI BEJNORDI, Paul Nicholas WHATMOUGH, Marco FEDERICI, Amir JALALIRAD
  • Publication number: 20260087386
    Abstract: Systems and techniques are described herein for processing tokens. For instance, a method for processing tokens is provided. The method may include processing a token at a router model to generate a recommendation a subset of expert models from a plurality of expert models to use for further processing of the token; selecting a number of expert models to use for the further processing of the token based on the recommendation of the subset of expert models and based on cached expert models of the plurality of expert models stored in a cache memory; and processing the token using the selected number of expert models.
    Type: Application
    Filed: January 10, 2025
    Publication date: March 26, 2026
    Inventors: Andrii SKLIAR, Babak EHTESHAMI BEJNORDI, Ties Jehan VAN ROZENDAAL, Marinus Willem VAN BAALEN, Markus NAGEL, Paul Nicholas WHATMOUGH
  • Publication number: 20260073199
    Abstract: Systems and techniques are described herein for adjusting weights of a machine learning (ML) model. For instance, a process can include generating a first matrix of quantized weight values by rounding values of an input matrix of weight values for the ML model; applying an activation function to a second matrix, the second matrix generated based on a third matrix and a fourth matrix of a first matrix pair; applying the activation function to a fifth matrix, the fifth matrix based on a sixth matrix and seventh matrix of a second matrix pair; generating a positive second matrix by applying a positive factor to the second matrix; generating a negative fifth matrix by applying a negative factor to the fifth matrix; and summing the first matrix of quantized weight values with the positive second matrix and the negative fifth matrix to generate an output matrix of quantized weight values.
    Type: Application
    Filed: September 11, 2024
    Publication date: March 12, 2026
    Inventors: Jaeseong YOU, Minseop PARK, Jinkyu LEE, Seunghan YANG, Yoonhyung LEE, Sweta PRIYADARSHI, Markus NAGEL
  • Patent number: 12561567
    Abstract: Various embodiments include methods and devices for neural network pruning. Embodiments may include receiving as an input a weight tensor for a neural network, increasing a level of sparsity of the weight tensor generating a sparse weight tensor, updating the neural network using the sparse weight tensor generating an updated weight tensor, decreasing a level of sparsity of the updated weight tensor generating a dense weight tensor, increasing the level of sparsity of the dense weight tensor the dense weight tensor generating a final sparse weight tensor, and using the neural network with the final sparse weight tensor to generate inferences. Some embodiments may include increasing a level of sparsity of a first sparse weight tensor generating a second sparse weight tensor, updating the neural network using the second sparse weight tensor generating a second updated weight tensor, and decreasing the level of sparsity the second updated weight tensor.
    Type: Grant
    Filed: November 23, 2021
    Date of Patent: February 24, 2026
    Assignee: QUALCOMM Incorporated
    Inventors: Suraj Srinivas, Tijmen Pieter Frederik Blankevoort, Andrey Kuzmin, Markus Nagel, Marinus Willem Van Baalen, Andrii Skliar
  • Publication number: 20260017561
    Abstract: Certain aspects of the present disclosure provide techniques and apparatus for machine learning. In an example method, a first plurality of quantization scales for a set of machine learning model parameters is accessed, and a shared quantization scale for the set of machine learning model parameters is accessed. A second plurality of quantization scales is generated based on the shared quantization scale and the first plurality of quantization scales. A dequantized set of machine learning model parameters is generated based on the shared quantization scale and the second plurality of quantization scales. A machine learning model output is generated based on the dequantized set of machine learning model parameters.
    Type: Application
    Filed: September 19, 2024
    Publication date: January 15, 2026
    Inventors: Nilesh Prasad PANDEY, Jun MA, Markus NAGEL, Kevin Lishing HSIEH, An CHEN, Chirag Sureshbhai PATEL, Abhijit KHOBARE, Liang ZHANG, Eric Wayne MAHURIN, Muralidhar Reddy AKULA, Joseph Binamira SORIAGA
  • Publication number: 20260010784
    Abstract: Certain aspects of the present disclosure provide techniques and apparatus for machine learning. In an example method, a weight tensor for a layer of a machine learning model is determined, where the weight tensor comprises per-block values in a first precision encoding. The weight tensor is upscaled to a second precision encoding having a higher precision than the first precision encoding to generate an upscaled weight tensor, and an input tensor for the layer of the machine learning model is accessed. An output tensor for the layer of the machine learning model is generated based on multiplying the upscaled weight tensor with the input tensor.
    Type: Application
    Filed: July 2, 2024
    Publication date: January 8, 2026
    Inventors: Weiliang ZENG, Marinus Willem VAN BAALEN, Markus NAGEL, Paul Nicholas WHATMOUGH, An CHEN, Liang ZHANG, Chirag Sureshbhai PATEL, Nilesh Prasad PANDEY, Jun MA, Kevin Lishing HSIEH, Abhijit KHOBARE, Joseph Binamira SORIAGA, Muralidhar Reddy AKULA, Peter John COUPERUS, Cedric BASTOUL
  • Patent number: 12499345
    Abstract: A processor-implemented method includes bit shifting a binary representation of a neural network parameter. The neural network parameter has fewer bits, b, than a number of hardware bits, B, supported by hardware that processes the neural network parameter. The bit shifting effectively multiplies the neural network parameter by 2B-b. The method also includes dividing a quantization scale by 2B-b to obtain an updated quantization scale. The method further includes quantizing the bit shifted binary representation with the updated quantization scale to obtain a value for the neural network parameter.
    Type: Grant
    Filed: January 30, 2023
    Date of Patent: December 16, 2025
    Assignee: QUALCOMM Incorporated
    Inventors: Marinus Willem Van Baalen, Brian Kahne, Eric Wayne Mahurin, Tijmen Pieter Frederik Blankevoort, Andrey Kuzmin, Andrii Skliar, Markus Nagel
  • Publication number: 20250356185
    Abstract: A processor-implemented method includes receiving an artificial neural network having a number of pre-trained weights. The method also includes training a subset of the number of pre-trained weights to obtain trained sparse adapter weights for obtaining a fine-tuned version of the artificial neural network. The subset of the number of pre-trained weights includes base model weights of a base model for the artificial neural network. The subset of the number of pre-trained weights is selected with a sparse mask of a sparse adapter. The method may also include replacing the subset of the number of pre-trained weights with the trained sparse adapter weights.
    Type: Application
    Filed: September 12, 2024
    Publication date: November 20, 2025
    Inventors: Kartikeya BHARDWAJ, Nilesh Prasad PANDEY, Sweta PRIYADARSHI, Shubhankar Mangesh BORSE, Shreya KADAMBI, Rafael Xavier ESTEVES, Viswanath GANAPATHY, Risheek GARREPALLI, Paul Nicholas WHATMOUGH, Marinus Willem VAN BAALEN, Markus NAGEL
  • Publication number: 20250356245
    Abstract: Certain aspects of the present disclosure provide techniques and apparatus for improved machine learning. In an example method, a first plurality of weights for a base model and a second plurality of weights for an adapter model associated with the base model are accessed. A quantized plurality of weights is generated based on the first plurality of weights, a first quantization scale for the first plurality of weights, and the second plurality of weights. A loss is generated based on processing training data using the quantized plurality of weights. An updated second plurality of weights is generated based on updating the second plurality of weights based on the loss. A machine learning model comprising quantized versions of the first plurality of weights and the updated second plurality of weights is deployed.
    Type: Application
    Filed: May 15, 2024
    Publication date: November 20, 2025
    Inventors: Yelysei BONDARENKO, Markus NAGEL, Riccardo DEL CHIARO
  • Publication number: 20250245494
    Abstract: Systems and techniques are described herein for quantizing a codebook used in the context of quantizing post-training parameters (e.g., vectors of weights) of a pre-trained model. For example, a device can perform rank reduction on a tensor of a codebook associated with parameters of a layer of a pre-trained machine learning model to generate a first tensor factor having a first shape and a second tensor factor having a second shape. The device can perform an optimization technique on the first tensor factor and the second tensor factor to minimize an output reconstruction error of the layer. The device can quantize the first tensor factor to generate a reduced size codebook.
    Type: Application
    Filed: August 30, 2024
    Publication date: July 31, 2025
    Inventors: Marinus Willem VAN BAALEN, Eric Wayne MAHURIN, Paul Nicholas WHATMOUGH, Andrey KUZMIN, Markus NAGEL, Tijmen Pieter Frederik BLANKEVOORT
  • Publication number: 20250245567
    Abstract: Systems and techniques are described for quantizing parameters (e.g., post-training vectors) associated with a pre-trained model. For example, a device can obtain a codebook for a group of weights of a pre-trained machine learning model. The device can determine a compression ratio based on the codebook and at least one of a vector quantization dimensionality, a group size, a codebook bit-width, or a scale group size. The device can quantize, via a vector quantization engine, the group of weights of the pre-trained machine learning model a plurality of columns at a time according to the compression ratio to generate a quantized pre-trained model.
    Type: Application
    Filed: October 22, 2024
    Publication date: July 31, 2025
    Inventors: Andrey KUZMIN, Marinus Willem VAN BAALEN, Paul Nicholas WHATMOUGH, Markus NAGEL, Tijmen Pieter Frederik BLANKEVOORT
  • Patent number: 12373697
    Abstract: Various embodiments include methods and devices for joint mixed-precision quantization and structured pruning. Embodiments may include determining whether a plurality of gates of quantization and pruning gates are selected for combination, and in response to determining that the plurality of gates are selected for combination, iteratively for each successive gate of the plurality of gates selected for combination quantizing a residual error of a quantized tensor to a scale of a next bit-width producing a residual error quantized tensor in which the next bit-width increases for each successive iteration, and adding the quantized tensor and the residual error quantized tensor producing a next quantized tensor in which the next quantized tensor has the next bit-width, and in which the next quantized tensor is the quantized tensor for a successive iteration.
    Type: Grant
    Filed: April 29, 2021
    Date of Patent: July 29, 2025
    Assignee: QUALCOMM Incorporated
    Inventors: Marinus Willem Van Baalen, Christos Louizos, Markus Nagel, Tijmen Pieter Frederik Blankevoort, Rana Ali Amjad
  • Publication number: 20250111232
    Abstract: An apparatus has one or more memories and one or more processor(s) coupled to the memories. The processor(s) is configured to estimate a local curvature of a loss landscape of a neural network. The processor(s) is also configured to dynamically allocate parameters to be removed from the neural network based on the local curvature. The processor(s) is further configured to update remaining weights of the neural network based on the parameters to be removed.
    Type: Application
    Filed: September 28, 2023
    Publication date: April 3, 2025
    Inventors: Tycho VAN DER OUDERAA, Markus NAGEL, Marinus Willem VAN BAALEN, Tijmen Pieter Frederik BLANKEVOORT
  • Patent number: 12242956
    Abstract: Various embodiments include methods and neural network computing devices implementing the methods for performing quantization in neural networks. Various embodiments may include equalizing ranges of weight tensors or output channel weights within a first layer of the neural network by scaling each of the output channel weights of the first layer by a corresponding scaling factor, and scaling each of a second adjacent layer's corresponding input channel weights by applying an inverse of the corresponding scaling factor to the input channel weights. The corresponding scaling factor may be determined using a black-box optimizer on a quantization error metric or based on heuristics, equalization of dynamic ranges, equalization of range extrema (minima or maxima), differential learning using straight through estimator (STE) methods and a local or global loss, or using an error metric for the quantization error and a black-box optimizer that minimizes the error metric with respect to the scaling.
    Type: Grant
    Filed: March 23, 2020
    Date of Patent: March 4, 2025
    Assignee: QUALCOMM Incorporated
    Inventors: Markus Nagel, Marinus Willem van Baalen, Tijmen Pieter Frederik Blankevoort
  • Patent number: 12206851
    Abstract: Techniques are described for compressing and decompressing data using machine learning systems. An example process can include receiving a plurality of images for compression by a neural network compression system. The process can include determining, based on a first image from the plurality of images, a first plurality of weight values associated with a first model of the neural network compression system. The process can include generating a first bitstream comprising a compressed version of the first plurality of weight values. The process can include outputting the first bitstream for transmission to a receiver.
    Type: Grant
    Filed: December 17, 2021
    Date of Patent: January 21, 2025
    Assignee: QUALCOMM INCORPORATED
    Inventors: Yunfan Zhang, Ties Jehan Van Rozendaal, Taco Sebastiaan Cohen, Markus Nagel, Johann Hinrich Brehmer
  • Publication number: 20250005452
    Abstract: Certain aspects of the present disclosure provide techniques and apparatus for mitigating weight oscillation during quantization-aware training. In one example, a method includes identifying oscillation of a parameter of a machine learning model during quantization-aware training of the machine learning model, and applying an oscillation mitigation procedure during the quantization-aware training of the machine learning model in response to identifying the oscillation, the oscillation mitigation procedure comprising at least one of oscillation dampening or parameter freezing.
    Type: Application
    Filed: January 24, 2023
    Publication date: January 2, 2025
    Inventors: Markus NAGEL, Marios FOURNARAKIS, Tijmen Pieter Frederik BLANKEVOORT, Yelysei BONDARENKO
  • Publication number: 20240386239
    Abstract: Certain aspects of the present disclosure provide techniques and apparatus for processing data using a transformer neural network. The method generally includes receiving an input for processing using a transformer neural network. An attention output is generated in the transformer neural network. Generally, the attention output may be generated such that outlier values for the attention output are attenuated in the transformer neural network. An output of the transformer neural network is generated based on the generated attention output.
    Type: Application
    Filed: October 6, 2023
    Publication date: November 21, 2024
    Inventors: Yelysei BONDARENKO, Markus NAGEL, Tijmen Pieter Frederik BLANKEVOORT
  • Publication number: 20240169708
    Abstract: Certain aspects of the present disclosure provide techniques and apparatus for delta quantization for video processing and other data streams with temporal content. An example method generally includes receiving image data including at least a first frame and a second frame, generating a first convolutional output based on a first frame using a machine learning model, generating a second convolutional output based on a difference between the first frame and the second frame using one or more quantizers of the machine learning model, generating a third convolutional output associated with the second frame as a combination of the first convolutional output and the second convolutional output, and performing image processing based on the first convolutional output and the third convolutional output.
    Type: Application
    Filed: June 20, 2023
    Publication date: May 23, 2024
    Inventors: Davide ABATI, Amirhossein HABIBIAN, Markus NAGEL