Patents by Inventor Christopher Lott
Christopher Lott has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Patent number: 12688235Abstract: Certain aspects of the present disclosure provide techniques and apparatus for generating a response to a query input in a generative artificial intelligence model. An example method generally includes receiving a plurality of sets of tokens generated based on an input prompt and a first generative artificial intelligence model, each set of tokens in the plurality of sets of tokens corresponding to a candidate response to the input prompt; selecting, using a second generative artificial intelligence model and recursive adjustment of a target distribution associated with the received plurality of sets of tokens, a set of tokens from the plurality of sets of tokens; and outputting the selected set of tokens as a response to the input prompt.Type: GrantFiled: January 7, 2025Date of Patent: July 21, 2026Assignee: QUALCOMM IncorporatedInventors: Christopher Lott, Mingu Lee, Wonseok Jeon, Roland Memisevic
-
Publication number: 20260178326Abstract: A processing-in-memory (PIM) device implements block quantization techniques for matrix-vector operations. The PIM device performs matrix-vector operations between portions of a weight matrix and an input vector, and copies results to a register. A read operation retrieves the copied results while additional matrix-vector operations are performed in parallel. The device may apply scaling factors to the results using multipliers within the PIM device. In some implementations, the weight matrix includes data columns and scaling factor columns interspersed at regular intervals. The scaling factors may be applied to accumulated results using parallel multiplication operations. Disclosed techniques enable efficient implementation of block quantization for applications such as Large Language Models while managing computational resources within the PIM architecture.Type: ApplicationFiled: December 20, 2024Publication date: June 25, 2026Inventors: Subbarao Palacharla, Christopher Lott, Rui Cao
-
Publication number: 20260170324Abstract: Techniques and apparatus for generating a response to an input prompt using efficient self-speculative decoding in a generative artificial intelligence model. An example method generally includes receiving an input prompt for processing. A forecast embedding representing one or more forecasted tokens responsive to the input prompt is generated. Generally, the one or more forecasted tokens include tokens speculatively decoded by a generative artificial intelligence model based on generation of an initial response token in response to the input prompt. A bias parameter for the input prompt is determined. Generally, the bias parameter includes an embedding representation representing an error metric between the one or more forecasted tokens and an accepted set of tokens responsive to the input prompt. Using the generative artificial intelligence model, a response to the input prompt is generated based on the input prompt, the forecast embedding, and the bias parameter, and the generated response is output.Type: ApplicationFiled: September 3, 2025Publication date: June 18, 2026Inventors: Raghavv GOEL, Mingu LEE, Wonseok JEON, Mukul GAGRANI, Junyoung PARK, Christopher LOTT
-
Publication number: 20260161571Abstract: Certain aspects of the present disclosure provide techniques and apparatus for machine learning. In an example method, a memory associated with a generative machine learning model is accessed while data is processed using the generative machine learning model. The memory stores, for each token of a set of tokens for, or generated by, the generative machine learning model, a respective key tensor and a respective value tensor corresponding to the token. An average key tensor of the key tensors corresponding to the set of tokens is determined. A respective cosine similarity metric is generated for each key tensor, based on the key tensor and the average key tensor. At least one key tensor corresponding to at least one token of the set of tokens is evicted from the memory in response to determining that the respective cosine similarity metric for the at least one key tensor satisfies a condition.Type: ApplicationFiled: October 28, 2025Publication date: June 11, 2026Inventors: Junyoung PARK, Dalton James JONES, Matthew James MORSE, Raghavv GOEL, Mingu LEE, Christopher LOTT
-
Publication number: 20260119905Abstract: Certain aspects of the present disclosure provide techniques and apparatus for generating a response to a query input into a generative artificial intelligence model. The method generally includes generating, based on an input query and a first generative model, a plurality of sets of tokens, each set of tokens in the plurality of sets of tokens corresponding to a candidate response to the input query; outputting, to a second generative model, the plurality of sets of tokens for verification; receiving, from the second generative model, an indication of a selected set of tokens from the plurality of sets of tokens based on the input query and the plurality of sets of tokens; and outputting the selected set of tokens as a response to the input query.Type: ApplicationFiled: October 2, 2023Publication date: April 30, 2026Inventors: Christopher LOTT, Mingu LEE, Joseph Binamira SORIAGA, Jilei HOU
-
Publication number: 20260099683Abstract: Disclosed are systems, apparatuses, processes, and computer-readable media for model training. A device may process, using a linear layer, an embedding generated from a first output token and input features to generate first features, wherein the first output token is generated by a previous iteration of a token predictor and wherein the input features are generated by a previous iteration of a decoding layer. A device may process, using the decoding layer, the first features to generate second features having first dimensions. A device may process, using a down-projection layer, the second features to generate third features having second dimensions smaller than the first dimensions. A device may generate, using the token predictor and the third features, a second output token.Type: ApplicationFiled: February 11, 2025Publication date: April 9, 2026Inventors: Mingu LEE, Wonseok JEON, Junyoung PARK, Kanghoon YOON, Christopher LOTT
-
Publication number: 20260099673Abstract: Disclosed are systems, apparatuses, processes, and computer-readable media for model training. A device may process, using a linear layer, an embedding generated from a first output token and input features to generate first features, wherein the first output token is generated by a previous iteration of a token predictor and wherein the input features are generated by a previous iteration of a decoding layer. A device may process, using the decoding layer, the first features to generate second features having first dimensions. A device may process, using a down-projection layer, the second features to generate third features having second dimensions smaller than the first dimensions. A device may generate, using the token predictor and the third features, a second output token.Type: ApplicationFiled: February 11, 2025Publication date: April 9, 2026Inventors: Mingu LEE, Wonseok JEON, Junyoung PARK, Kanghoon YOON, Christopher LOTT
-
Publication number: 20260087374Abstract: Certain aspects provide techniques and apparatus for executing queries in a computing system using machine learning models. An example method generally includes receiving a plan to satisfy a request in the computing system and event log data associated with execution of the plan. The plan generally specifies a first plurality of function calls at a first level of granularity. Using a plan refinement machine learning model, a refined plan is generated when the event log data indicates that execution of the generated plan results in one or more execution errors and the one or more execution errors are solvable. Generally, the refined plan specifies a second plurality of function calls at a second level of granularity, the second level of granularity being finer than the first level of granularity.Type: ApplicationFiled: September 20, 2024Publication date: March 26, 2026Inventors: Amr Mamoun MARTINI, Arvind Vardarajan SANTHANAM, Swagarika Jaharlal GIRI, Christopher LOTT
-
Patent number: 12579063Abstract: Certain aspects of the present disclosure provide techniques and apparatus for machine learning. In an example method, an input prompt comprising a set of tokens is accessed as input to a generative machine learning model. A first key tensor and a first value tensor are generated for a first token of the set of tokens, and the first key tensor and the first value tensor are stored in a memory. A first retention score is generated, for the first token, based on the first key tensor, the first value tensor, and a second token of the set of tokens. The first key tensor and the first value tensor are evicted from the memory in response to determining that the first retention score is a lowest retention score of the memory.Type: GrantFiled: September 5, 2024Date of Patent: March 17, 2026Assignee: QUALCOMM IncorporatedInventors: Raghavv Goel, Mukul Gagrani, Junyoung Park, Dalton James Jones, Mingu Lee, Wonseok Jeon, Matthew James Morse, Matthew Harper Langston, Christopher Lott
-
Publication number: 20260065048Abstract: Certain aspects of the present disclosure provide techniques and apparatus for generating a response to a query input in a generative artificial intelligence model. An example method generally includes receiving an input prompt for processing; generating a set of forecasted parameters for the input prompt using a parameter prediction model; generating, using a generative artificial intelligence model, a response to the input prompt based on the input prompt and the set of forecasted parameters; and outputting the generated response.Type: ApplicationFiled: December 18, 2024Publication date: March 5, 2026Inventors: Mingu LEE, Raghavv GOEL, Wonseok JEON, Mukul GAGRANI, Junyoung PARK, Christopher LOTT
-
Publication number: 20260050766Abstract: Certain aspects of the present disclosure provide techniques and apparatus for efficient inferencing using a machine learning model. An example method generally includes receiving an input including a set of tokens for processing by a transformer neural network. The set of tokens for processing by the transformer neural network is partitioned into a first set of tokens and a second set of tokens. Using at least one state space model, at least one compressed token representing the first set of tokens is generated. An output token is generated, using the transformer neural network, based on the compressed token and the second set of tokens. A response to the input is generated based on the output token.Type: ApplicationFiled: January 7, 2025Publication date: February 19, 2026Inventors: Mukul GAGRANI, Junyoung PARK, Raghavv GOEL, Dalton James JONES, Wonseok JEON, Matthew James MORSE, Matthew Harper LANGSTON, Mingu LEE, Christopher LOTT
-
Publication number: 20260044745Abstract: Certain aspects of the present disclosure provide techniques and apparatus for machine learning. In an example method, a machine learning model comprising a plurality of layers, and a set of input data for the machine learning model, are accessed. A combination of hyperparameters for the machine learning model is selected based on the set of input data, comprising selecting, for each respective layer of the plurality of layers, a respective cache size based on the input data. The machine learning model is deployed according to the combination of hyperparameters.Type: ApplicationFiled: August 8, 2024Publication date: February 12, 2026Inventors: Dalton James JONES, Junyoung PARK, Matthew James MORSE, Raghavv GOEL, Mukul GAGRANI, Mingu LEE, Matthew Harper LANGSTON, Pierre-David LETOURNEAU, Christopher LOTT
-
Publication number: 20260044449Abstract: Certain aspects of the present disclosure provide techniques and apparatus for machine learning. In an example method, an input prompt comprising a set of tokens is accessed as input to a generative machine learning model. A first key tensor and a first value tensor are generated for a first token of the set of tokens, and the first key tensor and the first value tensor are stored in a memory. A first retention score is generated, for the first token, based on the first key tensor, the first value tensor, and a second token of the set of tokens. The first key tensor and the first value tensor are evicted from the memory in response to determining that the first retention score is a lowest retention score of the memory.Type: ApplicationFiled: October 21, 2025Publication date: February 12, 2026Inventors: Raghavv GOEL, Mukul GAGRANI, Junyoung PARK, Dalton James JONES, Mingu LEE, Wonseok JEON, Matthew James MORSE, Matthew Harper LANGSTON, Christopher LOTT
-
Publication number: 20260017192Abstract: Certain aspects of the present disclosure provide techniques and apparatus for machine learning. In an example method, an input prompt comprising a set of tokens is accessed as input to a generative machine learning model. A first key tensor and a first value tensor are generated for a first token of the set of tokens, and the first key tensor and the first value tensor are stored in a memory. A first retention score is generated, for the first token, based on the first key tensor, the first value tensor, and a second token of the set of tokens. The first key tensor and the first value tensor are evicted from the memory in response to determining that the first retention score is a lowest retention score of the memory.Type: ApplicationFiled: September 5, 2024Publication date: January 15, 2026Inventors: Raghavv GOEL, Mukul GAGRANI, Junyoung PARK, Dalton James JONES, Mingu LEE, Wonseok JEON, Matthew James MORSE, Matthew Harper LANGSTON, Christopher LOTT
-
Publication number: 20260017323Abstract: Techniques and apparatus for efficiently adapting a machine learning model to perform a variety of tasks using different adapters are provided. An example method generally includes receiving an input including a sequence of tokens associated with at least an input prompt into a neural network. The sequence of tokens is generated by a transformer block and a first set of adapters associated with the transformer block. A second set of adapters associated with the transformer block is loaded. An output of the transformer block is generated based on a key-value cache associated with the input and on weights associated with the transformer block. An output of the second set of adapters associated with the transformer block is generated based on the key-value cache associated with the input and on adapter weights associated with the second set of adapters.Type: ApplicationFiled: December 4, 2024Publication date: January 15, 2026Inventors: Amr Mamoun MARTINI, Arvind Vardarajan SANTHANAM, Christopher LOTT
-
Publication number: 20260017564Abstract: Techniques and apparatus for efficiently adapting a machine learning model to perform tasks using adapters are provided. An example method generally includes receiving an input for processing by a transformer block in a neural network. An output of the transformer block is generated based on the received input and weights associated with the transformer block. An output of an adapter associated with the transformer block is generated based on a copy of the received input and adapter weights associated with the adapter. Key-value data associated with the output of the transformer block and key-value data associated with a combination of the output of the transformer block and the output of the adapter are stored in a cache for subsequent inferencing rounds. A response to the input is generated based on the combination of the output of the transformer block and the output of the adapter.Type: ApplicationFiled: December 4, 2024Publication date: January 15, 2026Inventors: Amr Mamoun MARTINI, Arvind Vardarajan SANTHANAM, Christopher LOTT
-
Patent number: 12524405Abstract: Certain aspects provide techniques and apparatus for executing queries in a computing system using machine learning models. An example method generally includes receiving a plan to satisfy a request in the computing system and event log data associated with execution of the plan. The plan generally specifies a first plurality of actions to be performed by the computing system at a first level of granularity. Using a plan refinement machine learning model, a refined plan is generated when the event log data indicates that execution of the generated plan results in one or more execution errors and the one or more execution errors are solvable. Generally, the refined plan specifies a second plurality of actions to be performed by the computing system at a second level of granularity, the second level of granularity being finer than the first level of granularity.Type: GrantFiled: September 20, 2024Date of Patent: January 13, 2026Assignee: Qualcomm IncorporatedInventors: Amr Mamoun Martini, Arvind Vardarajan Santhanam, Christopher Lott
-
Patent number: 12493827Abstract: A method for optimizing the compilation of a machine learning model to be executed on target edge devices is provided. Compute nodes of a plurality of compute nodes are allocated to a compiler optimization process for a compiler of said machine learning model. The machine learning model has a compute graph representation having nodes that are kernel operators necessary to execute the machine learning model and edges that connect said kernel operators to define precedence constraints. A round of optimization is scheduled for the process amongst the allocated compute nodes. At each allocated compute node a sequencing and scheduling solution is applied per round to obtain a performance metric for the machine learning model. From each compute node the performance metric is received and a solution that has the best performance metric is identified and implemented for execution of the machine learning model on the target edge devices.Type: GrantFiled: November 17, 2022Date of Patent: December 9, 2025Assignee: Qualcomm IncorporatedInventors: Weiliang Zeng, Christopher Lott, Edward Teague, Yang Yang, Joseph Binamira Soriaga
-
Publication number: 20250356184Abstract: Certain aspects of the present disclosure provide techniques and apparatus for improved machine learning. In an example method, a sequence of tokens is accessed as input to an attention operation. For a first token, an attention output is generated based on a window of tokens relative to the first token, comprising generating a first positional embedding for an influential token, generating a second positional embedding for the first token, and generating the attention output based on the first and second positional embeddings. For a second token, an attention output is generated based on a window of tokens relative to the second token, where the second window of tokens includes the first token, comprising generating a third positional embedding for the influential token, generating a fourth positional embedding for the second token, and generating the attention output based on the second, third, and fourth positional embeddings.Type: ApplicationFiled: May 17, 2024Publication date: November 20, 2025Inventors: Junyoung PARK, Mukul GAGRANI, Raghavv GOEL, Wonseok JEON, Mingu LEE, Christopher LOTT
-
Patent number: 12450486Abstract: A method performed by a computing device includes determining a partition for depth-first processing by a multi-layer artificial neural network (ANN) of the computing device. The computing device comprising a processor, on-chip memory, and off-chip memory. The first partition determined based on an amount of on-chip memory used by the first partition, an available amount of on-chip memory, and a size of a write back to the off-chip memory. The method also includes processing, at the device via the multi-layer ANN, an input, using the depth-first processing in accordance with the partition.Type: GrantFiled: December 14, 2020Date of Patent: October 21, 2025Assignee: QUALCOMM IncorporatedInventors: Piero Zappi, Jin Won Lee, Christopher Lott, Rexford Alan Hill