Patents by Inventor Ron Diamant

Ron Diamant has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Patent number: 12682232
    Abstract: A technique to compute statistics of data elements include serially inputting the data elements into a compute channel. The compute channel can generate a first running mean and a first running variance associated with data elements of the vector having odd sequence indices, and a second running mean and a second running variance associated with data elements of the vector having even sequence indices. Subsequent to serially inputting data elements into the compute channel, the first running mean and the second running mean are aggregated to generate a mean associated with the data elements of the vector, and the first running variance and the second running variance are aggregated to generate a variance associated with the data elements of the vector.
    Type: Grant
    Filed: September 26, 2022
    Date of Patent: July 14, 2026
    Assignee: Amazon Technologies, Inc.
    Inventors: Paul Gilbert Meyer, Sundeep Amirineni, Ron Diamant
  • Patent number: 12669980
    Abstract: A technique for matching the throughput between writing into and reading from a memory can include receiving, in parallel, computational results in a high precision format for storing into the memory at a first frequency, and storing the computational results in the memory. The technique may further include rounding the computational results using round-to-the-nearest-even or stochastic rounding to down-convert the computational results from the high precision format to a low precision format in parallel, and outputting the computational results in the low precision format in parallel from the memory at a second frequency.
    Type: Grant
    Filed: June 16, 2022
    Date of Patent: June 30, 2026
    Assignee: Amazon Technologies, Inc.
    Inventors: Kun Xu, Sundeep Amirineni, Paul Gilbert Meyer, Ron Diamant
  • Patent number: 12664430
    Abstract: A neural network processor is configured to execute instructions in parallel on different computing engines (CEs) to perform convolution operations on an input dataset. The input dataset is divided into overlapping chunks, including a first chunk and a second chunk. Each CE processes a last portion of a respective chunk to compute respective shared states and receives additional shared states for processing. The first chunk is the respective chunk for a first CE. The second chunk is the respective chunk for a second CE. The additional shared states received by the second CE are the respective shared states computed by the first CE. The second CE receives the respective shared states computed by the first CE as a substitute for intermediate states that would otherwise have been computed by the second CE using a first portion of the second chunk.
    Type: Grant
    Filed: April 30, 2024
    Date of Patent: June 23, 2026
    Assignee: Amazon Technologies, Inc.
    Inventors: Thiam Khean Hah, Randy Renfu Huang, Richard John Heaton, Ron Diamant, Vignesh Vivekraja
  • Patent number: 12645605
    Abstract: A computer-implemented method includes generating or receiving instruction code for executing by a computing device to implement a neural network model, where the instruction code includes a plurality of direct memory access (DMA) instructions for data transferring between a local memory of an accelerator of the computing device and a system memory of the computing device; modifying the instruction code to arrange sources or destinations of a group of DMA instructions of the plurality of DMA instructions into a contiguous block in the local memory; and replacing the group of DMA instructions with a single DMA instruction, wherein a source address or a destination address of the single DMA instruction is the contiguous block of the local memory.
    Type: Grant
    Filed: September 30, 2021
    Date of Patent: June 2, 2026
    Assignee: Amazon Technologies, Inc.
    Inventors: Ron Diamant, Yunxuan Yu, Taylor Goodhart, Robert Geva
  • Patent number: 12632401
    Abstract: A direct memory access (DMA) engine may receive a first indication that memory descriptors for a DMA operation are ready to be fetched from a memory. The DMA engine may prefetch the memory descriptors from the memory without waiting for memory locations specified in the memory descriptors to be ready for access. The DMA engine may receive a second indication that the memory locations specified in the memory descriptors are ready to be accessed. The DMA engine may execute the DMA operation based on the memory descriptors upon receiving the second indication.
    Type: Grant
    Filed: September 12, 2023
    Date of Patent: May 19, 2026
    Assignee: Amazon Technologies, Inc.
    Inventors: Kun Xu, Ron Diamant, Ilya Minkin, Raymond S. Whiteside
  • Patent number: 12632714
    Abstract: A processing engine array is provided with an interconnect mode of operation to use the array as an interconnect to move data elements to different locations in memory such as to perform a matrix transpose operation. In this interconnect mode of operation, although computations are still being performed in the array, the computations are not carried out to modify or change the values of the data elements, but are instead carried out to rearrange the data elements in memory. As such, the computations carried out in the interconnect mode of operation can deviate from the expected behavior of floating-point calculations. A mode selection signal can be used to provide the proper outputs of the processing elements of the array depending on the mode of operation.
    Type: Grant
    Filed: November 2, 2021
    Date of Patent: May 19, 2026
    Assignee: Amazon Technologies, Inc.
    Inventors: Sundeep Amirineni, Paul Gilbert Meyer, Ron Diamant, Qingrui Liu
  • Patent number: 12632693
    Abstract: A technique for packing matrix multiplications for concurrent execution in an integrated circuit device may include obtaining a description of a neural network model, and generating an intermediate representation of the neural network model. Matrix multiplication instructions in the intermediate representation of the neural network model can then be vectorized for concurrent execution on an integrated circuit device, and machine instructions can be generated for the integrated circuit device based on the vectorized matrix multiplication instructions.
    Type: Grant
    Filed: March 30, 2022
    Date of Patent: May 19, 2026
    Assignee: Amazon Technologies, Inc.
    Inventors: Jiading Gai, Tobias Joseph Kastulus Edler von Koch, Robert Geva, Paul Gilbert Meyer, Donald John Kretsch, Ron Diamant
  • Patent number: 12632409
    Abstract: Techniques for performing collective compute operations are described. A collective compute operation can be performed in a logical ring of processing ranks formed by a set of rank groups that each contain a number of processing ranks including a primary rank and one or more secondary ranks. At each rank group, a primary rank receives an incoming data slice via an intranode interconnect from a previous primary rank at multiple hops away on the logical ring. A data transfer is performed between the primary rank and each secondary rank of the rank group. An outgoing data slice is then transferred from the primary rank of the rank group to the next primary rank at multiple hops away on the logical ring.
    Type: Grant
    Filed: June 27, 2024
    Date of Patent: May 19, 2026
    Assignee: Amazon Technologies, Inc.
    Inventors: Yongseok Koh, Se Wang Oh, Zhaoqi Zhu, Ron Diamant
  • Patent number: 12632700
    Abstract: Systems and methods for performing improper input data detection are described. In one example, a system comprises: hardware circuits configured to receive input data and to perform computations of a neural network based on the input data to generate computation outputs; and an improper input detection circuit configured to: determine a relationship between the computation outputs of the hardware circuits and reference outputs; determine that the input data are improper based on the relationship; and perform an action based on determining that the input data are improper.
    Type: Grant
    Filed: May 5, 2023
    Date of Patent: May 19, 2026
    Assignee: Amazon Technologies, Inc.
    Inventors: Randy Renfu Huang, Richard John Heaton, Andrea Olgiati, Ron Diamant
  • Patent number: 12614110
    Abstract: A computer-implemented technique for optimizing self-attention masks is described. At compile time, a machine learning graph of an artificial intelligence model is analyzed. The machine learning graph includes a set of operators. Analysis includes identifying one or more mask operators and determining what fields of input tensors are masked. Optimizations at compile-time are used to eliminate instructions during training of the artificial intelligence model.
    Type: Grant
    Filed: September 28, 2022
    Date of Patent: April 28, 2026
    Assignee: Amazon Technologies, Inc.
    Inventors: Hongbin Zheng, Yunxuan Yu, Ron Diamant
  • Patent number: 12608335
    Abstract: Disclosed herein are techniques for obtaining weights for neural network computations. In one embodiment, an integrated circuit may include memory configured to store a first weight and a second weight; a row of processing elements comprising a first processing element and a second processing element, the first processing element comprising a first weight register, the second processing element comprising a second weight register, both of the first weight register and the second weight register being controllable by a weight load signal; and a controller configured to: provide the first weight from the memory to the row of processing elements; set the weight load signal to enable the first weight to propagate through the row to reach the first processing element; and set the weight load signal to store the first weight at the first weight register and a flush value at the second weight register.
    Type: Grant
    Filed: January 28, 2022
    Date of Patent: April 21, 2026
    Assignee: Amazon Technologies, Inc.
    Inventors: Dana Michelle Vantrease, Ron Diamant, Sundeep Amirineni
  • Patent number: 12596920
    Abstract: In one example, an apparatus comprises: a direct memory access (DMA) descriptor queue that stores DMA descriptors, each DMA descriptor including an indirect address; an address translation table that stores an address mapping between indirect addresses and physical addresses; and a DMA engine configured to: fetch a DMA descriptor from the DMA descriptor queue to the address translation table to translate a first indirect address of the DMA descriptor to a first physical address based on the address mapping, and perform a DMA operation based on executing the DMA descriptor to transfer data to or from the first physical address.
    Type: Grant
    Filed: May 4, 2023
    Date of Patent: April 7, 2026
    Assignee: Amazon Technologies, Inc.
    Inventors: Ilya Minkin, Ron Diamant, Kun Xu
  • Patent number: 12561260
    Abstract: Techniques to perform transpose operations in a crossbar circuit may include receiving, for a transpose write operation to transpose a data array, a set of write transactions from one or more data sources. Each write transaction can include an opcode and a row size of the data array being transposed. Data is written into the transpose memory in response to write transactions having an opcode indicating that the write transaction contains row data of the data array being transposed. When it is determined that the data array has been written into the transpose memory, the data array is outputted from the transposed memory in a transposed format to write the data array to a target memory.
    Type: Grant
    Filed: March 31, 2023
    Date of Patent: February 24, 2026
    Assignee: Amazon Technologies, Inc.
    Inventors: Patricio Kaplan, Ron Diamant
  • Publication number: 20260023529
    Abstract: Systems and methods are provided to perform multiply-accumulate operations of reduced precision numbers in a systolic array. Each row of the systolic array can receive reduced inputs from a respective reducer. The reducer can receive a particular input and generate multiple reduced inputs from the input. The reduced inputs can include reduced input data elements and/or a reduced weights. The systolic array may lack support for inputs with a first bit-length and the reducers may reduce the bit-length of a given input from the first bit-length to a second shorter bit-length and provide multiple reduced inputs with second shorter bit-length to the array. The systolic array may perform multiply-accumulate operations on each unique combination of the multiple reduced input data elements and the reduced weights to generate multiple partial outputs. The systolic array may sum the partial outputs to generate the output.
    Type: Application
    Filed: August 5, 2025
    Publication date: January 22, 2026
    Inventors: Paul Gilbert Meyer, Thomas A. Volpe, Ron Diamant, Joshua Wayne Bowman, Nishith Desai, Thomas Elmer
  • Patent number: 12530178
    Abstract: A technique for arranging matrix multiplications for concurrent execution in an integrated circuit device may include obtaining a representation of a data dependency graph of a neural network model. The data dependency graph may include having an accumulation group (AG) pack of accumulation groups (AGs), in which each of the AGs has one or more matrix multipartition instructions. A representation of a memory location base partition constraint graph of the AG pack can be generated, and an AG row group constraint graph can be generated based on the memory location base partition constraint graph. The AGs of the AG pack can then be assigned to tiles in an integrated circuit device based on the AG row group constraint graph.
    Type: Grant
    Filed: March 30, 2022
    Date of Patent: January 20, 2026
    Assignee: Amazon Technologies, Inc.
    Inventors: Jiading Gai, Tobias Joseph Kastulus Edler von Koch, Robert Geva, Paul Gilbert Meyer, Donald John Kretsch, Ron Diamant
  • Patent number: 12524360
    Abstract: Techniques to reduce direct memory access (DMA) overhead may include retrieving an address translation descriptor from a descriptor queue of a DMA engine, and updating an address translation table in the DMA engine with address translation information obtained from the location indicated by the address translation descriptor. A set of memory descriptors is then obtained from the descriptor queue. The set of memory descriptors can be processed by determining that the addresses in the set of memory descriptors are to be translated using the address translation table, and performing memory access operations by using the address translation table to translate the addresses in the set of memory descriptors.
    Type: Grant
    Filed: May 31, 2022
    Date of Patent: January 13, 2026
    Assignee: Amazon Technologies, Inc.
    Inventors: Kun Xu, Ilya Minkin, Ron Diamant
  • Patent number: 12518167
    Abstract: In one example, a method comprises: performing backward propagation computations for a second layer of a neural network to generate second weight gradients; splitting the second weight gradients into a plurality of subsets each associated with a second exchange operation over a computer network, a number of the second weight gradients included in the each subset being based on at least one of first characteristics of the computer network or second characteristics of the neural network; performing backward propagation computations for a first layer of the neural network to generate first weight gradients in parallel with at least one of the second exchange operations; performing a first exchange operation to exchange the first weight gradients after the at least one of the second exchange operations completes; and after the first exchange operation completes, perform the remaining second exchange operations to exchange the remaining subsets of the second weight gradients.
    Type: Grant
    Filed: September 30, 2019
    Date of Patent: January 6, 2026
    Assignee: Amazon Technologies, Inc.
    Inventors: Vignesh Vivekraja, Thiam Khean Hah, Randy Renfu Huang, Richard John Heaton, Ron Diamant
  • Patent number: 12511544
    Abstract: Systems and methods are provided to improve the memory throughput for storing and reading intermediate data computed by layers of a neural network during a training process. A compression operation can be performed by removing the zeros from the intermediate data and storing locations of the zeros before storing the intermediate data in the memory for a forward pass of the training process. The compressed data can be read from the memory for a backward pass of the training process and de-compressed by inserting zeros based on the stored locations. Additionally, a transpose operation can be performed before compression as a first atomic operation, or after de-compression as a second atomic operation.
    Type: Grant
    Filed: September 29, 2020
    Date of Patent: December 30, 2025
    Assignee: Amazon Technologies, Inc.
    Inventors: Kun Xu, Ron Diamant
  • Patent number: 12505339
    Abstract: An opportunistic approach is described to accelerate certain memory copy operations to be performed by a computing engine by executing instructions. An instruction for a memory copy operation can be identified that has a first data type with a smaller number of bits per data element than supported by the computing engine to perform the copy operation. The instruction can be replaced with another instruction that has a second data type with a higher number of bits per data element based on the alignment of the memory addresses for the copy operation and the total number of data elements to be copied. The second data type may not only accelerate the copy operation but also provide better utilization of the underlying hardware of the computing engine.
    Type: Grant
    Filed: September 7, 2021
    Date of Patent: December 23, 2025
    Assignee: Amazon Technologies, Inc.
    Inventors: Ron Diamant, Xin Tong
  • Patent number: 12499353
    Abstract: Systems and methods for performing hardware approximation of functions are provided. In one example, a system comprises a controller, a plurality of multiplexors, configurable arithmetic circuits, and a mapping table that stores a set of function parameters. According to a mode of operations, the controller may configure the plurality of multiplexors to forward the set of function parameters or a subset of the function parameters to the arithmetic circuits to compute an approximation result. In a case where the subset of the function parameters is forwarded to the arithmetic circuits, the controller may configure the arithmetic circuits to perform post-processing, such as quantization, of the approximation result.
    Type: Grant
    Filed: December 12, 2018
    Date of Patent: December 16, 2025
    Assignee: Amazon Technologies, Inc.
    Inventors: Ron Diamant, Sundeep Amirineni, Mohammad El-Shabani, Kenneth Wayne Patton