Patents by Inventor Aayush ANKIT

Aayush ANKIT has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Patent number: 12639253
    Abstract: Methods and systems for allreduce optimization on small tensors. In an embodiment, the present invention provides a system for hierarchical data aggregation includes a first module with interconnected chiplets linked via D2D links, enabling intra-module data sharing and reduce-scatter operations. A second module, also with interconnected chiplets and D2D links, performs a similar intra-module reduce-scatter operation. There are other embodiments as well.
    Type: Grant
    Filed: December 18, 2024
    Date of Patent: May 26, 2026
    Assignee: d-MATRIX CORPORATION
    Inventors: Aayush Ankit, Sudeep Bhoja
  • Publication number: 20250342134
    Abstract: An AI accelerator apparatus using in-memory compute chiplet devices. The apparatus includes a first semiconductor substrate having a plurality of chiplets, each of which includes a plurality of tiles. Each tile includes a plurality of slices, a central processing unit (CPU), and a hardware dispatch device. Each slice can include a digital in-memory compute (DIMC) device configured to perform high throughput computations. In particular, the DIMC device can be configured to accelerate the computations of attention functions for transformer-based models (a.k.a. transformers) applied to machine learning applications. The chiplets are in a full mesh connectivity configuration such that at least one of the die-to-die (D2D) interconnects of each chiplet is coupled to one of the D2D interconnects of each other chiplet using a non-diagonal link. The chiplets can also include other interfaces to facilitate communication between the chiplets, memory and a server or host system.
    Type: Application
    Filed: July 14, 2025
    Publication date: November 6, 2025
    Inventors: Jayaprakash BALACHANDRAN, Aayush ANKIT, Irene QUEK, Sudeep BHOJA
  • Publication number: 20250315666
    Abstract: A server system with AI accelerator apparatuses using in-memory compute chiplet devices. The system includes a plurality of multiprocessors each having at least a first server central processing unit (CPU) and a second server CPU, both of which are coupled to a plurality of switch devices. Each switch device is coupled to a plurality of AI accelerator apparatuses. The apparatus includes one or more chiplets, each of which includes a plurality of tiles. Each tile includes a plurality of slices, a CPU, and a hardware dispatch device. Each slice can include a digital in-memory compute (DIMC) device configured to perform high throughput computations. In particular, the DIMC device can be configured to accelerate the computations of attention functions for transformer-based models (a.k.a. transformers) applied to machine learning applications. A single input multiple data (SIMD) device configured to further process the DIMC output and compute softmax functions for the attention functions.
    Type: Application
    Filed: June 17, 2025
    Publication date: October 9, 2025
    Inventors: Jayaprakash Balachandran, Akhil Arunkumar, Aayush Ankit, Nithesh Kurella, Sudeep Bhoja
  • Patent number: 12386774
    Abstract: An AI accelerator apparatus using in-memory compute chiplet devices. The apparatus includes a first semiconductor substrate having a plurality of chiplets, each of which includes a plurality of tiles. Each tile includes a plurality of slices, a central processing unit (CPU), and a hardware dispatch device. Each slice can include a digital in-memory compute (DIMC) device configured to perform high throughput computations. In particular, the DIMC device can be configured to accelerate the computations of attention functions for transformer-based models (a.k.a. transformers) applied to machine learning applications. The chiplets are in a full mesh connectivity configuration such that at least one of the die-to-die (D2D) interconnects of each chiplet is coupled to one of the D2D interconnects of each other chiplet using a non-diagonal link. The chiplets can also include other interfaces to facilitate communication between the chiplets, memory and a server or host system.
    Type: Grant
    Filed: October 13, 2023
    Date of Patent: August 12, 2025
    Assignee: d-MATRIX CORPORATION
    Inventors: Jayaprakash Balachandran, Aayush Ankit, Irene Quek, Sudeep Bhoja
  • Patent number: 12353985
    Abstract: A server system with AI accelerator apparatuses using in-memory compute chiplet devices. The system includes a plurality of multiprocessors each having at least a first server central processing unit (CPU) and a second server CPU, both of which are coupled to a plurality of switch devices. Each switch device is coupled to a plurality of AI accelerator apparatuses. The apparatus includes one or more chiplets, each of which includes a plurality of tiles. Each tile includes a plurality of slices, a CPU, and a hardware dispatch device. Each slice can include a digital in-memory compute (DIMC) device configured to perform high throughput computations. In particular, the DIMC device can be configured to accelerate the computations of attention functions for transformer-based models (a.k.a. transformers) applied to machine learning applications. A single input multiple data (SIMD) device configured to further process the DIMC output and compute softmax functions for the attention functions.
    Type: Grant
    Filed: October 13, 2023
    Date of Patent: July 8, 2025
    Assignee: d-MATRIX CORPORATION
    Inventors: Jayaprakash Balachandran, Akhil Arunkumar, Aayush Ankit, Nithesh Kurella, Sudeep Bhoja
  • Publication number: 20250123984
    Abstract: An AI accelerator apparatus using in-memory compute chiplet devices. The apparatus includes a first semiconductor substrate having a plurality of chiplets, each of which includes a plurality of tiles. Each tile includes a plurality of slices, a central processing unit (CPU), and a hardware dispatch device. Each slice can include a digital in-memory compute (DIMC) device configured to perform high throughput computations. In particular, the DIMC device can be configured to accelerate the computations of attention functions for transformer-based models (a.k.a. transformers) applied to machine learning applications. The chiplets are in a full mesh connectivity configuration such that at least one of the die-to-die (D2D) interconnects of each chiplet is coupled to one of the D2D interconnects of each other chiplet using a non-diagonal link. The chiplets can also include other interfaces to facilitate communication between the chiplets, memory and a server or host system.
    Type: Application
    Filed: October 13, 2023
    Publication date: April 17, 2025
    Inventors: Jayaprakash BALACHANDRAN, Aayush Ankit, Irene Quek, Sudeep Bhoja
  • Publication number: 20250110885
    Abstract: A transformer compute apparatus and method of operation therefor. The apparatus receives matrix inputs in a first format and generates projection tokens from these inputs. Among others, the apparatus includes a first cache device configured for processing first projection tokens and a second cache device configured for processing second projection tokens. The first cache device stores the first projection tokens in a first cache region and stores these tokens converted to a second format in a second cache region. The second cache device stores the second projection tokens converted to the second format in a first cache region and stores the converted second projection tokens after being transposed. Then, a compute device performs various matrix computations with the converted first projection tokens and transposed second projection tokens. Re-processing data and expensive padding and de-padding operations for transposed storage and byte alignment can be avoided using this caching process.
    Type: Application
    Filed: November 18, 2024
    Publication date: April 3, 2025
    Inventors: Akhil ARUNKUMAR, Satyam Srivastava, Aayush Ankit
  • Patent number: 12182028
    Abstract: A transformer compute apparatus and method of operation therefor. The apparatus receives matrix inputs in a first format and generates projection tokens from these inputs. Among others, the apparatus includes a first cache device configured for processing first projection tokens and a second cache device configured for processing second projection tokens. The first cache device stores the first projection tokens in a first cache region and stores these tokens converted to a second format in a second cache region. The second cache device stores the second projection tokens converted to the second format in a first cache region and stores the converted second projection tokens after being transposed. Then, a compute device performs various matrix computations with the converted first projection tokens and transposed second projection tokens. Re-processing data and expensive padding and de-padding operations for transposed storage and byte alignment can be avoided using this caching process.
    Type: Grant
    Filed: September 28, 2023
    Date of Patent: December 31, 2024
    Assignee: d-MATRIX CORPORATION
    Inventors: Akhil Arunkumar, Satyam Srivastava, Aayush Ankit
  • Publication number: 20240303117
    Abstract: Operations of a workload are assigned to physical resources of a physical device array. The workload includes a graph of operations to be performed on a physical device array. The graph of operations is partitioned into subgraphs. Partitioning includes at least minimizing the quantity of subgraphs and maximizing resource utilization per subgraph. A logical mapping of the subgraph to logical processing engine (PE) units is generated using features of the subgraph and tiling factors of the logical PE units. The logical mapping is assigned to physical PE units of the physical device array at least by minimizing network traffic across the physical PE units. The operations of the subgraph are performed using the physical PE units to which the logical mapping is assigned. This process enhances the computational efficiency of the array when executing the workload.
    Type: Application
    Filed: February 24, 2023
    Publication date: September 12, 2024
    Inventors: Fanny NINA PARAVECINO, Michael Eric DAVIES, Abhishek Dilip KULKARNI, Md Aamir RAIHAN, Ankit MORE, Aayush ANKIT, Torsten HOEFLER, Douglas Christopher BURGER
  • Publication number: 20240090181
    Abstract: An immersion cooling server system with AI accelerator apparatuses using in-memory compute chiplet devices. This system includes one or more immersion tanks with heat transfer fluid and configured with at least a condenser device. A plurality of AI accelerator servers is immersed in the heat transfer fluid in a bottom portion of the tanks and is configured to process transformer workloads while cooled by the immersion cooling configuration. Each of the servers includes a plurality of multiprocessors each having at least a first server central processing unit (CPU) and a second server CPU, both of which are coupled to a plurality of switch devices. Each switch device is coupled to a plurality of AI accelerator apparatuses. The apparatus includes one or more chiplets, each of which includes a plurality of digital in-memory compute (DIMC) devices configured to perform high throughput matrix computations for transformer based models.
    Type: Application
    Filed: November 16, 2023
    Publication date: March 14, 2024
    Inventors: Jayaprakash BALACHANDRAN, Akhil ARUNKUMAR, Aayush ANKIT, Nithesh Kurella, Sudeep Bhoja
  • Patent number: 11929141
    Abstract: Sparsity-aware reconfiguration compute-in-memory (CIM) static random access memory (SRAM) systems are disclosed. In one aspect, a reconfigurable precision succession approximation register (SAR) analog-to-digital converter (ADC) that has the ability to form (n+m) bit precision using n-bit and m-bit sub-ADCs is provided. By controlling which sub-ADCs are used based on data sparsity, precision may be maintained as needed while providing a more energy efficient design.
    Type: Grant
    Filed: December 8, 2021
    Date of Patent: March 12, 2024
    Assignee: Purdue Research Foundation
    Inventors: Kaushik Roy, Amogh Agrawal, Mustafa Fayez Ahmed Ali, Indranil Chakraborty, Aayush Ankit, Utkarsh Saxena
  • Publication number: 20240037379
    Abstract: A server system with AI accelerator apparatuses using in-memory compute chiplet devices. The system includes a plurality of multiprocessors each having at least a first server central processing unit (CPU) and a second server CPU, both of which are coupled to a plurality of switch devices. Each switch device is coupled to a plurality of AI accelerator apparatuses. The apparatus includes one or more chiplets, each of which includes a plurality of tiles. Each tile includes a plurality of slices, a CPU, and a hardware dispatch device. Each slice can include a digital in-memory compute (DIMC) device configured to perform high throughput computations. In particular, the DIMC device can be configured to accelerate the computations of attention functions for transformer-based models (a.k.a. transformers) applied to machine learning applications. A single input multiple data (SIMD) device configured to further process the DIMC output and compute softmax functions for the attention functions.
    Type: Application
    Filed: October 13, 2023
    Publication date: February 1, 2024
    Inventors: Jayaprakash BALACHANDRAN, Akhil ARUNKUMAR, Aayush ANKIT, Nithesh Kurella, Sudeep Bhoja
  • Publication number: 20230178125
    Abstract: Sparsity-aware reconfiguration compute-in-memory (CIM) static random access memory (SRAM) systems are disclosed. In one aspect, a reconfigurable precision succession approximation register (SAR) analog-to-digital converter (ADC) that has the ability to form (n+m) bit precision using n-bit and m-bit sub-ADCs is provided. By controlling which sub-ADCs are used based on data sparsity, precision may be maintained as needed while providing a more energy efficient design.
    Type: Application
    Filed: December 8, 2021
    Publication date: June 8, 2023
    Inventors: Kaushik Roy, Amogh Agrawal, Mustafa Fayez Ahmed Ali, Indranil Chakraborty, Aayush Ankit, Utkarsh Saxena
  • Patent number: 11120605
    Abstract: A texture filtering unit and a method are disclosed that provide multiple variants of an approximate trilinear filtering operation. A texture sampling and filtering unit may be configured to determine a level-of-detail (LOD) value for a sample point in texture space, and select, based on the LOD value, a fine mip-level and a coarse mip-level from the mip-map. The closer of the two selected mip-levels to the sample point is determined, and farther of the two selected mip-levels from the sample point is determined. A first quad of texels in the closer mip-level and a second quad of texels in the farther mip-level are then determined. A total of five or fewer texels are selected from the first quad of texels and from the second quad of texels. A filtered value for the sample point is determined based on an approximate trilinear filtering operation on the selected texels.
    Type: Grant
    Filed: March 24, 2020
    Date of Patent: September 14, 2021
    Inventors: Anders M. Kugler, Aayush Ankit, Wilson Wai Lun Fung
  • Publication number: 20210217220
    Abstract: A texture filtering unit and a method are disclosed that provide multiple variants of an approximate trilinear filtering operation. A texture sampling and filtering unit may be configured to determine a level-of-detail (LOD) value for a sample point in texture space, and select, based on the LOD value, a fine mip-level and a coarse mip-level from the mip-map. The closer of the two selected mip-levels to the sample point is determined, and farther of the two selected mip-levels from the sample point is determined. A first quad of texels in the closer mip-level and a second quad of texels in the farther mip-level are then determined. A total of five or fewer texels are selected from the first quad of texels and from the second quad of texels. A filtered value for the sample point is determined based on an approximate trilinear filtering operation on the selected texels.
    Type: Application
    Filed: March 24, 2020
    Publication date: July 15, 2021
    Inventors: Anders M. KUGLER, Aayush ANKIT, Wilson Wai Lun FUNG