Patents by Inventor Aayush ANKIT
Aayush ANKIT has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Patent number: 12639253Abstract: Methods and systems for allreduce optimization on small tensors. In an embodiment, the present invention provides a system for hierarchical data aggregation includes a first module with interconnected chiplets linked via D2D links, enabling intra-module data sharing and reduce-scatter operations. A second module, also with interconnected chiplets and D2D links, performs a similar intra-module reduce-scatter operation. There are other embodiments as well.Type: GrantFiled: December 18, 2024Date of Patent: May 26, 2026Assignee: d-MATRIX CORPORATIONInventors: Aayush Ankit, Sudeep Bhoja
-
Publication number: 20250342134Abstract: An AI accelerator apparatus using in-memory compute chiplet devices. The apparatus includes a first semiconductor substrate having a plurality of chiplets, each of which includes a plurality of tiles. Each tile includes a plurality of slices, a central processing unit (CPU), and a hardware dispatch device. Each slice can include a digital in-memory compute (DIMC) device configured to perform high throughput computations. In particular, the DIMC device can be configured to accelerate the computations of attention functions for transformer-based models (a.k.a. transformers) applied to machine learning applications. The chiplets are in a full mesh connectivity configuration such that at least one of the die-to-die (D2D) interconnects of each chiplet is coupled to one of the D2D interconnects of each other chiplet using a non-diagonal link. The chiplets can also include other interfaces to facilitate communication between the chiplets, memory and a server or host system.Type: ApplicationFiled: July 14, 2025Publication date: November 6, 2025Inventors: Jayaprakash BALACHANDRAN, Aayush ANKIT, Irene QUEK, Sudeep BHOJA
-
Publication number: 20250315666Abstract: A server system with AI accelerator apparatuses using in-memory compute chiplet devices. The system includes a plurality of multiprocessors each having at least a first server central processing unit (CPU) and a second server CPU, both of which are coupled to a plurality of switch devices. Each switch device is coupled to a plurality of AI accelerator apparatuses. The apparatus includes one or more chiplets, each of which includes a plurality of tiles. Each tile includes a plurality of slices, a CPU, and a hardware dispatch device. Each slice can include a digital in-memory compute (DIMC) device configured to perform high throughput computations. In particular, the DIMC device can be configured to accelerate the computations of attention functions for transformer-based models (a.k.a. transformers) applied to machine learning applications. A single input multiple data (SIMD) device configured to further process the DIMC output and compute softmax functions for the attention functions.Type: ApplicationFiled: June 17, 2025Publication date: October 9, 2025Inventors: Jayaprakash Balachandran, Akhil Arunkumar, Aayush Ankit, Nithesh Kurella, Sudeep Bhoja
-
Patent number: 12386774Abstract: An AI accelerator apparatus using in-memory compute chiplet devices. The apparatus includes a first semiconductor substrate having a plurality of chiplets, each of which includes a plurality of tiles. Each tile includes a plurality of slices, a central processing unit (CPU), and a hardware dispatch device. Each slice can include a digital in-memory compute (DIMC) device configured to perform high throughput computations. In particular, the DIMC device can be configured to accelerate the computations of attention functions for transformer-based models (a.k.a. transformers) applied to machine learning applications. The chiplets are in a full mesh connectivity configuration such that at least one of the die-to-die (D2D) interconnects of each chiplet is coupled to one of the D2D interconnects of each other chiplet using a non-diagonal link. The chiplets can also include other interfaces to facilitate communication between the chiplets, memory and a server or host system.Type: GrantFiled: October 13, 2023Date of Patent: August 12, 2025Assignee: d-MATRIX CORPORATIONInventors: Jayaprakash Balachandran, Aayush Ankit, Irene Quek, Sudeep Bhoja
-
Patent number: 12353985Abstract: A server system with AI accelerator apparatuses using in-memory compute chiplet devices. The system includes a plurality of multiprocessors each having at least a first server central processing unit (CPU) and a second server CPU, both of which are coupled to a plurality of switch devices. Each switch device is coupled to a plurality of AI accelerator apparatuses. The apparatus includes one or more chiplets, each of which includes a plurality of tiles. Each tile includes a plurality of slices, a CPU, and a hardware dispatch device. Each slice can include a digital in-memory compute (DIMC) device configured to perform high throughput computations. In particular, the DIMC device can be configured to accelerate the computations of attention functions for transformer-based models (a.k.a. transformers) applied to machine learning applications. A single input multiple data (SIMD) device configured to further process the DIMC output and compute softmax functions for the attention functions.Type: GrantFiled: October 13, 2023Date of Patent: July 8, 2025Assignee: d-MATRIX CORPORATIONInventors: Jayaprakash Balachandran, Akhil Arunkumar, Aayush Ankit, Nithesh Kurella, Sudeep Bhoja
-
Publication number: 20250123984Abstract: An AI accelerator apparatus using in-memory compute chiplet devices. The apparatus includes a first semiconductor substrate having a plurality of chiplets, each of which includes a plurality of tiles. Each tile includes a plurality of slices, a central processing unit (CPU), and a hardware dispatch device. Each slice can include a digital in-memory compute (DIMC) device configured to perform high throughput computations. In particular, the DIMC device can be configured to accelerate the computations of attention functions for transformer-based models (a.k.a. transformers) applied to machine learning applications. The chiplets are in a full mesh connectivity configuration such that at least one of the die-to-die (D2D) interconnects of each chiplet is coupled to one of the D2D interconnects of each other chiplet using a non-diagonal link. The chiplets can also include other interfaces to facilitate communication between the chiplets, memory and a server or host system.Type: ApplicationFiled: October 13, 2023Publication date: April 17, 2025Inventors: Jayaprakash BALACHANDRAN, Aayush Ankit, Irene Quek, Sudeep Bhoja
-
Publication number: 20250110885Abstract: A transformer compute apparatus and method of operation therefor. The apparatus receives matrix inputs in a first format and generates projection tokens from these inputs. Among others, the apparatus includes a first cache device configured for processing first projection tokens and a second cache device configured for processing second projection tokens. The first cache device stores the first projection tokens in a first cache region and stores these tokens converted to a second format in a second cache region. The second cache device stores the second projection tokens converted to the second format in a first cache region and stores the converted second projection tokens after being transposed. Then, a compute device performs various matrix computations with the converted first projection tokens and transposed second projection tokens. Re-processing data and expensive padding and de-padding operations for transposed storage and byte alignment can be avoided using this caching process.Type: ApplicationFiled: November 18, 2024Publication date: April 3, 2025Inventors: Akhil ARUNKUMAR, Satyam Srivastava, Aayush Ankit
-
Patent number: 12182028Abstract: A transformer compute apparatus and method of operation therefor. The apparatus receives matrix inputs in a first format and generates projection tokens from these inputs. Among others, the apparatus includes a first cache device configured for processing first projection tokens and a second cache device configured for processing second projection tokens. The first cache device stores the first projection tokens in a first cache region and stores these tokens converted to a second format in a second cache region. The second cache device stores the second projection tokens converted to the second format in a first cache region and stores the converted second projection tokens after being transposed. Then, a compute device performs various matrix computations with the converted first projection tokens and transposed second projection tokens. Re-processing data and expensive padding and de-padding operations for transposed storage and byte alignment can be avoided using this caching process.Type: GrantFiled: September 28, 2023Date of Patent: December 31, 2024Assignee: d-MATRIX CORPORATIONInventors: Akhil Arunkumar, Satyam Srivastava, Aayush Ankit
-
Publication number: 20240303117Abstract: Operations of a workload are assigned to physical resources of a physical device array. The workload includes a graph of operations to be performed on a physical device array. The graph of operations is partitioned into subgraphs. Partitioning includes at least minimizing the quantity of subgraphs and maximizing resource utilization per subgraph. A logical mapping of the subgraph to logical processing engine (PE) units is generated using features of the subgraph and tiling factors of the logical PE units. The logical mapping is assigned to physical PE units of the physical device array at least by minimizing network traffic across the physical PE units. The operations of the subgraph are performed using the physical PE units to which the logical mapping is assigned. This process enhances the computational efficiency of the array when executing the workload.Type: ApplicationFiled: February 24, 2023Publication date: September 12, 2024Inventors: Fanny NINA PARAVECINO, Michael Eric DAVIES, Abhishek Dilip KULKARNI, Md Aamir RAIHAN, Ankit MORE, Aayush ANKIT, Torsten HOEFLER, Douglas Christopher BURGER
-
Publication number: 20240090181Abstract: An immersion cooling server system with AI accelerator apparatuses using in-memory compute chiplet devices. This system includes one or more immersion tanks with heat transfer fluid and configured with at least a condenser device. A plurality of AI accelerator servers is immersed in the heat transfer fluid in a bottom portion of the tanks and is configured to process transformer workloads while cooled by the immersion cooling configuration. Each of the servers includes a plurality of multiprocessors each having at least a first server central processing unit (CPU) and a second server CPU, both of which are coupled to a plurality of switch devices. Each switch device is coupled to a plurality of AI accelerator apparatuses. The apparatus includes one or more chiplets, each of which includes a plurality of digital in-memory compute (DIMC) devices configured to perform high throughput matrix computations for transformer based models.Type: ApplicationFiled: November 16, 2023Publication date: March 14, 2024Inventors: Jayaprakash BALACHANDRAN, Akhil ARUNKUMAR, Aayush ANKIT, Nithesh Kurella, Sudeep Bhoja
-
Patent number: 11929141Abstract: Sparsity-aware reconfiguration compute-in-memory (CIM) static random access memory (SRAM) systems are disclosed. In one aspect, a reconfigurable precision succession approximation register (SAR) analog-to-digital converter (ADC) that has the ability to form (n+m) bit precision using n-bit and m-bit sub-ADCs is provided. By controlling which sub-ADCs are used based on data sparsity, precision may be maintained as needed while providing a more energy efficient design.Type: GrantFiled: December 8, 2021Date of Patent: March 12, 2024Assignee: Purdue Research FoundationInventors: Kaushik Roy, Amogh Agrawal, Mustafa Fayez Ahmed Ali, Indranil Chakraborty, Aayush Ankit, Utkarsh Saxena
-
Publication number: 20240037379Abstract: A server system with AI accelerator apparatuses using in-memory compute chiplet devices. The system includes a plurality of multiprocessors each having at least a first server central processing unit (CPU) and a second server CPU, both of which are coupled to a plurality of switch devices. Each switch device is coupled to a plurality of AI accelerator apparatuses. The apparatus includes one or more chiplets, each of which includes a plurality of tiles. Each tile includes a plurality of slices, a CPU, and a hardware dispatch device. Each slice can include a digital in-memory compute (DIMC) device configured to perform high throughput computations. In particular, the DIMC device can be configured to accelerate the computations of attention functions for transformer-based models (a.k.a. transformers) applied to machine learning applications. A single input multiple data (SIMD) device configured to further process the DIMC output and compute softmax functions for the attention functions.Type: ApplicationFiled: October 13, 2023Publication date: February 1, 2024Inventors: Jayaprakash BALACHANDRAN, Akhil ARUNKUMAR, Aayush ANKIT, Nithesh Kurella, Sudeep Bhoja
-
Publication number: 20230178125Abstract: Sparsity-aware reconfiguration compute-in-memory (CIM) static random access memory (SRAM) systems are disclosed. In one aspect, a reconfigurable precision succession approximation register (SAR) analog-to-digital converter (ADC) that has the ability to form (n+m) bit precision using n-bit and m-bit sub-ADCs is provided. By controlling which sub-ADCs are used based on data sparsity, precision may be maintained as needed while providing a more energy efficient design.Type: ApplicationFiled: December 8, 2021Publication date: June 8, 2023Inventors: Kaushik Roy, Amogh Agrawal, Mustafa Fayez Ahmed Ali, Indranil Chakraborty, Aayush Ankit, Utkarsh Saxena
-
Patent number: 11120605Abstract: A texture filtering unit and a method are disclosed that provide multiple variants of an approximate trilinear filtering operation. A texture sampling and filtering unit may be configured to determine a level-of-detail (LOD) value for a sample point in texture space, and select, based on the LOD value, a fine mip-level and a coarse mip-level from the mip-map. The closer of the two selected mip-levels to the sample point is determined, and farther of the two selected mip-levels from the sample point is determined. A first quad of texels in the closer mip-level and a second quad of texels in the farther mip-level are then determined. A total of five or fewer texels are selected from the first quad of texels and from the second quad of texels. A filtered value for the sample point is determined based on an approximate trilinear filtering operation on the selected texels.Type: GrantFiled: March 24, 2020Date of Patent: September 14, 2021Inventors: Anders M. Kugler, Aayush Ankit, Wilson Wai Lun Fung
-
Publication number: 20210217220Abstract: A texture filtering unit and a method are disclosed that provide multiple variants of an approximate trilinear filtering operation. A texture sampling and filtering unit may be configured to determine a level-of-detail (LOD) value for a sample point in texture space, and select, based on the LOD value, a fine mip-level and a coarse mip-level from the mip-map. The closer of the two selected mip-levels to the sample point is determined, and farther of the two selected mip-levels from the sample point is determined. A first quad of texels in the closer mip-level and a second quad of texels in the farther mip-level are then determined. A total of five or fewer texels are selected from the first quad of texels and from the second quad of texels. A filtered value for the sample point is determined based on an approximate trilinear filtering operation on the selected texels.Type: ApplicationFiled: March 24, 2020Publication date: July 15, 2021Inventors: Anders M. KUGLER, Aayush ANKIT, Wilson Wai Lun FUNG