Patents by Inventor William Peter Ehrett
William Peter Ehrett has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Publication number: 20260087091Abstract: A processor includes a plurality of processing elements. The processor is configured to execute a work graph including a plurality of nodes representing kernels executable by one or more processing elements of the plurality of processing elements. A first processing element of the one or more of the processing elements associated with a first node of the plurality of nodes is configured to assign each logical division of a sparse input matrix to a bin of a plurality of bins. Responsive to a dispatch condition associated with a bin of the plurality of bins, the processor is configured to dispatch a workgroup to at least a second processing element associated with at least a second node of the plurality of nodes corresponding to the bin. The workgroup includes a plurality of work items based on one or more logical divisions of the sparse input matrix assigned to the bin.Type: ApplicationFiled: September 25, 2024Publication date: March 26, 2026Inventors: William Peter Ehrett, Bradford Michael Beckmann, Paul Fridtjof Trojahn, Dominik Joerg Baumeister, Fabian Robert Sebastian Wildgrube
-
Publication number: 20260003576Abstract: In aspects of near-memory random and pattern-based data generation, a system includes a number generator circuit configured to generate a sequence of numbers, a memory chip configured to store the sequence of numbers, and a memory interface configured to enable communication between the number generator circuit and the memory chip. In one or more implementations, the number generator circuit includes a random number generator circuit configured to generate the sequence of numbers as a sequence of random numbers. Additionally, or alternatively, the number generator circuit includes a pattern fill function configured to generate the sequence of numbers based on a pattern. In other aspects of near-memory random and pattern-based data generation, a memory device includes a base layer, a memory interface, and a number generator circuit interleaved among the base layer and the memory interface.Type: ApplicationFiled: June 26, 2024Publication date: January 1, 2026Applicant: Advanced Micro Devices, Inc.Inventors: William Peter Ehrett, Nuwan S. Jayasena, Yasuko Eckert, Gabriel Hsiuwei Loh
-
Patent number: 12461872Abstract: A semiconductor device, referred to herein as a Globally Interconnected Operations (GIO) layer, provides global operations in the form of global data reduction for one or more PE arrays. The GIO layer includes processing elements that perform global data reduction on processing results from one or more PE arrays. The GIO layer includes connectors that allow it to be arranged in a 3D stack with one or more PE arrays, for example, on top of or beneath a PE array. This allows reduction operations to be implemented across PE arrays using an efficient topology with superior flexibility, scalability, latency and/or power characteristics that is customizable for particular use cases at assembly time, without requiring costly and time-consuming redesign of PE arrays, and without being constrained by particular PE array designs.Type: GrantFiled: June 30, 2023Date of Patent: November 4, 2025Assignee: Advanced Micro Devices, Inc.Inventors: William Peter Ehrett, Anthony Gutierrez, Vedula Venkata Srikant Bharadwaj, Karthik Ramu Sangaiah, Prachi Shukla, Sriseshan Srikanth, Ganesh Dasika, John Kalamatianos
-
Publication number: 20250199850Abstract: An apparatus and method for efficiently scheduling kernels for execution in a computing system. In various implementations, a computing system includes a cache and a processing circuit with multiple compute circuits and a scheduler. The scheduler groups kernels into scheduling groups where each scheduling group includes particular kernels of the multiple kernels that access a same data set different from a data set of another scheduling group. Each of these scheduling groups is referred to as a “cohort.” The scheduler accesses completion time estimates of kernels of the cohorts. Using the completion time estimates, the number of kernels currently executing, and the number of remaining kernels that have not yet begun execution of each currently scheduled cohort, the scheduler determines whether to immediately schedule a next cohort or delay scheduling the next cohort. By doing so, the scheduler balances throughput and cache contention.Type: ApplicationFiled: December 14, 2023Publication date: June 19, 2025Inventors: Vinay Bharadwaj Ramakrishnaiah, Bradford Michael Beckmann, William Peter Ehrett
-
Publication number: 20250123846Abstract: A processing unit includes a plurality of processing cores and is configured to arrange a sparse matrix for parallel performance by the cores on different rows of the matrix at least in part by calculating a respective quantity of non-zero elements in each row, assigning each row to a respective collection according to the respective quantity of non-zero elements for the row, wherein the processing unit is configured to assign at least one first row of the sparse matrix to respective collections of in parallel with assigning at least one second row of the sparse matrix to respective collections, and performing at least one mathematical operation on at least a first collection of the plurality of collections in parallel with performing the at least one mathematical operation on at least a second collection of the plurality of collections.Type: ApplicationFiled: October 12, 2023Publication date: April 17, 2025Applicant: Advanced Micro Devices, Inc.Inventors: William Peter Ehrett, Muhammad Osama, Bradford Beckmann
-
Publication number: 20250103395Abstract: A computer-implemented method for dynamic resource management can include evaluating, by at least one processor, whether a priority of one or more processes associated with a request for one or more shared resources meets a threshold condition. The method can additionally include determining, by the at least one processor and in response to an evaluation that the priority meets the threshold condition, whether the one or more shared resources is available to meet the request. The method can further include completing, by the at least one processor and in response to a determination that the one or more shared resources is available, execution of the one or more processes. Various other methods, systems, and computer-readable media are also disclosed.Type: ApplicationFiled: September 27, 2023Publication date: March 27, 2025Applicant: Advanced Micro Devices, Inc.Inventors: Bradford Beckmann, Matthew David Sinclair, Vinay Bharadwaj Ramakrishnaiah, William Peter Ehrett
-
Publication number: 20250004963Abstract: A semiconductor device, referred to herein as a Globally Interconnected Operations (GIO) layer, provides global operations in the form of global data reduction for one or more PE arrays. The GIO layer includes processing elements that perform global data reduction on processing results from one or more PE arrays. The GIO layer includes connectors that allow it to be arranged in a 3D stack with one or more PE arrays, for example, on top of or beneath a PE array. This allows reduction operations to be implemented across PE arrays using an efficient topology with superior flexibility, scalability, latency and/or power characteristics that is customizable for particular use cases at assembly time, without requiring costly and time-consuming redesign of PE arrays, and without being constrained by particular PE array designs.Type: ApplicationFiled: June 30, 2023Publication date: January 2, 2025Inventors: William Peter Ehrett, Anthony Gutierrez, Vedula Venkata Srikant Bharadwaj, Karthik Ramu Sangaiah, Prachi Shukla, Sriseshan Srikanth, Ganesh Dasika, John Kalamatianos
-
Publication number: 20240330045Abstract: A technique for scheduling executing items on a highly parallel processing architecture is provided. The technique includes identifying a plurality of execution items that share data, as indicated by having matching commonality metadata; identifying an execution unit for executing the plurality of execution items together; and scheduling the plurality of execution items for execution together on the execution unit.Type: ApplicationFiled: March 31, 2023Publication date: October 3, 2024Applicant: Advanced Micro Devices, Inc.Inventors: William Peter Ehrett, Vinay Bharadwaj Ramakrishnaiah, Bradford M. Beckmann