Patents by Inventor David K. Li
David K. Li has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Patent number: 12724615Abstract: Techniques are disclosed relating to accessing source data in single-instruction multiple-thread (SIMT) pipelines. In some embodiments, multiple categories of operand resource circuits are configured to provide operands for instructions executed by processor pipeline circuitry. Per-resource arbitration circuitry may arbitrate between the SIMT execution slots for access to different operand resources. Source access circuitry may access operand data from operand resources based on source capture commands and source control circuitry may prior to a first SIMT group winning arbitration for all its operands, send a source capture command to the source access circuitry in response to the first SIMT group winning arbitration at the per-resource arbitration circuitry for a first operand resource. Instruction control circuitry may send an instruction release command down the processor pipeline circuitry for the first SIMT group, in response to the first SIMT group winning arbitration for all its operands.Type: GrantFiled: October 10, 2024Date of Patent: September 1, 2026Assignee: Apple Inc.Inventors: David K. Li, Ana Lucia Rescala Loper, Arun Bansal, Chance C. Coats, Zeran Zhu
-
Patent number: 12625821Abstract: Techniques are disclosed relating to eviction control for cache lines that store register data. In some embodiments, memory hierarchy circuitry is configured to provide memory backing for register operand data in one or more cache circuits. Lock circuitry may control a first set of lock indicators for a set of registers for a first thread, including to assert one or more lock indicators for registers that are indicated, by decode circuitry, as being utilized by decoded instructions of the first thread. The lock circuitry may preserve register operand data in the one or more cache circuits, including to prevent eviction of a given cache line from a cache circuit based on an asserted lock indicator. The lock circuitry may clear the first set of lock indicators in response to a reset event. Disclosed techniques may advantageously retain relevant register information in the cache with limited control circuit area.Type: GrantFiled: November 27, 2024Date of Patent: May 12, 2026Assignee: Apple Inc.Inventors: Jonathan M. Redshaw, Winnie W. Yeung, Benjiman L. Goodman, David K. Li, Zelin Zhang, Yoong Chert Foo
-
Patent number: 12625810Abstract: Techniques are disclosed relating to cache control in cache hierarchies. In some embodiments, processor execution circuitry is configured to perform operations on input operand data from a first-level cache, including a first operation that reads first data from an entry in the first-level cache and signals an invalidation of the first data. Control circuitry may set an indicator, in response to the first operation, to indicate that the entry in the first-level cache has a pending invalidation (e.g., a last-use indicator). The control circuitry may, in response to a second operation overwriting the entry in the first-level cache while the indicator is set, clear the indicator without invalidating a corresponding entry in a second-level cache. This may advantageously reduce invalidate operations and bandwidth to the second-level cache.Type: GrantFiled: October 10, 2024Date of Patent: May 12, 2026Assignee: Apple Inc.Inventors: David K. Li, Yoong Chert Foo, Benjiman L. Goodman, Chance C. Coats
-
Publication number: 20260087093Abstract: Techniques are disclosed relating to hardware acceleration for matrix operations. In some embodiments, processor circuitry configured to execute a first matrix multiply instruction and a second matrix multiply instruction included in an execution thread. Control circuitry generates a first set of operations, including multiple dot product and multiple accumulate operations, for the first matrix multiply instruction and a second set of operations, including multiple dot product and multiple accumulate operations, for the second matrix multiply instruction. The control circuitry specifies a first execution rate for the first set of operations and a second, different execution rate for the second set of operations. Matrix acceleration circuitry performs the sets of operations at the specified rates. The rates may be based on thermal measurements.Type: ApplicationFiled: January 9, 2025Publication date: March 26, 2026Inventors: Angel E. Socarras, Anurag Choudhury, Arvind Arunasalam, Benoit M. Jubelin, David K. Li, Evan R. Lissoos
-
Publication number: 20260086944Abstract: Techniques are disclosed relating to cache control in cache hierarchies. In some embodiments, processor execution circuitry is configured to perform operations on input operand data from a first-level cache, including a first operation that reads first data from an entry in the first-level cache and signals an invalidation of the first data. Control circuitry may set an indicator, in response to the first operation, to indicate that the entry in the first-level cache has a pending invalidation (e.g., a last-use indicator). The control circuitry may, in response to a second operation overwriting the entry in the first-level cache while the indicator is set, clear the indicator without invalidating a corresponding entry in a second-level cache. This may advantageously reduce invalidate operations and bandwidth to the second-level cache.Type: ApplicationFiled: October 10, 2024Publication date: March 26, 2026Inventors: David K. Li, Yoong Chert Foo, Benjiman L. Goodman, Chance C. Coats
-
Publication number: 20260079715Abstract: Techniques are disclosed relating to accessing source data in single-instruction multiple-thread (SIMT) pipelines. In some embodiments, multiple categories of operand resource circuits are configured to provide operands for instructions executed by processor pipeline circuitry. Per-resource arbitration circuitry may arbitrate between the SIMT execution slots for access to different operand resources. Source access circuitry may access operand data from operand resources based on source capture commands and source control circuitry may prior to a first SIMT group winning arbitration for all its operands, send a source capture command to the source access circuitry in response to the first SIMT group winning arbitration at the per-resource arbitration circuitry for a first operand resource. Instruction control circuitry may send an instruction release command down the processor pipeline circuitry for the first SIMT group, in response to the first SIMT group winning arbitration for all its operands.Type: ApplicationFiled: October 10, 2024Publication date: March 19, 2026Inventors: David K. Li, Ana Lucia Rescala Loper, Arun Bansal, Chance C. Coats, Zeran Zhu
-
Publication number: 20250103292Abstract: Techniques are disclosed relating to integrated circuits that support matrix operations. In various embodiments, an integrated circuit comprises a dot product accumulate circuit that includes a dot product circuit configured to determine a dot product of a first vector and a second vector, and an adder circuit coupled to an output of the dot product circuit and configured to add a result of the dot product and an accumulation value. The integrated circuit further includes an accumulator cache coupled to an input of the adder circuit and an output of the adder circuit. The accumulator cache is configured to provide the accumulation value to the adder circuit and store a result of the add as a subsequent accumulation value for a subsequent dot product accumulate operation.Type: ApplicationFiled: January 29, 2024Publication date: March 27, 2025Inventors: David K. Li, Christopher A. Burns, Daniel E. Barnard, Evan R. Lissoos
-
Publication number: 20250094357Abstract: Techniques are disclosed relating to eviction control for cache lines that store register data. In some embodiments, memory hierarchy circuitry is configured to provide memory backing for register operand data in one or more cache circuits. Lock circuitry may control a first set of lock indicators for a set of registers for a first thread, including to assert one or more lock indicators for registers that are indicated, by decode circuitry, as being utilized by decoded instructions of the first thread. The lock circuitry may preserve register operand data in the one or more cache circuits, including to prevent eviction of a given cache line from a cache circuit based on an asserted lock indicator. The lock circuitry may clear the first set of lock indicators in response to a reset event. Disclosed techniques may advantageously retain relevant register information in the cache with limited control circuit area.Type: ApplicationFiled: November 27, 2024Publication date: March 20, 2025Inventors: Jonathan M. Redshaw, Winnie W. Yeung, Benjiman L. Goodman, David K. Li, Zelin Zhang, Yoong Chert Foo
-
Patent number: 12182037Abstract: Techniques are disclosed relating to eviction control for cache lines that store register data. In some embodiments, memory hierarchy circuitry is configured to provide memory backing for register operand data in one or more cache circuits. Lock circuitry may control a first set of lock indicators for a set of registers for a first thread, including to assert one or more lock indicators for registers that are indicated, by decode circuitry, as being utilized by decoded instructions of the first thread. The lock circuitry may preserve register operand data in the one or more cache circuits, including to prevent eviction of a given cache line from a cache circuit based on an asserted lock indicator. The lock circuitry may clear the first set of lock indicators in response to a reset event. Disclosed techniques may advantageously retain relevant register information in the cache with limited control circuit area.Type: GrantFiled: February 23, 2023Date of Patent: December 31, 2024Assignee: Apple Inc.Inventors: Jonathan M. Redshaw, Winnie W. Yeung, Benjiman L. Goodman, David K. Li, Zelin Zhang, Yoong Chert Foo
-
Publication number: 20240289282Abstract: Techniques are disclosed relating to eviction control for cache lines that store register data. In some embodiments, memory hierarchy circuitry is configured to provide memory backing for register operand data in one or more cache circuits. Lock circuitry may control a first set of lock indicators for a set of registers for a first thread, including to assert one or more lock indicators for registers that are indicated, by decode circuitry, as being utilized by decoded instructions of the first thread. The lock circuitry may preserve register operand data in the one or more cache circuits, including to prevent eviction of a given cache line from a cache circuit based on an asserted lock indicator. The lock circuitry may clear the first set of lock indicators in response to a reset event. Disclosed techniques may advantageously retain relevant register information in the cache with limited control circuit area.Type: ApplicationFiled: February 23, 2023Publication date: August 29, 2024Inventors: Jonathan M. Redshaw, Winnie W. Yeung, Benjiman L. Goodman, David K. Li, Zelin Zhang, Yoong Chert Foo
-
Publication number: 20160378497Abstract: Embodiments of systems, methods, and apparatuses for thread selection and reservation station binding are disclosed. In an embodiment, an apparatus includes allocation hardware including reservation station binding logic to bind an operation to one of a plurality of reservation stations. In an embodiment, an apparatus includes thread selection logic to select a thread to be processed by a pipeline stage, wherein the thread selection logic to evaluate a plurality of conditions to select a thread, wherein the conditions include if a thread is active, if a thread has operations in an instruction queue, if a thread has available resources, and if a thread has no known stall.Type: ApplicationFiled: June 26, 2015Publication date: December 29, 2016Inventors: Roger Gramunt, Rammohan Padmanabhan, Gerardo A. Fernandez, David K. Li, Julio Gaga, Michael Yang, Jonathan C. Hall
-
Patent number: 7600103Abstract: Apparatus, systems and methods for speculative scheduling of uops after allocation are disclosed including an apparatus having logic to schedule a micro-operation (uop) for execution before source data of the uop is ready. The apparatus further includes logic to cancel dispatching of the uop for execution if the source data is invalid. Other implementations are disclosed.Type: GrantFiled: June 30, 2006Date of Patent: October 6, 2009Assignee: Intel CorporationInventors: Avinash Sodani, Rahul Kulkarni, David K. Li
-
Publication number: 20080005535Abstract: Apparatus, systems and methods for speculative scheduling of uops after allocation are disclosed including an apparatus having logic to schedule a micro-operation (uop) for execution before source data of the uop is ready. The apparatus further includes logic to cancel dispatching of the uop for execution if the source data is invalid. Other implementations are disclosed.Type: ApplicationFiled: June 30, 2006Publication date: January 3, 2008Inventors: Avinash Sodani, Rahul Kulkarni, David K. Li
-
Patent number: 6567337Abstract: A pulsed circuit topology to perform a memory array write operation. A write enable pulse width control circuit is responsive to a pulsed clock signal to generate a pulsed write enable signal and a write data path circuit is provided to output a write data signal. The write enable pulse width control circuit and the write data path circuit together control a write operation to a memory cell.Type: GrantFiled: June 30, 2000Date of Patent: May 20, 2003Assignee: Intel CorporationInventors: Milo D. Sprague, David K. Li, Robert J. Murray