Patents Examined by Corey S Faherty
  • Patent number: 12730637
    Abstract: An apparatus and method for efficiently scheduling instructions for a parallel data processing circuit. In various implementations, a computing system includes a parallel data processing circuit with multiple compute circuits, each uses multiple single instruction multiple data (SIMD) circuits. Each compute circuit includes a scheduler for selecting instructions to issue to the SIMD circuits. During execution, a thread executes an instruction that provides a point of synchronization. Examples are the wait instruction and the barrier instruction. A control circuit accesses the metrics indicating hardware behavior of the corresponding wave. Based on these metrics, the control circuit generates a prediction of the amount of time before the point of synchronization completes. For example, the prediction indicates how soon each of the other threads of the corresponding wave are to arrive at the point of synchronization. The prediction is used to update control flow of the thread.
    Type: Grant
    Filed: September 26, 2024
    Date of Patent: September 8, 2026
    Assignees: Advanced Micro Devices, Inc., ATI Technologies ULC
    Inventors: Johnathan Alsop, Bradford M. Beckmann
  • Patent number: 12730638
    Abstract: An apparatus comprises exception return state register storage, and processing circuitry. In response to a guarded control stack (GCS) exception return state push instruction, the processing circuitry obtains exception return state information from the exception return state register storage and push the state information to a GCS data structure. In response to a GCS exception return state pop instruction, the processing circuitry obtains GCS-protected exception return state information from the GCS data structure. In at least one operating state, the processing circuitry detects, in response to an attempt to modify the exception return state information stored in the exception return state register storage, whether an exception return state lock parameter is in a locked state or an unlocked state, and signals a fault when it is in the locked state.
    Type: Grant
    Filed: March 17, 2023
    Date of Patent: September 8, 2026
    Assignee: Arm Limited
    Inventors: Simon John Craske, John Michael Horley
  • Patent number: 12730646
    Abstract: Branch target buffer structures are provided. A device can include a hierarchy of branch target buffers storing entries corresponding to branch instructions, the hierarchy of branch target buffers including respective branch target buffers that have progressively slower access times. The device can include a first program counter configured to generate a first program counter value associated with a next instruction of an executing application. The device can include a second program counter configured to predict a second program counter value that is associated with a subsequent instruction of the executing application that is after the next instruction. The device can include first branch prediction circuitry configured to populate a branch target buffer of the branch target buffers based on the second program counter value.
    Type: Grant
    Filed: September 28, 2023
    Date of Patent: September 8, 2026
    Assignee: Microsoft Technology Licensing, LLC
    Inventors: Julio Gago Alonso, Santiago Galan, Antonio Juan Hormigo, Ivan Pizarro
  • Patent number: 12730649
    Abstract: Apparatuses, computer programs and methods are disclosed, relating 2D arrays of data elements and a 2D array of processing elements. A first 2D array of data elements provides data values for processing by each processing element and a second 2D array of data elements provides control values controlling the processing. The 2D array of processing elements has a data flow direction across the 2D array of processing elements, and the data flow direction proceeds from a starting set of processing elements of the 2D array of processing elements. For each processing element not in the starting set of processing elements, the data processing operation preformed takes as an operand a respective data value provided by a neighbouring data element of the first 2D array of data elements and selection of the neighbouring processing element is controlled by a corresponding source control value in the second 2D array of data elements.
    Type: Grant
    Filed: September 19, 2024
    Date of Patent: September 8, 2026
    Assignee: Arm Limited
    Inventor: Alejandro Martinez Vicente
  • Patent number: 12730633
    Abstract: Apparatuses, systems, and techniques to perform a matrix multiply-accumulate (MMA) instruction to exclusively store information to be used by the MMA instruction. In at least one embodiment, a processor retrieves matrix information from a memory that exclusively stores and performs a multiplication computation using said matrix information.
    Type: Grant
    Filed: August 29, 2024
    Date of Patent: September 8, 2026
    Assignee: NVIDIA Corporation
    Inventors: Harold Carter Edwards, Vijay Harshad Thakkar, Gokul Ramaswamy Hirisave Chandra Shekhara, Edward H. Gornish, Rishkul Kulkarni, Maciej Piotr Tyrlik, Sean Jeffrey Treichler, Chao Li
  • Patent number: 12730647
    Abstract: An apparatus comprises execution circuitry configured to execute a given instruction to produce a given data value. Value analysis circuitry is configured to perform an analysis of the given data value produced by the execution circuitry to determine at least one property of the given data value, and register allocation circuitry is configured to make a register allocation decision regarding storage of the given data value in a physical register file in dependence on the analysis of the given data value.
    Type: Grant
    Filed: October 1, 2024
    Date of Patent: September 8, 2026
    Assignee: Arm Limited
    Inventors: Rami Mohammad Al Sheikh, Rodney Wayne Smith, Kiran Ravi Seth, Michael David Achenbach
  • Patent number: 12717630
    Abstract: A processing system includes one or more network-attached hostless accelerator (NAHA) units each having an integrated circuit. The integrated circuit of each NAHA unit includes a memory and a first network interface controller configured to communicatively couple the memory to a network such that the memory of the NAHA unit is communicatively coupled to one or more memories of other NAHA units via the network. Additionally, a NAHA unit of the processing system includes one or more processor cores configured to execute an instruction to generate a result. Further, the one or more processor cores of the NAHA unit are configured to store the result in the one or more memories of the other NAHA units.
    Type: Grant
    Filed: March 14, 2023
    Date of Patent: August 25, 2026
    Assignee: Advanced Micro Devices, Inc.
    Inventor: Mazda Sabony
  • Patent number: 12717752
    Abstract: In various examples, systems and methods are disclosed that relate to programming multi-dimensional single instruction, multiple data (SIMD) processors (also referred to as an accelerator). In one example, a processor can obtain instructions to be performed by the accelerator. The processor can determine one or more operations to be performed by the accelerator based at least on the instructions and generate a set of accelerator instructions. In examples, the processor can then provide data associated with the accelerator instructions to cause the accelerator to perform at least a portion of the one or more operations.
    Type: Grant
    Filed: July 31, 2024
    Date of Patent: August 25, 2026
    Assignee: NVIDIA Corporation
    Inventors: Andrew Peter Taussig, Ravi Pratap Singh, Ching-Yu Hung, Sreenivas Krishnan, Jagadeesh Sankaran, Yen-Te Shih
  • Patent number: 12699564
    Abstract: Segment load operations are performed by processing data through an anything-to-anything mux, and sections writing elements to respective storage locations based on corresponding indices of the elements and the storage locations. Once all of the elements are loaded into the correct storage location, each location is read again with the elements of that storage location being sent through the mux, arranged) into the correct order, and written back to the same register.
    Type: Grant
    Filed: January 10, 2025
    Date of Patent: August 4, 2026
    Assignee: Imagination Technologies Limited
    Inventor: Peter Vrabel
  • Patent number: 12699568
    Abstract: Disclosed are techniques for runtime adaptive prefetching in a many-core system. In an aspect, a method for runtime adaptive prefetching in a many-core system may include periodically performing the following steps: determining, for a first processor core in a many-core system, a workload classification based on at least one performance indicator of the first processor core; determining a first prefetching configuration from a plurality of prefetching configurations based on the workload classification; and configuring at least the first processor core according to the first prefetching configuration.
    Type: Grant
    Filed: April 18, 2024
    Date of Patent: August 4, 2026
    Assignee: Ampere Computing LLC
    Inventors: Erika Susana Alcorta Lozano, Mahesh Jagdish Madhav, Raymond Scott Tetrick
  • Patent number: 12681724
    Abstract: Embodiments of systems, apparatuses, and methods for chained fused multiply add. In some embodiments, an apparatus includes a decoder to decode a single instruction having an opcode, a destination field representing a destination operand, a first source field representing a plurality of packed data source operands of a first type that have packed data elements of a first size, a second source field representing a plurality of packed data source operands that have packed data elements of a second size, and a field for a memory location that stores a scalar value. A register file having a plurality of packed data registers includes registers for the plurality of packed data source operands that have packed data elements of a first size, the source operands that have packed data elements of a second size, and the destination operand.
    Type: Grant
    Filed: August 26, 2024
    Date of Patent: July 14, 2026
    Assignee: Intel Corporation
    Inventors: Jesus Corbal, Robert Valentine, Roman S. Dubtsov, Nikita A. Shustrov, Mark J. Charney, Dennis R. Bradford, Milind B. Girkar, Edward T. Grochowski, Thomas D. Fletcher, Warren E. Ferguson
  • Patent number: 12681725
    Abstract: Techniques are disclosed involving selection circuitry for stored operands. An embodiment of an apparatus includes a random-access storage element array and a permute network. The storage element array is configured to store a set of input data to an operation in a set of entries allocated among one or more execution lanes. The permute network is connected to each entry of the set of entries and is configured to select from among the set of entries to provide operands to one or more source inputs of execution circuitry configured to perform the operation. For a given source input, the permute network provides connection to only a subset of the entries that are allocated to a given execution lane. In a further embodiment, the permute network is configured to support multiple modes of the operation.
    Type: Grant
    Filed: August 28, 2024
    Date of Patent: July 14, 2026
    Assignee: Apple Inc.
    Inventors: Anurag Choudhury, Evan R. Lissoos
  • Patent number: 12681806
    Abstract: Disclosed is a method for resetting configurable units in a reconfigurable processor with an array of configurable units and a force-quit controller on an integrated circuit substrate. The array includes multiple sub-arrays of configurable units. The method involves receiving a force-quit command at the force-quit controller and generating force-quit control signals to reset configurable units in a specific sub-array of the plurality of sub-arrays. The specific sub-array contains the force-quit controller. This method enhances the efficiency and reliability of reconfigurable processors by enabling targeted and controlled resets of configurable units within the reconfigurable processor architecture.
    Type: Grant
    Filed: July 30, 2024
    Date of Patent: July 14, 2026
    Assignee: SambaNova Systems, Inc.
    Inventor: Manish K Shah
  • Patent number: 12670123
    Abstract: A hardware accelerator is disclosed that can flexibly be configured to support differing data types and differing operation flows. The hardware accelerator includes a plurality of fixed tensor operation logic units, tensor operation pipeline logic configured to receive from the processor a pipeline command including a software-defined tensor operation pipeline definition defining a plurality of tensor operation stages in a tensor operation pipeline and associated predetermined tensor operations to be performed at each of the defined tensor operation stages. The hardware accelerator is further configured to receive tensor data to be computed by the tensor operation pipeline, and implement the tensor operation pipeline to perform the tensor operations in each of the tensor operation stages on the tensor data, to thereby produce a tensor operation pipeline result for the tensor data, and output the tensor operation pipeline result to the processor.
    Type: Grant
    Filed: July 10, 2024
    Date of Patent: June 30, 2026
    Assignee: Microsoft Technology Licensing, LLC
    Inventors: Ashraf Ayman Michail, Li Zhang, Nitin Naresh Garegrat, Thomas Craig Savell
  • Patent number: 12657027
    Abstract: At least one instruction storage coupled with a fetch unit including sets of fetch circuitry each having a same plurality of pipeline stages. The sets of fetch circuitry perform fetch operations to fetch blocks of instructions from the at least one instruction storage. Stall circuitry, in response to an indication of a hazard for a given pipeline stage of a first set of fetch circuitry, retains a fetch operation for a first block of instructions at the given pipeline stage, and zero or more fetch operations for zero or more corresponding blocks of instructions at zero or more preceding pipeline stages of the first set of fetch circuitry, until the hazard has been removed. The stall circuitry advances a fetch operation for a second block of instructions from the given pipeline stage of a second set of fetch circuitry, during an initial cycle of the one or more cycles.
    Type: Grant
    Filed: April 2, 2022
    Date of Patent: June 16, 2026
    Assignee: Intel Corporation
    Inventors: Eliyah Kilada, Ammon Christiansen, Ariel Fabien Sabba, Christopher Celio, Ankur Groen, Muhammad Faisal Azeem, Malihe Ahmadi, Rangeen Basu Roy Chowdhury
  • Patent number: 12650841
    Abstract: Systems, methods, and apparatuses for implementing capability informed prefetches are described.
    Type: Grant
    Filed: April 2, 2022
    Date of Patent: June 9, 2026
    Assignee: Intel Corporation
    Inventor: Scott D. Constable
  • Patent number: 12645454
    Abstract: Techniques for shared data prefetch are described. An exemplary instruction for shared data prefetch includes at least one field for an opcode, at least one field for a source operand to provide a memory address at least a byte of data, wherein the opcode is to indicate that circuitry is to fetch of a line of data from memory at the provided address that contains the byte specified with the source operand and store that byte in at least a cache local to a requester, wherein the byte of data is to be stored in a shared state.
    Type: Grant
    Filed: September 25, 2021
    Date of Patent: June 2, 2026
    Assignee: Intel Corporation
    Inventors: Christopher Hughes, Zhe Wang, Dan Baum, Alexander Heinecke, Evangelos Georganas, Lingxiang Xiang, Joseph Nuzman, Ritu Gupta
  • Patent number: 12639256
    Abstract: Embodiments herein describe a hardware accelerator with an array of data processing engines (DPEs) which includes a controller (e.g., a microcontroller) for multiple columns of the array. The controllers can be hardened circuitry that executes software code (or firmware) that controls the hardware accelerator. In one embodiment, the task of the controller is to control and orchestrate the functions performed by the hardware accelerator.
    Type: Grant
    Filed: July 23, 2024
    Date of Patent: May 26, 2026
    Assignee: XILINX, INC.
    Inventors: Juan J. Noguera Serra, David Patrick Clarke, Javier Cabezas Rodriguez, Mikhail Asiatici, Patrick Schlangen
  • Patent number: 12632259
    Abstract: Techniques for using soft-barrier hints are described. An example includes a synchronous microthreading (SyMT) co-processor coupled to a logical processor to execute a plurality of microthreads, with each microthread having an independent register state, upon an execution of an instruction to enter into SyMT mode, wherein the SyMT co-processor is further to support a soft-barrier hint instruction in code which when processed by a microthread is to pause execution of the microthread to be resumed based at least in part on a data structure having at least one entry, the entry to include an instruction pointer of the soft-barrier hint instruction and a count of microthreads that have encountered the soft-barrier hint instruction at the instruction pointer.
    Type: Grant
    Filed: April 2, 2022
    Date of Patent: May 19, 2026
    Assignee: Intel Corporation
    Inventors: Shreesha Srinath, Jonathan Pearce, David B. Sheffield, Ching-Kai Liang, Jeffrey Cook
  • Patent number: 12625706
    Abstract: Embodiments of this application disclose an instruction translation method. The method includes: obtaining a return instruction of a function call instruction; obtaining a first address mapping result based on a second address indicated in the return instruction; storing the first address mapping result in a running stack space; and obtaining a first translation result of the return instruction, where the first translation result is a binary translation result of the return instruction, and the second translation result indicates to obtain, from a target location, an instruction indicated by the first address mapping result and execute the instruction. In this application, a running stack space of a source program is reused, thereby saving a storage space. In addition, an address of a return instruction does not need to be checked each time the return instruction is translated, thereby reducing overheads during translation and increasing program running efficiency.
    Type: Grant
    Filed: September 26, 2024
    Date of Patent: May 12, 2026
    Assignee: HUAWEI TECHNOLOGIES CO., LTD.
    Inventors: Xianzhe Liu, Jianjiang Zeng, Yandong Lv