Patents by Inventor Duane E. Galbi

Duane E. Galbi has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Patent number: 12645634
    Abstract: Methods, apparatus, and software and for hardware microservices accelerated in other processing units (XPUs). The apparatus may be a platform including a System on Chip (SOC) and an XPU, such as a Field Programmable Gate Array (FPGA). The FPGA is configured to implement one or more Hardware (HW) accelerator functions associated with HW microservices. Execution of microservices is split between a software front-end that executes on the SOC and a hardware backend comprising the HW accelerator functions. The software front-end offloads a portion of a microservice and/or associated workload to the HW microservice backend implemented by the accelerator functions. An XPU or FPGA proxy is used to provide the microservice front-ends with shared access to HW accelerator functions, and schedules/multiplexes access to the HW accelerator functions using, e.g., telemetry data generated by the microservice front-ends and/or the HW accelerator functions.
    Type: Grant
    Filed: December 13, 2021
    Date of Patent: June 2, 2026
    Assignee: Intel Corporation
    Inventors: Susanne M. Balle, Duane E. Galbi, Andrzej Kuriata, Sundar Nadathur, Nagabhushan Chitlur, Francesc Guim Bernat, Alexander Bachmutsky
  • Patent number: 12639219
    Abstract: A memory system has a configurable mapping of address space of a memory array to address of a memory access command. In response to a memory access command, a memory device can apply a traditional mapping of the command address to the address space, or can apply an address remapping to remap the command address to different address space.
    Type: Grant
    Filed: September 24, 2021
    Date of Patent: May 26, 2026
    Assignee: Intel Corporation
    Inventors: Duane E. Galbi, Kuljit S. Bains
  • Publication number: 20260133852
    Abstract: Examples described herein relate to an interface and a processor, coupled to the interface, that is configured to: offload decompression of codebook compressed data to a device, wherein the data comprises weight data, wherein the codebook compressed data comprises data represented by code values, and wherein the code values utilize less memory than the corresponding data. In some examples, the device comprises a direct memory access (DMA) engine. In some examples, the device comprises an accelerator to perform matrix multiplication or a decoder.
    Type: Application
    Filed: December 19, 2025
    Publication date: May 14, 2026
    Inventors: Duane E. GALBI, Ellick CHAN, Susanne M. BALLE
  • Publication number: 20260119061
    Abstract: Examples described herein relate to allocate a first of multiple memory-mapped input/output regions to a first sub-non-uniform memory access cluster and allocate a second of the memory-mapped input/output regions to a second sub-non-uniform memory access cluster. In some examples, the second of the memory-mapped input/output regions can be allocated to a third sub-non-uniform memory access cluster. In some examples, a sub-non-uniform memory access cluster includes a core and a caching agent.
    Type: Application
    Filed: December 24, 2025
    Publication date: April 30, 2026
    Inventors: Duane E. GALBI, Jeffrey HAGAN, Robert FABER, George Leonard TKACHUK, Susanne M. BALLE, Hugh WILKINSON
  • Patent number: 12572417
    Abstract: A translation cache and configurable error checking and correction (“ECC”) memory reduces ECC memory overhead. The translation cache supports a configurable ECC memory capable of storing a portion of a cache line, along with any ECC data, in corresponding parts of memory devices to reduce the ECC memory overhead in a memory subsystem. The corresponding parts include any same one of an upper, lower, left or right part of memory devices in a memory module, including dynamic random access memory (“DRAM”) devices in a dual inline memory module (“DIMM”).
    Type: Grant
    Filed: September 23, 2021
    Date of Patent: March 10, 2026
    Assignee: Intel Corporation
    Inventors: Duane E. Galbi, Wim Heirman, Dimitrios Ziakas
  • Patent number: 12450008
    Abstract: Methods, apparatus, and software for remote storage of hardware microservices hosted on other processing units (XPUs) and SOC-XPU Platforms. The apparatus may be a platform including a System on Chip (SOC) and an XPU, such as a Field Programmable Gate Array (FPGA). Software, via execution on the SOC, enables the platform to pre-provision storage space on a remote storage node and assign the storage space to the platform, wherein the pre-provisioned storage space includes one or more container images to be implemented as one or more hardware (HW) microservice front-ends. The XPU/FPGA is configured to implement one or more accelerator functions used to accelerate HW microservice backend operations that are offloaded from the one or more HW microservice front-ends. The platform is also configured to pre-provision a remote storage volume containing worker node components and access and persistently store worker node components.
    Type: Grant
    Filed: December 21, 2021
    Date of Patent: October 21, 2025
    Assignee: Intel Corporation
    Inventors: Andrzej Kuriata, Susanne M. Balle, Duane E. Galbi, Sundar Nadathur, Nagabhushan Chitlur, Francesc Guim Bernat, Alexander Bachmutsky
  • Publication number: 20250252055
    Abstract: Examples described herein relate to a processor that includes a core and a cache, coupled to the core. In some examples, the core is to perform an instruction of a process to specify loads of data from a source to destination regions of caches of a group of target cores. In some examples, the destination regions includes multiple different cache regions and wherein cores of the group of target cores have at least one respective cache.
    Type: Application
    Filed: April 24, 2025
    Publication date: August 7, 2025
    Inventors: Duane E. GALBI, Christopher J. HUGHES, Andrew J. HERDRICH, Simon C. STEELY, JR., Chen DAN
  • Publication number: 20240370699
    Abstract: Examples described herein relate to a processor to process constant weight values and key value entries associated with a first transformer kernel of a large language model (LLM) neural network and a circuitry. The circuitry is to: during processing of the constant weight values and key value entries associated with the first transformer kernel of the LLM neural network, pre-fetch constant weight values and key value entries associated with a second transformer kernel of the LLM neural network into a buffer.
    Type: Application
    Filed: July 5, 2024
    Publication date: November 7, 2024
    Inventors: Duane E. GALBI, Matthew Joseph ADILETTA, Matthew James ADILETTA
  • Patent number: 12111775
    Abstract: Examples described herein relate to an apparatus that includes at least two processing units and a memory hub coupled to the at least two processing units. In some examples, the memory hub includes a home agent. In some examples, the memory hub is to perform a memory access request involving a memory device, a first processing unit among the at least two processing units is to send the memory access request to the memory hub. In some examples, the first processing unit is to offload at least some but not all home agent operations to the home agent of the memory hub. In some examples, the first processing unit comprises a second home agent and wherein the second home agent is to perform the at least some but not all home agent operations before the offload of at least some but not all home agent operations to the home agent of the memory hub.
    Type: Grant
    Filed: March 25, 2021
    Date of Patent: October 8, 2024
    Assignee: Intel Corporation
    Inventors: Duane E. Galbi, Matthew J. Adiletta, Hugh Wilkinson, Patrick Connor
  • Patent number: 12099408
    Abstract: An apparatus is described. The apparatus includes a memory controller having logic circuitry to write a unit of write data into a plurality of memory chips according to a striping pattern that includes multiple protected sub words, each protected sub word including a smaller portion of the unit of write data and error correction coding (ECC) information calculated from the smaller portion of the unit of write data.
    Type: Grant
    Filed: December 23, 2020
    Date of Patent: September 24, 2024
    Assignee: Intel Corporation
    Inventors: Duane E. Galbi, Matthew J. Adiletta
  • Publication number: 20230333921
    Abstract: Examples described herein relate to a host interface and circuitry. In some examples, the circuitry, when coupled to a physical device, is to: perform operations of a hypervisor. In some examples, the host interface is configured to route first communications to the circuitry instead of the physical device and route second communications to the physical device. In some examples, the physical device is accessible as a virtual device via the host interface.
    Type: Application
    Filed: February 21, 2023
    Publication date: October 19, 2023
    Inventors: Noam ELATI, Piotr UMINSKI, Boris KLEIMAN, Lloyd DCRUZ, Bradley A. BURRES, Salma Mirza JOHNSON, Thomas E. WILLIS, Duane E. GALBI
  • Patent number: 11699471
    Abstract: An apparatus is described. The apparatus includes logic circuitry to multiplex on a data bus a first data burst, a second data burst, a third data burst and a fourth data burst having different respective base target addresses that respectively target a first memory rank, a second memory rank, a third memory rank and a fourth memory rank. A first data transfer for the first data burst occurs on a first edge of a first pulse of a data strobe signal for the data bus and a second data transfer for the second data burst occurs on a second edge of the first pulse of the data strobe signal. A third data transfer for the third data burst occurs on a first edge of a second pulse of the data strobe signal for the data bus and a fourth data transfer for the fourth data burst occurs on a second edge of the second pulse. The second pulse immediately follows the first pulse on the data strobe signal.
    Type: Grant
    Filed: September 23, 2020
    Date of Patent: July 11, 2023
    Assignee: Intel Corporation
    Inventors: Duane E. Galbi, Bill Nale
  • Publication number: 20230215493
    Abstract: Methods and apparatus for Cross DRAM DIMM sub-channel pairing. Memory channels on a memory controller or System on a Chip (SoC) are segmented into two subchannels, each including Command and Address (C/A) signals, DQ (data) lines. Under different solutions the two subchannels may share a command-bus clock or use separate command-bus clocks. Some approaches use subchannels from different memory channels to provide the C/A and DQ lines for two subchannels to a given DIMM. One solution implements an additional command-bus clock on the DIMM connector repurposing existing MCR pins to provide command-bus clock signals to a Registered Clock Driver (RCD) to allow the subchannels to be fully independent. Another solution is the pair every other DRAM controller to the same command-bus clock. Other solutions employ Skip-1, Skip-2, and Skip-3 configurations under which the clocks for the DDR-IO circuitry are not logically co-located with the subchannel IO circuitry.
    Type: Application
    Filed: March 13, 2023
    Publication date: July 6, 2023
    Inventors: Duane E. GALBI, Matthew J. ADILETTA, Mohammad M. RASHID, Todd HINCK, Vijaya K. BODDU
  • Publication number: 20230185760
    Abstract: Methods, apparatus, and software and for hardware microservices accelerated in other processing units (XPUs). The apparatus may be a platform including a System on Chip (SOC) and an XPU, such as a Field Programmable Gate Array (FPGA). The FPGA is configured to implement one or more Hardware (HW) accelerator functions associated with HW microservices. Execution of microservices is split between a software front-end that executes on the SOC and a hardware backend comprising the HW accelerator functions. The software front-end offloads a portion of a microservice and/or associated workload to the HW microservice backend implemented by the accelerator functions. An XPU or FPGA proxy is used to provide the microservice front-ends with shared access to HW accelerator functions, and schedules/multiplexes access to the HW accelerator functions using, e.g., telemetry data generated by the microservice front-ends and/or the HW accelerator functions.
    Type: Application
    Filed: December 13, 2021
    Publication date: June 15, 2023
    Inventors: Susanne M. BALLE, Duane E. GALBI, Andrzej KURIATA, Sundar NADATHUR, Nagabhushan CHITLUR, Francesc GUIM BERNAT, Alexander BACHMUTSKY
  • Publication number: 20220321434
    Abstract: Reliability and performance of a data center is increased by processing telemetry data in a network device in the data center. A Telemetry Correlation Engine (TCE) in the network device correlates host level telemetry received from a compute node with low-level network device telemetry collected in the network device to identify performance bottlenecks for microservices based applications. The Telemetry Correlation Engine processes and analyzes the telemetry data from the compute node and network statistics available in the network device.
    Type: Application
    Filed: June 24, 2022
    Publication date: October 6, 2022
    Inventors: Andrzej KURIATA, Francesc GUIM BERNAT, Karthik KUMAR, Susanne M. BALLE, Alexander BACHMUTSKY, Duane E. GALBI, Nagabhushan CHITLUR, Sundar NADATHUR
  • Publication number: 20220321491
    Abstract: Examples described herein relate to a network interface device that includes circuitry to process data and circuitry to split a received flow of a mixture of control and data content and provide the control content to a control plane processor and provide the data content for access to the circuitry to process data, wherein the mixture of control and data content are received as part of a Remote Procedure Call. In some examples, provide the control content to a control plane processor, the circuitry is to remove data content from a received packet and include an indicator of a location of removed data content in the received packet.
    Type: Application
    Filed: June 20, 2022
    Publication date: October 6, 2022
    Inventors: Susanne M. BALLE, Shihwei CHIEN, Duane E. GALBI, Nagabhushan CHITLUR
  • Publication number: 20220206864
    Abstract: Examples described herein relate to causing execution of a workload on a device based on characteristics of the device and based on metadata associated with the device identifying execution requirements and software and hardware compatibilities between the device and a platform environment. In some examples, an accelerator device is selected to execute a workload based on characteristics of the accelerator device and based on software and hardware compatibilities between the device and a platform environment of the accelerator device.
    Type: Application
    Filed: March 14, 2022
    Publication date: June 30, 2022
    Inventors: Sundar NADATHUR, Susanne M. BALLE, Andrzej KURIATA, Duane E. GALBI, Nagabhushan CHITLUR, Francesc GUIM BERNAT, Alexander BACHMUTSKY
  • Publication number: 20220113911
    Abstract: Methods, apparatus, and software for remote storage of hardware microservices hosted on other processing units (XPUs) and SOC-XPU Platforms. The apparatus may be a platform including a System on Chip (SOC) and an XPU, such as a Field Programmable Gate Array (FPGA). Software, via execution on the SOC, enables the platform to pre-provision storage space on a remote storage node and assign the storage space to the platform, wherein the pre-provisioned storage space includes one or more container images to be implemented as one or more hardware (HW) microservice front-ends. The XPU/FPGA is configured to implement one or more accelerator functions used to accelerate HW microservice backend operations that are offloaded from the one or more HW microservice front-ends. The platform is also configured to pre-provision a remote storage volume containing worker node components and access and persistently store worker node components.
    Type: Application
    Filed: December 21, 2021
    Publication date: April 14, 2022
    Inventors: Andrzej KURIATA, Susanne M. BALLE, Duane E. GALBI, Sundar NADATHUR, Nagabhushan CHITLUR, Francesc GUIM BERNAT, Alexander BACHMUTSKY
  • Publication number: 20220012195
    Abstract: A memory system has a configurable mapping of address space of a memory array to address of a memory access command. A controller provides command and enable information specific to a memory device. The command and enable information can cause the memory device to apply a traditional mapping of the command address to the address space, or can cause the memory device to apply an address remapping to remap the command address to different address space.
    Type: Application
    Filed: September 24, 2021
    Publication date: January 13, 2022
    Inventors: Duane E. GALBI, Kuljit S. BAINS
  • Publication number: 20220012126
    Abstract: A translation cache and configurable error checking and correction (“ECC”) memory reduces ECC memory overhead. The translation cache supports a configurable ECC memory capable of storing a portion of a cache line, along with any ECC data, in corresponding parts of memory devices to reduce the ECC memory overhead in a memory subsystem. The corresponding parts include any same one of an upper, lower, left or right part of memory devices in a memory module, including dynamic random access memory (“DRAM”) devices in a dual inline memory module (“DIMM”).
    Type: Application
    Filed: September 23, 2021
    Publication date: January 13, 2022
    Inventors: Duane E. GALBI, Wim HEIRMAN, Dimitrios ZIAKAS