Patents by Inventor Pierre Boudier
Pierre Boudier has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Publication number: 20260195959Abstract: A processing system comprises a memory and a processor configured to store multiple data elements of non-power-of-two bit width across plural registers with power-of-two bit width. The processor stores each bit of the data elements in separate registers, with the first bit in a first register, second bit in a second register, and so on. Arithmetic operations are performed on corresponding bits across registers. The number of data elements equals the register bit width. An arithmetic logic unit concurrently operates on the data elements using bits stored across registers. The system enables efficient storage and processing of low-precision values, optimizing cache usage and reducing wasted bits. This novel approach improves performance for applications like artificial intelligence using narrow data types.Type: ApplicationFiled: December 23, 2025Publication date: July 9, 2026Applicant: Intel CorporationInventors: Pierre Boudier, Sebastian Björn Herholz, Matthaeus Georg Chajdas, Graham John Sellers, Marcus Rogowsky
-
Publication number: 20260179302Abstract: A technique for graphics processing is disclosed. The technique includes processing an allocation command for a graphics resource object, wherein the graphics resource object is a representation of a data store, the allocation command reserving address space for the data store as a sparse data store with uncommitted memory; and processing a commitment command for the graphics resource object that specifies a range of the data store and includes an indicator of commitment state, wherein processing the commitment command sets a commitment state of memory for the specified range according to the indicator.Type: ApplicationFiled: February 11, 2026Publication date: June 25, 2026Applicant: Advanced Micro Devices, Inc.Inventors: Graham Sellers, Eric Zolnowski, Pierre Boudier, Juraj Obert
-
Patent number: 12620160Abstract: A system and method for performing graphics processing is provided. The system and method includes processing an allocation command for a buffer object; reserving processor address space for a data store of the buffer object with uncommitted physical memory in response to the allocation command including a null parameter, and reserving processor address space for a data store of the buffer object with committed physical memory in response to the allocation command including a non-null parameter.Type: GrantFiled: May 19, 2023Date of Patent: May 5, 2026Assignee: Advanced Micro Devices, Inc.Inventors: Graham Sellers, Eric Zolnowski, Pierre Boudier, Juraj Obert
-
Publication number: 20260037477Abstract: Post-synchronization operations in multi-tile processor computing is described. An example of an apparatus an apparatus includes a memory to store data for processing, including data for an application; and one or more processors including a graphical processing unit (GPU), the GPU including multiple compute engine tiles including multiple processing resources, and a dispatcher for dispatching kernels for processing by the compute engine tiles, wherein each of the compute engine tiles is to write a signal to a location in the memory upon the compute engine tile completing processing of a partition of a first kernel, wherein the location is a same location for each of the plurality of compute engine tiles.Type: ApplicationFiled: August 5, 2024Publication date: February 5, 2026Applicant: Intel CorporationInventors: Michal Mrozek, Vasanth Ranganathan, Pierre Boudier, Jeffery S. Boles, Aditya Navale, Hema Chand Nalluri
-
Publication number: 20250342384Abstract: Quantum-based dispatch of workgroups is described. An example of an apparatus includes a computer memory to store data for processing, including data for an application; and one or more processors including a graphical processing unit (GPU), the GPU including multiple chiplets, each of the multiple chiplets including compute containers and a cache, each compute container including a plurality of processing resources, and a dispatcher for dispatching workgroups to the processing resources of the GPU, wherein dispatching workgroups includes dispatching workgroups for the application according to a selected workgroup quantum, the selected workgroup quantum having a certain size and shape.Type: ApplicationFiled: May 2, 2024Publication date: November 6, 2025Applicant: Intel CorporationInventors: Milind Nemlekar, Changwon Rhee, Vasanth Ranganathan, Maxim Kazakov, Pierre Boudier, Wei Xiong, Moshe Maor, Deepak N K, Jain Philip, Abhishek Kumar Singh, Michal Mrozek
-
Publication number: 20250284567Abstract: One embodiment provides a multi-chiplet graphics processor comprising a plurality of chiplets, where a chiplet of the plurality of chiplets comprise a memory interface, processing resources configured to execute threads of a kernel, and thread dispatch circuitry to facilitate dispatch of threads of the kernel to the processing resources. The processing resources are configured to execute threads of a first kernel, receive dispatch of threads of a second kernel for execution before completion of the first kernel as threads of the first kernel retire, execute a first phase of the second kernel during completion of execution of the first kernel, via a thread of the first kernel, signal an event via an uncached write to a global memory, and execute a second phase of the second kernel based on detection of the event via an uncached read from the global memory.Type: ApplicationFiled: February 20, 2025Publication date: September 11, 2025Applicant: Intel CorporationInventors: Jeffery S. Boles, Altug Koker, Deepak Vembar, Lakshminarayanan Striramassarma, Michal Mrozek, Pierre Boudier, Oren Kaidar
-
Publication number: 20250265762Abstract: A system that includes a graphics processing unit (GPU) comprising multiple processors and circuitry to: parse a first queue of the multiple queues; at an arbitration point in the first queue, select a second queue of the multiple queues to parse based on a priority level of the second queue and a head of line blocking condition of the second queue; and based on identification of a thread spawning instruction, enqueue the thread spawning instruction for execution by at least one processor of the multiple processors.Type: ApplicationFiled: February 16, 2024Publication date: August 21, 2025Inventors: Michal MROZEK, Pierre BOUDIER, Jeffery S. BOLES, AMAN, Vasanth RANGANATHAN, Aditya NAVALE, William DAMON, Rebecca DAVID, Hema C. NALLURI, Antonio VALLES
-
Publication number: 20230290035Abstract: A system and method for performing graphics processing is provided. The system and method includes processing an allocation command for a buffer object; reserving processor address space for a data store of the buffer object with uncommitted physical memory in response to the allocation command including a null parameter, and reserving processor address space for a data store of the buffer object with committed physical memory in response to the allocation command including a non-null parameter.Type: ApplicationFiled: May 19, 2023Publication date: September 14, 2023Applicant: Advanced Micro Devices, Inc.Inventors: Graham Sellers, Eric Zolnowski, Pierre Boudier, Juraj Obert
-
Patent number: 11676321Abstract: A method and system for performing graphics processing is provided. The method and system includes storing stencil buffer values in a stencil buffer; generating either or both of a reference value and a source value in a fragment shader; comparing the stencil buffer values against the reference value; and processing a fragment based on the comparing the stencil buffer values against the reference value.Type: GrantFiled: June 29, 2020Date of Patent: June 13, 2023Assignee: Advanced Micro Devices, Inc.Inventors: Graham Sellers, Eric Zolnowski, Pierre Boudier, Juraj Obert
-
Patent number: 10909739Abstract: In various embodiments, a parallel processor implements a graphics processing pipeline that generates rendered images. In operation, the parallel processor causes execution threads to execute a task shading program on an input mesh to generate a task shader output specifying a mesh shader count. The parallel processor then generates mesh shader identifiers, where the total number of the mesh shader identifiers equals the mesh shader count. For each mesh shader identifier, the parallel processor invokes a mesh shader based on the mesh shader identifier and the task shader output to generate geometry associated with the mesh shader identifier. Subsequently, the parallel processor performs operations on the geometries associated with the mesh shader identifiers to generate a rendered image. Advantageously, unlike conventional graphics processing pipelines, the performance of the graphics processing pipeline is not limited by a primitive distributor.Type: GrantFiled: January 26, 2018Date of Patent: February 2, 2021Assignee: NVIDIA CorporationInventors: Ziyad Hakura, Yury Uralsky, Christoph Kubisch, Pierre Boudier, Henry Moreton
-
Patent number: 10878611Abstract: In various embodiments, a deduplication application pre-processes index buffers for a graphics processing pipeline that generates rendered images via a shading program. In operation, the deduplication application causes execution threads to identify a set of unique vertices specified in an index buffer based on an instruction. The deduplication application then generates a vertex buffer and an indirect index buffer based on the set of unique vertices. The vertex buffer and the indirect index buffer are associated with a portion of an input mesh. The graphics processing pipeline then renders a first frame and a second frame based on the vertex buffer, the indirect index buffer, and the shading program. Advantageously, the graphics processing pipeline may re-use the vertex buffer and indirect index buffer until the topology of the input mesh changes.Type: GrantFiled: January 26, 2018Date of Patent: December 29, 2020Assignee: NVIDIA CorporationInventors: Ziyad Hakura, Yury Uralsky, Christoph Kubisch, Pierre Boudier, Henry Moreton
-
Publication number: 20200327715Abstract: A method and system for performing graphics processing is provided. The method and system includes storing stencil buffer values in a stencil buffer; generating either or both of a reference value and a source value in a fragment shader; comparing the stencil buffer values against the reference value; and processing a fragment based on the comparing the stencil buffer values against the reference value.Type: ApplicationFiled: June 29, 2020Publication date: October 15, 2020Applicant: Advanced Micro Devices, Inc.Inventors: Graham Sellers, Eric Zolnowski, Pierre Boudier, Juraj Obert
-
Patent number: 10699464Abstract: Methods for enabling graphics features in processors are described herein. Methods are provided to enable trinary built-in functions in the shader, allow separation of the graphics processor's address space from the requirement that all textures must be physically backed, enable use of a sparse buffer allocated in virtual memory, allow a reference value used for stencil test to be generated and exported from a fragment shader, provide support for use specific operations in the stencil buffers, allow capture of multiple transform feedback streams, allow any combination of streams for rasterization, allow a same set of primitives to be used with multiple transform feedback streams as with a single stream, allow rendering to be directed to layered framebuffer attachments with only a vertex and fragment shader present, and allow geometry to be directed to one of an array of several independent viewport rectangles without a geometry shader.Type: GrantFiled: October 31, 2016Date of Patent: June 30, 2020Assignee: Advanced Micro Devices, Inc.Inventors: Graham Sellers, Eric Zolnowski, Pierre Boudier, Juraj Obert
-
Patent number: 10600229Abstract: In various embodiments, a parallel processor implements a graphics processing pipeline that generates rendered images via a shading program. In operation, the parallel processor causes a first set of execution threads to execute the shading program on a first portion of the input mesh to generate first geometry stored in an on-chip memory. The parallel processor also causes a second set of execution threads to execute the mesh shading program on a second portion of the input mesh to generate second geometry stored in the on-chip memory. Subsequently, the parallel processor reads the first geometry and the second geometry from the on-chip memory, and performs operations on the first geometry and the second geometry to generate a rendered image derived from the input mesh. Advantageously, unlike conventional graphics processing pipelines, the performance of the graphics processing pipeline is not limited by a primitive distributor.Type: GrantFiled: January 26, 2018Date of Patent: March 24, 2020Assignee: NVIDIA CorporationInventors: Ziyad Hakura, Yury Uralsky, Christoph Kubisch, Pierre Boudier, Henry Moreton
-
Publication number: 20190236829Abstract: In various embodiments, a deduplication application pre-processes index buffers for a graphics processing pipeline that generates rendered images via a shading program. In operation, the deduplication application causes execution threads to identify a set of unique vertices specified in an index buffer based on an instruction. The deduplication application then generates a vertex buffer and an indirect index buffer based on the set of unique vertices. The vertex buffer and the indirect index buffer are associated with a portion of an input mesh. The graphics processing pipeline then renders a first frame and a second frame based on the vertex buffer, the indirect index buffer, and the shading program. Advantageously, the graphics processing pipeline may re-use the vertex buffer and indirect index buffer until the topology of the input mesh changes.Type: ApplicationFiled: January 26, 2018Publication date: August 1, 2019Inventors: Ziyad HAKURA, Yury URALSKY, Christoph KUBISCH, Pierre BOUDIER, Henry MORETON
-
Publication number: 20190236828Abstract: In various embodiments, a parallel processor implements a graphics processing pipeline that generates rendered images. In operation, the parallel processor causes execution threads to execute a task shading program on an input mesh to generate a task shader output specifying a mesh shader count. The parallel processor then generates mesh shader identifiers, where the total number of the mesh shader identifiers equals the mesh shader count. For each mesh shader identifier, the parallel processor invokes a mesh shader based on the mesh shader identifier and the task shader output to generate geometry associated with the mesh shader identifier. Subsequently, the parallel processor performs operations on the geometries associated with the mesh shader identifiers to generate a rendered image. Advantageously, unlike conventional graphics processing pipelines, the performance of the graphics processing pipeline is not limited by a primitive distributor.Type: ApplicationFiled: January 26, 2018Publication date: August 1, 2019Inventors: Ziyad HAKURA, Yury URALSKY, Christoph KUBISCH, Pierre BOUDIER, Henry MORETON
-
Publication number: 20190236827Abstract: In various embodiments, a parallel processor implements a graphics processing pipeline that generates rendered images via a shading program. In operation, the parallel processor causes a first set of execution threads to execute the shading program on a first portion of the input mesh to generate first geometry stored in an on-chip memory. The parallel processor also causes a second set of execution threads to execute the mesh shading program on a second portion of the input mesh to generate second geometry stored in the on-chip memory. Subsequently, the parallel processor reads the first geometry and the second geometry from the on-chip memory, and performs operations on the first geometry and the second geometry to generate a rendered image derived from the input mesh. Advantageously, unlike conventional graphics processing pipelines, the performance of the graphics processing pipeline is not limited by a primitive distributor.Type: ApplicationFiled: January 26, 2018Publication date: August 1, 2019Inventors: Ziyad HAKURA, Yury URALSKY, Christoph KUBISCH, Pierre BOUDIER, Henry MORETON
-
Patent number: 10019829Abstract: Methods for enabling graphics features in processors are described herein. Methods are provided to enable trinary built-in functions in the shader, allow separation of the graphics processor's address space from the requirement that all textures must be physically backed, enable use of a sparse buffer allocated in virtual memory, allow a reference value used for stencil test to be generated and exported from a fragment shader, provide support for use specific operations in the stencil buffers, allow capture of multiple transform feedback streams, allow any combination of streams for rasterization, allow a same set of primitives to be used with multiple transform feedback streams as with a single stream, allow rendering to be directed to layered framebuffer attachments with only a vertex and fragment shader present, and allow geometry to be directed to one of an array of several independent viewport rectangles without a geometry shader.Type: GrantFiled: June 7, 2013Date of Patent: July 10, 2018Assignee: Advanced Micro Devices, Inc.Inventors: Graham Sellers, Pierre Boudier, Juraj Obert
-
Patent number: 9830163Abstract: Methods, apparatuses, and computer readable media are disclosed for control flow on a heterogeneous computer system. The method may include a first processor of a first type, for example a CPU, requesting a first kernel be executed on a second processor of a second type, for example a GPU, to process first work items. The method may include the GPU executing the first kernel to process the first work items. The first kernel may generate second work items. The GPU may execute a second kernel to process the generated second work items. The GPU may dispatch producer kernels when space is available in a work buffer. The GPU may dispatch consumer kernels to process work items in the work buffer when the work buffer has available work items. The GPU may be configured to determine a number of processing elements to execute the first kernel and the second kernel.Type: GrantFiled: June 7, 2013Date of Patent: November 28, 2017Assignee: Advanced Micro Devices, Inc.Inventor: Pierre Boudier
-
Publication number: 20170061670Abstract: Methods for enabling graphics features in processors are described herein. Methods are provided to enable trinary built-in functions in the shader, allow separation of the graphics processor's address space from the requirement that all textures must be physically backed, enable use of a sparse buffer allocated in virtual memory, allow a reference value used for stencil test to be generated and exported from a fragment shader, provide support for use specific operations in the stencil buffers, allow capture of multiple transform feedback streams, allow any combination of streams for rasterization, allow a same set of primitives to be used with multiple transform feedback streams as with a single stream, allow rendering to be directed to layered framebuffer attachments with only a vertex and fragment shader present, and allow geometry to be directed to one of an array of several independent viewport rectangles without a geometry shader.Type: ApplicationFiled: October 31, 2016Publication date: March 2, 2017Applicant: Advanced Micro Devices, Inc.Inventors: Graham Sellers, Eric Zolnowski, Pierre Boudier, Juraj Obert