Patents by Inventor Ching-Yu Hung

Ching-Yu Hung has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Publication number: 20260203122
    Abstract: In various examples, a VPU and associated components may be optimized to improve VPU performance and throughput. For example, the VPU may include a min/max collector, automatic store predication functionality, a SIMD data path organization that allows for inter-lane sharing, a transposed load/store with stride parameter functionality, a load with permute and zero insertion functionality, hardware, logic, and memory layout functionality to allow for two point and two by two point lookups, and per memory bank load caching capabilities. In addition, decoupled accelerators may be used to offload VPU processing tasks to increase throughput and performance, and a hardware sequencer may be included in a DMA system to reduce programming complexity of the VPU and the DMA system. The DMA and VPU may execute a VPU configuration mode that allows the VPU and DMA to operate without a processing controller for performing dynamic region based data movement operations.
    Type: Application
    Filed: January 7, 2026
    Publication date: July 16, 2026
    Applicant: NVIDIA Corporation
    Inventors: Ravi P. Singh, Ching-Yu Hung, Jagadeesh Sankaran, Ahmad Itani, Yen-Te Shih
  • Publication number: 20260129150
    Abstract: Approaches presented herein provide for generation of alternate views from disparity data captured for one or more objects in a scene. The generation can be performed using an embedded processor with DMA memory access, or other limited capacity hardware. An intermediate representation can be generated that is a 2D histogram view of the disparity data. This intermediate representation can be transformed, using the embedded processor, to an alternate view image, such as a bird's eye view image. Morphological or similar filtering can be performed on the one or more objects in the intermediate representation using the same size filter, regardless of distance from a camera plane used to capture the disparity data.
    Type: Application
    Filed: November 5, 2024
    Publication date: May 7, 2026
    Inventors: Branislav Kisacanin, Ching-Yu Hung
  • Publication number: 20260129149
    Abstract: Approaches presented herein provide for generation of alternate views from disparity data captured for one or more objects in a scene. The generation can be performed using an embedded processor with DMA memory access, or other limited capacity hardware. An intermediate representation can be generated that is a 2D histogram view of the disparity data. This intermediate representation can be transformed, using the embedded processor, to an alternate view image, such as a bird's eye view image. Morphological or similar filtering can be performed on the one or more objects in the intermediate representation using the same size filter, regardless of distance from a camera plane used to capture the disparity data.
    Type: Application
    Filed: November 5, 2024
    Publication date: May 7, 2026
    Inventors: Branislav Kisacanin, Ching-Yu Hung
  • Patent number: 12602244
    Abstract: In various examples, a VPU and associated components may be optimized to improve VPU performance and throughput. For example, the VPU may include a min/max collector, automatic store predication functionality, a SIMD data path organization that allows for inter-lane sharing, a transposed load/store with stride parameter functionality, a load with permute and zero insertion functionality, hardware, logic, and memory layout functionality to allow for two point and two by two point lookups, and per memory bank load caching capabilities. In addition, decoupled accelerators may be used to offload VPU processing tasks to increase throughput and performance, and a hardware sequencer may be included in a DMA system to reduce programming complexity of the VPU and the DMA system. The DMA and VPU may execute a VPU configuration mode that allows the VPU and DMA to operate without a processing controller for performing dynamic region based data movement operations.
    Type: Grant
    Filed: August 2, 2021
    Date of Patent: April 14, 2026
    Assignee: NVIDIA Corporation
    Inventors: Ravi P Singh, Ching-Yu Hung, Jagadeesh Sankaran, Ahmad Itani, Yen-Te Shih
  • Patent number: 12572387
    Abstract: In various examples, a VPU and associated components may be optimized to improve VPU performance and throughput. For example, the VPU may include a min/max collector, automatic store predication functionality, a SIMD data path organization that allows for inter-lane sharing, a transposed load/store with stride parameter functionality, a load with permute and zero insertion functionality, hardware, logic, and memory layout functionality to allow for two point and two by two point lookups, and per memory bank load caching capabilities. In addition, decoupled accelerators may be used to offload VPU processing tasks to increase throughput and performance, and a hardware sequencer may be included in a DMA system to reduce programming complexity of the VPU and the DMA system. The DMA and VPU may execute a VPU configuration mode that allows the VPU and DMA to operate without a processing controller for performing dynamic region based data movement operations.
    Type: Grant
    Filed: October 17, 2023
    Date of Patent: March 10, 2026
    Assignee: NVIDIA Corporation
    Inventors: Ravi P. Singh, Ching-Yu Hung, Jagadeesh Sankaran, Ahmad Itani, Yen-Te Shih
  • Publication number: 20260051160
    Abstract: Aspects of this technical solution can increase speed of processing in low-latency application areas, while maintaining integrity of image feature recognition at those higher speeds. For example, in image-processing environments associated with autonomous navigation (e.g., driving), a large volume of image data is to be rapidly and accurately processed to maintain reliable and up-to-date models of a physical environment. For example, embodiments in accordance with this disclosure can provide high-speed and accurate image feature recognition of input frame data beyond the capability of CPU processing or general GPU processing to achieve.
    Type: Application
    Filed: August 28, 2024
    Publication date: February 19, 2026
    Applicant: NVIDIA Corporation
    Inventors: Yi LU, Ching-Yu HUNG, Hanjie MEI, Yen-Te SHIH, Chao LYU
  • Publication number: 20260051013
    Abstract: Aspects of this technical solution can increase speed of processing in low-latency application areas, while maintaining integrity of image feature recognition at those higher speeds. For example, in image-processing environments associated with autonomous navigation (e.g., driving), a large volume of image data is to be rapidly and accurately processed to maintain reliable and up-to-date models of a physical environment. For example, embodiments in accordance with this disclosure can provide high-speed and accurate image feature recognition of input frame data beyond the capability of CPU processing or general GPU processing to achieve.
    Type: Application
    Filed: August 28, 2024
    Publication date: February 19, 2026
    Applicant: NVIDIA Corporation
    Inventors: Yi LU, Ching-Yu HUNG, Hanjie MEI, Yen-Te SHIH, Chao LYU
  • Publication number: 20260051161
    Abstract: Aspects of this technical solution can increase speed of processing in low-latency application areas, while maintaining integrity of image feature recognition at those higher speeds. For example, in image-processing environments associated with autonomous navigation (e.g., driving), a large volume of image data is to be rapidly and accurately processed to maintain reliable and up-to-date models of a physical environment. For example, embodiments in accordance with this disclosure can provide high-speed and accurate image feature recognition of input frame data beyond the capability of CPU processing or general GPU processing to achieve.
    Type: Application
    Filed: August 28, 2024
    Publication date: February 19, 2026
    Applicant: NVIDIA Corporation
    Inventors: Yi LU, Ching-Yu HUNG, Hanjie MEI, Yen-Te SHIH, Chao LYU
  • Publication number: 20260051162
    Abstract: Aspects of this technical solution can increase speed of processing and lower computational hardware complexity in motion detection, while maintaining integrity of motion detection across image frames. For example, in image-processing environments associated with autonomous or semi-autonomous navigation (e.g., driving, robotic navigation, etc.), a large volume of image data is to be rapidly and accurately processed to maintain reliable and up-to-date models of a physical environment. Thus, embodiments in accordance with this disclosure can provide high-speed and accurate motion detection of input frame data.
    Type: Application
    Filed: August 29, 2024
    Publication date: February 19, 2026
    Applicant: NVIDIA Corporation
    Inventors: Hanjie MEI, Chao LYU, Yen-Te SHIH, Admad ITANI, Ching-Yu HUNG
  • Publication number: 20260037478
    Abstract: In various examples, systems and methods are disclosed that relate to programming multi-dimensional single instruction, multiple data (SIMD) processors (also referred to as an accelerator). In one example, a processor can obtain instructions to be performed by the accelerator. The processor can determine one or more operations to be performed by the accelerator based at least on the instructions and generate a set of accelerator instructions. In examples, the processor can then provide data associated with the accelerator instructions to cause the accelerator to perform at least a portion of the one or more operations.
    Type: Application
    Filed: July 31, 2024
    Publication date: February 5, 2026
    Applicant: NVIDIA Corporation
    Inventors: Andrew Peter TAUSSIG, Ravi Pratap SINGH, Ching-Yu HUNG, Sreenivas KRISHNAN, Jagadeesh SANKARAN, Yen-Te SHIH
  • Publication number: 20260037480
    Abstract: In various examples, systems and methods are disclosed that relate to performing inter-accelerator data transfers. For example, an accelerator such as a pixel processing engine (PPE) can include multiple processing engines (PEs). The PEs can be arranged in a two-dimensional array and each PE can be configured to receive data in the registers of the PEs. The PEs can transfer data between registers of the same or different PEs. The PEs can also be configured to perform transfers and operations in sequence to perform complex functions such as filtering.
    Type: Application
    Filed: July 31, 2024
    Publication date: February 5, 2026
    Applicant: NVIDIA Corporation
    Inventors: Ching-Yu Hung, Sreenivas Krishnan, Divya Ojha, Jagadeesh Sankaran, Yen-Te Shih, Ravi Pratap Singh, Andrew Peter Taussig, Arun Visweswaraiah
  • Publication number: 20260038079
    Abstract: In various examples, systems and methods are disclosed relating to coordinating and synchronizing the actions of different types of processors with low latency. Different types of processors may perform better at different types of tasks. By coordinating the processing of a one-dimensional processor such as a vector processing unit (VPU) and the processing of a two-dimensional processor such as a pixel processing engine (PPE), an overall speed of task completion can be improved.
    Type: Application
    Filed: July 31, 2024
    Publication date: February 5, 2026
    Applicant: NVIDIA Corporation
    Inventors: Sreenivas Krishnan, Ching-Yu Hung, Ahmad Itani, Jagadeesh Sankaran, Yen-Te Shih, Ravi Pratap Singh, Andrew Peter Taussig, Jeremy Chan
  • Publication number: 20260037264
    Abstract: Aspects of this technical solution can increase processing speed in low-latency application areas, while maintaining integrity of error correction detection at those higher speeds. For example, in image-processing environments associated with autonomous navigation (e.g., driving), a large volume of image data is to be rapidly and accurately processed to maintain reliable and up-to-date models of a physical environment. Thus, embodiments in accordance with this disclosure can provide high-speed and accurate error correction of input frame data beyond the capability of CPU processing or general GPU processing to achieve.
    Type: Application
    Filed: July 31, 2024
    Publication date: February 5, 2026
    Applicant: NVIDIA Corporation
    Inventors: Lachlan Francis DOWLING, Dawid Stanislaw PAJAK, Ching-Yu HUNG
  • Publication number: 20260037457
    Abstract: Aspects of this technical solution can provide at least a technical improvement to reading and writing data between a memory device and a processor, including, for example, by providing a technical solution to configure one or more load streams with stream sizes configured based on relative speed of a processor and a memory. For example, this technical solution can provide a technical improvement to processing speed of computations by a processor with data obtained from or stored to a memory device. For example, a system in accordance with this technical solution can provide a decoupled load store unit (DLSU) distinct from a processor and a memory device to prefetch a sufficient amount of data from a memory device into a stream buffer of a DLSU, to provide instructions to a processor at a rate that eliminates waiting by the processor for memory over one or more cycles.
    Type: Application
    Filed: July 31, 2024
    Publication date: February 5, 2026
    Applicant: NVIDIA Corporation
    Inventors: Ravi Pratap Singh, Sreenivas Krishnan, Ching-Yu Hung, Jagadeesh Sankaran, Yen-Te Shih, Andrew Peter Taussig
  • Publication number: 20260037330
    Abstract: In various examples, systems and methods are disclosed that relate to processing data using accelerators in a system on a chip. For example, a plurality of processing elements (PEs) can interconnect to form a processing engine, and a control system can control operation of the PEs based at least on the connections between the PEs. In some examples, the PEs can receive sub-inputs and transfer the sub-inputs to one or more other PEs to enable performance of the instructed operations. In examples, once the PEs complete the instructed operations, the sub-inputs can be transferred out of the processing engine.
    Type: Application
    Filed: July 31, 2024
    Publication date: February 5, 2026
    Applicant: NVIDIA Corporation
    Inventors: Ching-Yu Hung, Sreenivas Krishnan, Jagadeesh Sankaran, Yen-Te Shih, Ravi Pratap Singh, Andrew Peter Taussig
  • Publication number: 20250103529
    Abstract: In various examples, a VPU and associated components may be optimized to improve VPU performance and throughput. For example, the VPU may include a min/max collector, automatic store predication functionality, a SIMD data path organization that allows for inter-lane sharing, a transposed load/store with stride parameter functionality, a load with permute and zero insertion functionality, hardware, logic, and memory layout functionality to allow for two point and two by two point lookups, and per memory bank load caching capabilities. In addition, decoupled accelerators may be used to offload VPU processing tasks to increase throughput and performance, and a hardware sequencer may be included in a DMA system to reduce programming complexity of the VPU and the DMA system. The DMA and VPU may execute a VPU configuration mode that allows the VPU and DMA to operate without a processing controller for performing dynamic region based data movement operations.
    Type: Application
    Filed: December 5, 2024
    Publication date: March 27, 2025
    Inventors: Ahmad Itani, Yen-Te Shih, Jagadeesh Sankaran, Ravi P. Singh, Ching-Yu Hung
  • Patent number: 12204475
    Abstract: In various examples, a VPU and associated components may be optimized to improve VPU performance and throughput. For example, the VPU may include a min/max collector, automatic store predication functionality, a SIMD data path organization that allows for inter-lane sharing, a transposed load/store with stride parameter functionality, a load with permute and zero insertion functionality, hardware, logic, and memory layout functionality to allow for two point and two by two point lookups, and per memory bank load caching capabilities. In addition, decoupled accelerators may be used to offload VPU processing tasks to increase throughput and performance, and a hardware sequencer may be included in a DMA system to reduce programming complexity of the VPU and the DMA system. The DMA and VPU may execute a VPU configuration mode that allows the VPU and DMA to operate without a processing controller for performing dynamic region based data movement operations.
    Type: Grant
    Filed: December 9, 2022
    Date of Patent: January 21, 2025
    Assignee: NVIDIA Corporation
    Inventors: Ahmad Itani, Yen-Te Shih, Jagadeesh Sankaran, Ravi P Singh, Ching-Yu Hung
  • Patent number: 12118353
    Abstract: In various examples, a VPU and associated components may be optimized to improve VPU performance and throughput. For example, the VPU may include a min/max collector, automatic store predication functionality, a SIMD data path organization that allows for inter-lane sharing, a transposed load/store with stride parameter functionality, a load with permute and zero insertion functionality, hardware, logic, and memory layout functionality to allow for two point and two by two point lookups, and per memory bank load caching capabilities. In addition, decoupled accelerators may be used to offload VPU processing tasks to increase throughput and performance, and a hardware sequencer may be included in a DMA system to reduce programming complexity of the VPU and the DMA system. The DMA and VPU may execute a VPU configuration mode that allows the VPU and DMA to operate without a processing controller for performing dynamic region based data movement operations.
    Type: Grant
    Filed: August 2, 2021
    Date of Patent: October 15, 2024
    Assignee: NVIDIA Corporation
    Inventors: Ching-Yu Hung, Ravi P Singh, Jagadeesh Sankaran, Yen-Te Shih, Ahmad Itani
  • Patent number: 12099439
    Abstract: In various examples, a VPU and associated components may be optimized to improve VPU performance and throughput. For example, the VPU may include a min/max collector, automatic store predication functionality, a SIMD data path organization that allows for inter-lane sharing, a transposed load/store with stride parameter functionality, a load with permute and zero insertion functionality, hardware, logic, and memory layout functionality to allow for two point and two by two point lookups, and per memory bank load caching capabilities. In addition, decoupled accelerators may be used to offload VPU processing tasks to increase throughput and performance, and a hardware sequencer may be included in a DMA system to reduce programming complexity of the VPU and the DMA system. The DMA and VPU may execute a VPU configuration mode that allows the VPU and DMA to operate without a processing controller for performing dynamic region based data movement operations.
    Type: Grant
    Filed: August 2, 2021
    Date of Patent: September 24, 2024
    Assignee: NVIDIA Corporation
    Inventors: Ching-Yu Hung, Ravi P Singh, Jagadeesh Sankaran, Yen-Te Shih, Ahmad Itani
  • Patent number: 12093539
    Abstract: In various examples, a VPU and associated components may be optimized to improve VPU performance and throughput. For example, the VPU may include a min/max collector, automatic store predication functionality, a SIMD data path organization that allows for inter-lane sharing, a transposed load/store with stride parameter functionality, a load with permute and zero insertion functionality, hardware, logic, and memory layout functionality to allow for two point and two by two point lookups, and per memory bank load caching capabilities. In addition, decoupled accelerators may be used to offload VPU processing tasks to increase throughput and performance, and a hardware sequencer may be included in a DMA system to reduce programming complexity of the VPU and the DMA system. The DMA and VPU may execute a VPU configuration mode that allows the VPU and DMA to operate without a processing controller for performing dynamic region based data movement operations.
    Type: Grant
    Filed: December 21, 2022
    Date of Patent: September 17, 2024
    Assignee: NVIDIA Corporation
    Inventors: Ching-Yu Hung, Ravi P Singh, Jagadeesh Sankaran, Yen-Te Shih, Ahmad Itani