Patents by Inventor Prashant Laddha
Prashant Laddha has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Publication number: 20260086769Abstract: Techniques for efficient multi-dimensional data processing. For example, front end circuitry sorts tuples across a plurality of tuple buffers to provide conflict-free access to a corresponding plurality of input data banks without memory access conflicts, each tuple to associate an input data element of a plurality of input data elements with a corresponding weight data element of a weight tensor and a corresponding output data element of an output tensor; and execution circuitry to perform multiply-accumulate operations using a subset of the tuples, the execution circuitry to perform parallel multiplications with a corresponding subset of input data elements of the plurality of input data elements and a corresponding subset of weight data elements of the weight tensor indicated by the subset of the tuples, the execution circuitry to access the subset of input data elements from different input data banks of the plurality of input data banks without memory conflicts.Type: ApplicationFiled: December 2, 2025Publication date: March 26, 2026Applicant: Intel CorporationInventors: Deepali Garg, Baishik Biswas, Prashant Laddha, Om Ji Omer, Sreenivas Subramoney
-
Publication number: 20260082065Abstract: Real-time neural video codecs face significant latency and energy bottlenecks due to pixel-level grid sampling, which requires irregular, fine-grained memory accesses and limits efficient hardware acceleration. To address this, a sub-tile-based grid sampling technique is disclosed herein. The technique determines super tile sizes using motion vector gradients, neural network parameters, and available on-chip memory. A super tile is split into sub-tiles by detecting motion boundaries through motion vector analysis, where a sub-tile has homogeneous motion vectors. For each sub-tile, a reference bounding box is computed to enable efficient block transfers of reference data, and per-pixel metadata is generated for feature interpolation. The pipelined, parallelizable solution reduces number of memory accesses and computational overhead, compared to existing pixel-based techniques.Type: ApplicationFiled: November 24, 2025Publication date: March 19, 2026Applicant: Intel CorporationInventors: Prashant Laddha, Om Ji Omer, Arnab Raha, Deepak Abraham Mathaikutty
-
Patent number: 12487856Abstract: An embodiment of an apparatus comprises a hardware accelerator to perform a three-dimensional (3D) point cloud data access operation, and circuitry coupled to the hardware accelerator to control the hardware accelerator to perform the 3D point cloud data access operation in response to a request. Other embodiments are disclosed and claimed.Type: GrantFiled: October 25, 2022Date of Patent: December 2, 2025Assignee: Intel CorporationInventors: Gurpreet S. Kalsi, Om Ji Omer, Prashant Laddha, Kamlesh R. Pillai, Anirud Thyagharajan, Meenal Kudalkar, Krishnan Ananthanarayanan, Sreenivas Subramoney
-
Publication number: 20250308074Abstract: Corner table creation for graphics processing is described. An example of an apparatus includes corner data generator circuitry, including a circuit to generate a plurality of vertex-corner lists for a plurality of portions of a triangle mesh, wherein a vertex-corner list for a portion includes an index for a vertex, a count of corners for the vertex, and a list of corners associated with the vertex, and a circuit to receive the vertex-corner lists and generate one or more edge hash maps, each edge hash map including corner indices for edges of the triangle mesh.Type: ApplicationFiled: March 29, 2024Publication date: October 2, 2025Applicant: Intel CorporationInventors: Prashant Laddha, Kamlesh Pillai, Om Ji Omer
-
Publication number: 20250021819Abstract: Systems, apparatus, articles of manufacture, and methods for quality and capacity-aware grouped query attention are disclosed. To accomplish such groupings, example instructions cause a machine to create a plurality of groups of query heads present in a key value cache using an evolutionary algorithm based on at least two objectives, quantify an amount of error introduced by a first group of query heads in the plurality of groups of query heads, and retain the query heads of the first group of query heads in a non-grouped arrangement when the error meets an error threshold.Type: ApplicationFiled: September 27, 2024Publication date: January 16, 2025Applicant: Intel CorporationInventors: Vinay Joshi, Om Ji Omer, Prashant Laddha, Shambhavi Sinha
-
Patent number: 12189559Abstract: Exemplary embodiments maintain spatial locality of the data being processed by a sparse CNN. The spatial locality is maintained by reordering the data to preserve spatial locality. The reordering may be performed on data elements and on data for groups of co-located data elements referred to herein as “chunks”. Thus, the data may be reordered into chunks, where each chunk contains data for spatially co-located data elements, and in addition, chunks may be organized so that spatially located chunks are together. The use of chunks helps to reduce the need to re-fetch data during processing. Chunk sizes may be chosen based on the memory constraints of the processing logic (e.g., cache sizes).Type: GrantFiled: June 26, 2020Date of Patent: January 7, 2025Assignee: Intel CorporationInventors: Anirud Thyagharajan, Prashant Laddha, Om Omer, Sreenivas Subramoney
-
Patent number: 11875555Abstract: A computer model is trained to classify regions of a space (e.g., a pixel of an image or a voxel of a point cloud) according to a multi-label classification. To improve the model's accuracy, the model's self-confidence is determined with respect to its own predictions of regions in a training space. The self-confidence is determined based on the class predictions, such as a difference between the highest-predicted class and a second-highest-predicted class. When these are similar, it may reflect areas for potential improvement by focusing training on these low-confidence areas. Additional training may be performed by including modified training data in subsequent training iterations that focuses on low-confidence areas. As another example, additional training may be performed using the self-confidence to modify a classification loss used to refine parameters of the model.Type: GrantFiled: November 24, 2021Date of Patent: January 16, 2024Assignee: Intel CorporationInventors: Anirud Thyagharajan, Prashant Laddha, Benjamin Ummenhofer, Om Ji Omer
-
Patent number: 11783170Abstract: Systems, apparatuses and methods may provide for technology that decodes data via an instruction that indicates a number of rulebooks to be processed, an input feature size, an output feature size, and a plurality of feature map base addresses, rearranges spatially distributed voxel output feature maps in the decoded data based on weight planes, and performs a channel-wise multiply-accumulate (MAC) operation on the rearranged spatially distributed voxel output feature maps to obtain an output, wherein the channel-wise MAC operation is performed as partial accumulations by a plurality of processing elements.Type: GrantFiled: January 25, 2023Date of Patent: October 10, 2023Assignee: INTEL CORPORATIONInventors: Kamlesh Pillai, Gurpreet Singh Kalsi, Sreenivas Subramoney, Prashant Laddha, Om Ji Omer
-
Publication number: 20230169319Abstract: Systems, apparatuses and methods may provide for technology that decodes data via an instruction that indicates a number of rulebooks to be processed, an input feature size, an output feature size, and a plurality of feature map base addresses, rearranges spatially distributed voxel output feature maps in the decoded data based on weight planes, and performs a channel-wise multiply-accumulate (MAC) operation on the rearranged spatially distributed voxel output feature maps to obtain an output, wherein the channel-wise MAC operation is performed as partial accumulations by a plurality of processing elements.Type: ApplicationFiled: January 25, 2023Publication date: June 1, 2023Inventors: Kamlesh Pillai, Gurpreet Singh Kalsi, Sreenivas Subramoney, Prashant Laddha, Om Ji Omer
-
Patent number: 11620818Abstract: Systems, apparatuses and methods may provide for technology that decodes data via an instruction that indicates a number of rulebooks to be processed, an input feature size, an output feature size, and a plurality of feature map base addresses, rearranges spatially distributed voxel output feature maps in the decoded data based on weight planes, and performs a channel-wise multiply-accumulate (MAC) operation on the rearranged spatially distributed voxel output feature maps to obtain an output, wherein the channel-wise MAC operation is performed as partial accumulations by a plurality of processing elements.Type: GrantFiled: December 22, 2020Date of Patent: April 4, 2023Assignee: INTEL CORPORATIONInventors: Kamlesh Pillai, Gurpreet Singh Kalsi, Sreenivas Subramoney, Prashant Laddha, Om Ji Omer
-
Publication number: 20220148311Abstract: Systems, apparatuses and methods may provide for technology that identifies a plurality of segments based on semantic features and instance features associated with a scene, fuses the plurality of segments into a plurality of instances, and selects classification labels for the plurality of instances. In one example, the plurality of segments is fused into the plurality of instances via a learnable self-attention based network.Type: ApplicationFiled: January 24, 2022Publication date: May 12, 2022Inventors: Anirud Thyagharajan, Prashant Laddha, Benjamin Ummenhofer, Om Ji Omer
-
Publication number: 20220084310Abstract: A computer model is trained to classify regions of a space (e.g., a pixel of an image or a voxel of a point cloud) according to a multi-label classification. To improve the model's accuracy, the model's self-confidence is determined with respect to its own predictions of regions in a training space. The self-confidence is determined based on the class predictions, such as a difference between the highest-predicted class and a second-highest-predicted class. When these are similar, it may reflect areas for potential improvement by focusing training on these low-confidence areas. Additional training may be performed by including modified training data in subsequent training iterations that focuses on low-confidence areas. As another example, additional training may be performed using the self-confidence to modify a classification loss used to refine parameters of the model.Type: ApplicationFiled: November 24, 2021Publication date: March 17, 2022Applicant: Intel CorporationInventors: Anirud Thyagharajan, Prashant Laddha, Benjamin Ummenhofer, Om Ji Omer
-
Patent number: 11238309Abstract: An example apparatus for selecting keypoints in image includes a keypoint detector to detect keypoints in a plurality of received images. The apparatus also includes a score calculator to calculate a keypoint score for each of the detected keypoints based on a descriptor score indicating descriptor invariance. The apparatus includes a keypoint selector to select keypoints based on the calculated keypoint scores. The apparatus also further includes a descriptor calculator to calculate descriptors for each of the selected keypoints. The apparatus also includes a descriptor matcher to match corresponding descriptors between images in the plurality of received images. The apparatus further also includes a feature tracker to track a feature in the plurality of images based on the matched descriptors.Type: GrantFiled: December 26, 2018Date of Patent: February 1, 2022Assignee: Intel CorporationInventors: Dipan Kumar Mandal, Gurpreet Kalsi, Om J Omer, Prashant Laddha, Sreenivas Subramoney
-
Patent number: 11189000Abstract: An embodiment of an image processor device includes technology to fetch a feature point data set from outside a local memory, locally store three or more fetched feature point data sets in the local memory, compute orientation information for each fetched feature point data set, compute first descriptor information based on the computed orientation information and a first locally stored feature point data set in parallel with a fetch and local store of a second feature point data set in the local memory, and compute second descriptor information based on the computed orientation information and the second locally stored feature point data set in parallel with the compute of the first descriptor information. Other embodiments are disclosed and claimed.Type: GrantFiled: June 24, 2019Date of Patent: November 30, 2021Assignee: Intel CorporationInventors: Gopi Neela, Dipan Kumar Mandal, Gurpreet S. Kalsi, Prashant Laddha, Om J. Omer, Anirud Thyagharajan, Srivatsava Jandhyala
-
Publication number: 20210110187Abstract: Systems, apparatuses and methods may provide for technology that decodes data via an instruction that indicates a number of rulebooks to be processed, an input feature size, an output feature size, and a plurality of feature map base addresses, rearranges spatially distributed voxel output feature maps in the decoded data based on weight planes, and performs a channel-wise multiply-accumulate (MAC) operation on the rearranged spatially distributed voxel output feature maps to obtain an output, wherein the channel-wise MAC operation is performed as partial accumulations by a plurality of processing elements.Type: ApplicationFiled: December 22, 2020Publication date: April 15, 2021Inventors: Kamlesh Pillai, Gurpreet Singh Kalsi, Sreenivas Subramoney, Prashant Laddha, Om Ji Omer
-
Publication number: 20210090328Abstract: Systems, apparatuses and methods provide technology for optimizing processing of sparse data, such as 3D pointcloud data sets. The technology may include generating a locality-aware rulebook based on an input unstructured sparse data set, such as a 3D pointcloud data set, the locality-aware rulebook storing spatial neighborhood information for active voxels in the input unstructured sparse data set, computing an average receptive field (ARF) value based on the locality aware rulebook, and determining, from a plurality of tile size and loop order combinations, a tile size and loop order combination for processing the unstructured sparse data based on the computed ARF value. The technology may also include providing the locality-aware rulebook and the tile size and loop order combination to a compute engine such as a neural network, the compute engine to process the unstructured sparse data using the locality aware rulebook and the tile size and loop order combination.Type: ApplicationFiled: December 7, 2020Publication date: March 25, 2021Inventors: Prashant Laddha, Anirud Thyagharajan, Om Ji Omer, Sreenivas Subramoney
-
Publication number: 20200327396Abstract: Exemplary embodiments maintain spatial locality of the data being processed by a sparse CNN. The spatial locality is maintained by reordering the data to preserve spatial locality. The reordering may be performed on data elements and on data for groups of co-located data elements referred to herein as “chunks”. Thus, the data may be reordered into chunks, where each chunk contains data for spatially co-located data elements, and in addition, chunks may be organized so that spatially located chunks are together. The use of chunks helps to reduce the need to re-fetch data during processing. Chunk sizes may be chosen based on the memory constraints of the processing logic (e.g., cache sizes).Type: ApplicationFiled: June 26, 2020Publication date: October 15, 2020Applicant: Intel CorporationInventors: Anirud Thyagharajan, Prashant Laddha, Om Omer, Sreenivas Subramoney
-
Publication number: 20190333183Abstract: An embodiment of an image processor device includes technology to fetch a feature point data set from outside a local memory, locally store three or more fetched feature point data sets in the local memory, compute orientation information for each fetched feature point data set, compute first descriptor information based on the computed orientation information and a first locally stored feature point data set in parallel with a fetch and local store of a second feature point data set in the local memory, and compute second descriptor information based on the computed orientation information and the second locally stored feature point data set in parallel with the compute of the first descriptor information. Other embodiments are disclosed and claimed.Type: ApplicationFiled: June 24, 2019Publication date: October 31, 2019Applicant: Intel CorporationInventors: Gopi Neela, Dipan Kumar Mandal, Gurpreet S. Kalsi, Prashant Laddha, Om J. Omer, Anirud Thyagharajan, Srivatsava Jandhyala
-
Publication number: 20190171909Abstract: An example apparatus for selecting keypoints in image includes a keypoint detector to detect keypoints in a plurality of received images. The apparatus also includes a score calculator to calculate a keypoint score for each of the detected keypoints based on a descriptor score indicating descriptor invariance. The apparatus includes a keypoint selector to select keypoints based on the calculated keypoint scores. The apparatus also further includes a descriptor calculator to calculate descriptors for each of the selected keypoints. The apparatus also includes a descriptor matcher to match corresponding descriptors between images in the plurality of received images. The apparatus further also includes a feature tracker to track a feature in the plurality of images based on the matched descriptors.Type: ApplicationFiled: December 26, 2018Publication date: June 6, 2019Applicant: INTEL CORPORATIONInventors: Dipan Kumar Mandal, Gurpreet Kalsi, Om J. Omer, Prashant Laddha, Sreenivas Subramoney
-
Patent number: 9445049Abstract: Presented herein are techniques for detecting whether a presentation video stream in a conference call includes motion video. In an embodiment, this can be done by detecting the presence of audio. The presence of audio suggests that the presentation video stream may include motion video. If audio is not detected, then the presentation video stream is encoded at a first frame rate. If audio is detected, then the presentation video stream is encoded at a second, higher frame rate to accommodate the motion video. The higher frame rate allows for a better viewing experience by conference participants.Type: GrantFiled: August 20, 2014Date of Patent: September 13, 2016Assignee: Cisco Technology, Inc.Inventors: Manju Kumari Meghwani, Satheesh Babu Sudarsanan, Prashant Laddha