Patents by Inventor Liu Ren

Liu Ren has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Patent number: 12731081
    Abstract: Methods and systems for training an autonomous driving, agent-centric vison-language planning (VLP) machine learning model. Image data is obtained from a vehicle-mounted camera, encompassing details about agents situated within the external environment. Via image processing, the system identifies these agents within the environment. A Bird's Eye View (BEV) representation of the surroundings is then generated, encapsulating BEV features including spatiotemporal information linked to the vehicle and the recognized agents. Executing the VLP model begins by first extracting agent-wise BEV features from the BEV, wherein the agent-wise BEV features are associated with respective agents in the environment. Agent-wise text features are extracted from natural language text prompts. A contrastive learning model derives similarities between the agent-wise BEV features and the agent-wise text features.
    Type: Grant
    Filed: November 10, 2023
    Date of Patent: September 8, 2026
    Assignee: Robert Bosch GmbH
    Inventors: Chenbin Pan, Burhaneddin Yaman, Tommaso Nesti, Abhirup Mallik, Liu Ren
  • Patent number: 12718548
    Abstract: A method for finetuning a pretrained language model for performing a visual navigation task is described. Scene data is provided that describes a plurality of scenes. Scene graphs that represent the plurality of scenes are derived based on the scene data. For different combinations of a given starting room and target object, a ground truth shortest path from the starting room to the target object in the scene is determined based on the scene graph. Based on the scene data and the scene graphs, natural language prompts are derived that prompt the language model to predict, given a current room and the target object in the scene, a next navigation step of a shortest path from the starting room to the target object in the scene. Together, the natural language prompts and the ground truth shortest paths are used to finetune the pretrained language model.
    Type: Grant
    Filed: July 26, 2024
    Date of Patent: August 25, 2026
    Assignee: Robert Bosch GmbH
    Inventors: Yuliang Guo, Christian Juette, Xinyu Huang, Liu Ren
  • Patent number: 12711740
    Abstract: A method includes encoding a set of hierarchical text prompts to define a set of text embeddings, where the set of hierarchical text prompt defines a primary informative prompt and a secondary informative prompt associated with the primary informative prompt. The method further includes encoding an input image to define a plurality of feature representations, changing a value of one or more identified feature representations among the plurality of feature representations to mask the one or more identified feature representation and define a general feature representation of the input image based on a class-specific threshold indicative of boundary between a class-specific feature and a general feature. The method further includes classifying the input image based on an out-of-distribution (OOD) score determined using a similarity analysis of the general feature representation and the set of text embeddings.
    Type: Grant
    Filed: July 10, 2024
    Date of Patent: August 18, 2026
    Assignee: Robert Bosch GmbH
    Inventors: Sima Behpour, Thang Doan, Xin Li, Wenbin He, Liang Gou, Liu Ren
  • Patent number: 12700184
    Abstract: The disclosure presents an efficient method for improving the aesthetic quality of noisy room segments corresponding to a scanned environment, such as those used by mobile robots to navigate an environment to perform a task. Room polygons are extracted from the noisy room segments and visual scores of the extracted room polygons are assessed. Based on the visual score, low scoring room polygons are removed or merged. A polygon mesh is formed from the room segments, which is then simplified and aligned using efficient methods. Finally, room segments and boundaries between rooms are recovered to generate refined room segments from the polygon mesh. Compared to the noisy room segments, the refined room segments are better suited for visual presentation to users, such as in mobile applications for operating such mobile robots.
    Type: Grant
    Filed: May 14, 2024
    Date of Patent: August 4, 2026
    Assignee: Robert Bosch GmbH
    Inventors: Lincan Zou, Christian Juette, Liu Ren
  • Patent number: 12682630
    Abstract: Methods and systems for Few-Shot Class-Incremental Learning (FSCIL) that utilizes a combination of Session Specific Prompts (SSP) and hyperbolic distance metrics to enhance session-wise learning and representation of image-text pairings across differing classes. The methods and systems include a base training session where both text and image features are projected into hyperbolic space for accurate class pairing using a cross-entropy loss function. Subsequent incremental sessions incorporate previously learned SSPs to retain and augment the separability of classes while minimizing the trainable parameters. This enhances performance in image-text classification tasks by leveraging a minimalistic approach, achieving higher accuracy with fewer trainable parameters compared to traditional models.
    Type: Grant
    Filed: June 7, 2024
    Date of Patent: July 14, 2026
    Assignee: Robert Bosch GmbH
    Inventors: Thang Doan, Sima Behpour, Xin Li, Wenbin He, Liang Gou, Liu Ren
  • Publication number: 20260187139
    Abstract: A computer-implemented method and system relate to digital image retrieval and data curation. The data curation may relate to training a machine learning model on at least one specific task. A vocabulary of visual concepts is generated for a specific task using a target dataset. The vocabulary includes a representative image embedding for each visual concept. Precomputed image embeddings are retrieved from a vector database. Each precomputed image embedding is decomposed into a linear combination of the visual concepts. For each precomputed image embedding, a set of weights is generated based on the vocabulary. Each weight is indicative of a prominence of a respective representative image embedding. The set of weights of each precomputed image embedding is stored in an enhanced vector database. A set of digital images is retrievable from the enhanced vector database in response to a query.
    Type: Application
    Filed: December 9, 2025
    Publication date: July 2, 2026
    Inventors: Xin Li, Clint Sebastian, Frederik Zilly, Wenbin He, Liu Ren
  • Patent number: 12664757
    Abstract: A computer-implemented system and method relates to language-guided self-supervised semantic segmentation. A modified image is generated by performing data augmentation on a source image. A machine learning model generates first pixel embeddings based on the modified image. First segment embeddings are generated using the first pixel embeddings. A pretrained vision-language model generates second pixel embeddings based on the source image. Second segment embeddings are generated by applying segment contour data from the first pixel embeddings to the second pixel embeddings after the data augmentation is performed on the second pixel embeddings. Embedding consistent loss data is generated by comparing the first segment embeddings in relation to the second segment embeddings. Combined loss data is generated that includes the embedding consistent loss data. Parameters of the machine learning model are updated based on the combined loss data.
    Type: Grant
    Filed: May 12, 2023
    Date of Patent: June 23, 2026
    Assignee: Robert Bosch GmbH
    Inventors: Wenbin He, Suphanut Jamonnak, Liang Gou, Liu Ren
  • Patent number: 12657357
    Abstract: Methods and systems for determining a 6D pose of an object in an image are disclosed. In embodiments, an input image is received from a sensor, wherein the input image includes an object in the image. A trained image encoder transforms the input image into a normal map and an instance segmentation map. The normal map is encoded with pointwise 2D features. A 3D CAD model is selected from memory that resembles the object in the image. The 3D CAD model is encoded with pointwise 3D features. The pointwise 2D features are matched with the pointwise 3D features to obtain correspondences between the 2D features and the 3D features. The 6D pose of the object is then determined based on the correspondences.
    Type: Grant
    Filed: January 31, 2022
    Date of Patent: June 16, 2026
    Assignee: Robert Bosch GmbH
    Inventors: Yuliang Guo, Xinyu Huang, Liu Ren
  • Publication number: 20260148536
    Abstract: A method includes splitting an input image into a plurality of patches with each patch corresponding to a distinct region of the input image using a vision transformer. The input image is defined using a model output of a vision model. The method further includes defining a plurality of position embeddings including a position embedding for each of the plurality of patches and for the input image as a whole using the vision transformer, labeling identified regions of the original image based on the estimated loss map to define a labeled image; and outputting a test performance qualifier indicating expected performance of the vision model when the vision model is part of the vision system. The test performance qualifier is calculated using a weighted analysis based on the image loss level and the regional loss level for each patch provided with the labeled image.
    Type: Application
    Filed: November 27, 2024
    Publication date: May 28, 2026
    Inventors: Sanbao Su, Xin Li, Thang Doan, Sima Behpour, Wenbin He, Liang Gou, Liu Ren
  • Publication number: 20260146867
    Abstract: A method of generating high definition (HD) maps for a vehicle based on standard definition (SD) maps includes, using a processor, receiving an SD map corresponding to an environment, receiving one or more aerial images corresponding to the environment, predicting and labeling features contained within the SD map using at least one pairing of the SD map and a respective aerial image of the one or more aerial images, generating a first HD map based on the predicted and labeled features contained within the SD map and the one or more aerial images, and transmitting the first HD map to the vehicle for use in controlling autonomous driving functions.
    Type: Application
    Filed: November 25, 2024
    Publication date: May 28, 2026
    Inventors: David Fernando PAZ RUIZ, Yuliang GUO, Xinyu HUANG, Liu REN
  • Patent number: 12639935
    Abstract: The workflow generally comprises three phases: an Explainable Data Slice-Finding phase, a Slice Summarization and Annotation phase, and a Slice Error Mitigation phase. In the Explainable Data Slice-Finding phase, the workflow employs pixel attributions to create interpretable features of the dataset. In the Slice Summarization and Annotation phase, the workflow transforms the generated features into visualizations including a ‘Data Slice Mosaic.’ Finally, in the Slice Error Mitigation phase, the workflow leverages the annotation and user-verified spuriousness to mitigate slice errors in the machine learning model.
    Type: Grant
    Filed: November 30, 2023
    Date of Patent: May 26, 2026
    Assignee: Robert Bosch GmbH
    Inventors: Xiwei Xuan, Jorge Henrique Piazentin Ono, Liang Gou, Liu Ren
  • Publication number: 20260141479
    Abstract: A computer-implemented method and system relate to an image encoder that receives a digital image as input. The image encoder generates a weight map using a preceding feature map. The preceding feature map is generated using pixels of the digital image. The weight map is generated based on lie data associated with the digital image. A homographic transformation is interpolated between two planar projections of the digital image using at least the weight map and a homography matrix. The homography matrix provides a mapping between the two planar projections of the digital image. Homographic transformed kernels are generated by applying the homographic transformation to convolution kernels. The homographic transformed kernels are applied to the preceding feature map to conduct convolution on different plane regions appearing in the digital image and generate a new feature map, which is used for a computer vision task involving three-dimensional (3D) perception.
    Type: Application
    Filed: November 18, 2024
    Publication date: May 21, 2026
    Inventors: Yuliang Guo, Ruoyu Wang, Cheng Zhao, Xinyu Huang, Liu Ren, Abhinav Kumar
  • Publication number: 20260134605
    Abstract: Methods and systems for executing an online Gaussian Splatting model for simultaneous localization and mapping of a surrounding 3D space are disclosed. The model is configured to receive an image-based data sample that depicts a first field-of-view of the 3D space, and, using a 3D Gaussian map of the model, render both a new image-based data sample that depicts a new field-of-view that is different from the first field-of-view and render corresponding language features. By incorporating a hierarchical encoder and a Contrastive Language-Image Pre-training (CLIP) model into the architecture of the online Gaussian Splatting model, the overall architecture is configured to operate at near real-time.
    Type: Application
    Filed: November 14, 2024
    Publication date: May 14, 2026
    Inventors: Yuliang GUO, Saimouli KATRAGADDA, Xinyu HUANG, Liu REN
  • Patent number: 12626478
    Abstract: A method and device for performing a perception task are disclosed. The method and device incorporate a dense regression model. The dense regression model advantageously incorporates a distortion-free convolution technique that is designed to accommodate and appropriately handle the varying levels of distortion in omnidirectional images across different regions. In addition to distortion-free convolution, the dense regression model further utilizes a transformer that incorporates an spherical self-attention that use distortion-free image embedding to compute an appearance attention and uses spherical distance to compute a positional attention.
    Type: Grant
    Filed: December 13, 2021
    Date of Patent: May 12, 2026
    Assignee: Robert Bosch GmbH
    Inventors: Yuliang Guo, Zhixin Yan, Yuyan Li, Xinyu Huang, Liu Ren
  • Publication number: 20260116419
    Abstract: A system for training a drive system for an autonomous vehicle (AV) includes one or more computing devices configured to output a first drive plan and a second drive plan using an initial autonomous drive system (ADS) and an initial behavior foundation system (BFS). The computing devices is configured to generate a system-to-system (SS) loss between the initial ADS and the initial BFS using data from the respective system, and generate a module task loss for each system using respective drive plan and respective ground truth data. The computing devices is configured to adjust tunable parameters of the initial ADS and/or tunable parameters of the initial BFS to reduce a total loss provided by the SS loss and the module task loss. The initial ADS and/or the initial BFS is outputted as a trained drive system to be employed for the AV in response to the total loss being reduced.
    Type: Application
    Filed: October 31, 2024
    Publication date: April 30, 2026
    Inventors: Abhirup Mallik, Yunsheng Ma, Feng Tao, Xin Ye, Chenbin Pan, Burhaneddin Yaman, Liu Ren
  • Publication number: 20260120475
    Abstract: A Bird's Eye View (BEV)-based object detection framework for autonomous vehicles. The disclosed embodiments pretrain a diffusion model system on BEV representations generated from sensor data, such as cameras, LiDARs, and radars. The pretrained diffusion model system may be integrated into the BEV generation network to denoise BEV features through a supervision loss mechanism during training. This approach enhances the quality of BEV representations used in downstream tasks, such as object detection and trajectory prediction, without introducing any incremental computational cost during run-time.
    Type: Application
    Filed: October 31, 2024
    Publication date: April 30, 2026
    Inventors: Xin YE, Burhaneddin YAMAN, Feng TAO, Chenbin PAN, Abhirup MALLIK, Liu REN
  • Publication number: 20260111708
    Abstract: Methods for developing and managing long-term memory solutions for large language models (LLMs) within a context of providing agents of the LLM as a service are disclosed. Following task-related communications between LLM agents and users of the service, information pertaining to domain knowledge, user preferences, and success or not in completing the requested task is distilled into data samples by a reflections agent of the service. The data samples are then stored into a long-term memory database that is accessible by LLM agents in the future, such that the agents can recall information of previous interactions in order to more efficiently perform new tasks for users.
    Type: Application
    Filed: October 23, 2024
    Publication date: April 23, 2026
    Inventors: Jiajing GUO, Vikram MOHANTY, Jorge Henrique PIAZENTIN ONO, Wenbin HE, Liu REN
  • Publication number: 20260091792
    Abstract: Methods and systems for training an end-to-end autonomous driving system using a vision-language planning (VLP) machine learning model in a closed-loop environment. Images associated with an environment about a vehicle are generated, and a BEV model is executed to generate a BEV view based on the images. A planning model predicts navigation trajectories based on the BEV. The VLP model enhances the system by extracting vision-based planning features, generating text prompts, and employing a language encoder to create text-based expectation features. A contrastive learning model identifies similarities between vision and text features, boosting the performance of the BEV and planning models. The system undergoes closed-loop evaluation in a simulated environment, capturing metrics to refine the autonomous driving system.
    Type: Application
    Filed: September 27, 2024
    Publication date: April 2, 2026
    Inventors: Feng TAO, Xin YE, Burhaneddin YAMAN, Chenbin PAN, Abhirup MALLIK, Liu Ren
  • Patent number: 12592038
    Abstract: A computer-implemented method and system relate to computer vision. A first semantic map of an environment is three-dimensional (3D). A foreground scene and a background scene are generated individually using the semantic data of the first semantic map. The foreground scene contains foreground components of the first semantic map. The background scene contains background components of the first semantic map. A machine learning model generates an enhanced background view by completing incomplete regions of the background components. Input data is received to modify the background components, the foreground components, or both. A second semantic map is generated in 3D using the enhanced background view, the foreground components, and the input data. The second semantic map is 3D. Virtual camera data is generated using the second semantic map. The virtual camera data includes at least new image data and corresponding new depth data.
    Type: Grant
    Filed: May 24, 2024
    Date of Patent: March 31, 2026
    Assignee: Robert Bosch GmbH
    Inventors: Cheng Zhao, Yuliang Guo, Ruoyu Wang, Xinyu Huang, Liu Ren
  • Patent number: 12585250
    Abstract: In some implementations, the device may receive, for a plurality of stations, processing times indicating a time required for a part to be processed by each station, and waiting times indicating how long the part waited before moving to a subsequent one of the plurality of stations. In addition, the device may determine, cycle times for a predetermined window of time, where the cycle times indicates an average number of parts processed by the plurality of stations during the predetermined window of time. The device may determine one of the stations as a potential bottleneck station. Moreover, the device may display, to a user, the potential bottleneck station as a visualization which includes the processing time, the waiting time, and the cycle time of the potential bottleneck station. Also, the device may receive, from the user, a user feedback related to the potential bottleneck station.
    Type: Grant
    Filed: September 14, 2023
    Date of Patent: March 24, 2026
    Assignee: Robert Bosch GmbH
    Inventors: Jiajing Guo, Liang Gou, Samuel Kimport, Liu Ren