Patents by Inventor Kezhen Chen

Kezhen Chen has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Patent number: 12731196
    Abstract: Implementations are disclosed for fusing multiple modalities of data into a multimodal feature embedding and then processing the multimodal feature embedding using various downstream processes for training and/or inference purposes. In various implementations, multiple different modalities of agricultural data about an agricultural parcel may be obtained. Each modality of agricultural data may be processed based on a respective modality-specific encoder to generate a respective modality-specific embedding. The plurality of modality-specific embeddings may be processed based on a multimodal fusion machine learning model to generate a multimodal feature embedding that represents the agricultural parcel. In some implementations, the multimodal feature embedding may be processed using downstream computer process(es) to generate agricultural prediction(s) about the agricultural parcel.
    Type: Grant
    Filed: September 28, 2023
    Date of Patent: September 8, 2026
    Assignee: Deere & Company
    Inventors: Yawen Zhang, Kezhen Chen, Jinmeng Rao, Xiaoyuan Guo, Jie Yang, Luis Pazos Outon
  • Patent number: 12651447
    Abstract: Implementations improve object classification/detection by leveraging visual attributes. An image depicting instance(s) of object class(es) is obtained with a textual snippet that includes: noun(s) identifying target object class(es); and adjective(s) describing target visual attribute(s). The textual snippet may be encoded as text embedding(s) that represent target object class(es) and visual attribute(s) in a shared embedding space. The image may be processed using an image encoder to generate image encoder output tokens (IEOTs) that are used to generate object visual embedding(s) in the shared embedding space. The text embedding(s) and the object visual embeddings may be used to classify the IEOTs as depicting an instance of the target object class(es) having target visual attribute(s). The IEOTs may also be processed using a localization head to predict annotation(s) for the digital image.
    Type: Grant
    Filed: May 5, 2023
    Date of Patent: June 9, 2026
    Assignee: Deere & Company
    Inventors: Kezhen Chen, Xiaoyuan Guo, Jie Yang, Yueqi Li
  • Publication number: 20250148555
    Abstract: Implementations are disclosed for fusing multiple modalities of data into a multimodal feature embedding and then processing the multimodal feature embedding using various downstream processes for training and/or inference purposes. In various implementations, multiple different modalities of agricultural data about an agricultural parcel may be obtained. Each modality of agricultural data may be processed based on a respective modality-specific encoder to generate a respective modality-specific embedding. The plurality of modality-specific embeddings may be processed based on a multimodal fusion machine learning model to generate a multimodal feature embedding that represents the agricultural parcel. In some implementations, the multimodal feature embedding may be processed using downstream computer process(es) to generate agricultural prediction(s) about the agricultural parcel.
    Type: Application
    Filed: September 28, 2023
    Publication date: May 8, 2025
    Inventors: Yawen Zhang, Kezhen Chen, Jinmeng Rao, Xiaoyuan Guo, Jie Yang, Luis Pazos Outon
  • Publication number: 20250107479
    Abstract: Systems and methods herein present a multi-modal model, or insight miner, for agricultural use. The model is used for generating actionable insights from time-series agricultural data. To do so, a control system of a farming machine includes the insight miner and a training module. The miner ingests agricultural information (e.g., time series field measurements), identifies trends in the time-series data (e.g., increasing, decreasing, etc.), and determines potential farming actions to take based on identified trends. To do so, the model generates visual representations of the time-series data, inputs the visual representations into a model to obtain a natural language description of that data, and enact farming actions based on the insights obtained. The training module trains the model by generating visual representation and natural language pairs, which are used to train the model.
    Type: Application
    Filed: October 2, 2024
    Publication date: April 3, 2025
    Inventors: Yawen Zhang, Yunkai Zhang, Kezhen Chen, Ming Zheng, Jinmeng Rao, Xiaoyuan Guo, Hongxu Ma, Jie Yang, Christopher Grant Padwick