Patents by Inventor Shen YAN

Shen YAN has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Publication number: 20250259465
    Abstract: One example aspect of the present disclosure is directed to a streaming model for video processing tasks, such as, for example, dense video captioning. Thanks to a memory mechanism, the proposed streaming model does not require access to all input frames concurrently in order to process the video. Moreover, thanks to a new streaming decoding algorithm, the proposed model can produce outputs causally without processing the entire input sequence. The streaming model is inherently suited to processing long videos—as it ingests frames sequentially (e.g., one at a time or in small batches). Moreover, as the output is streamed, intermediate predictions can be produced before processing the full video. This property means that the streaming model can be applied to process live video streams, as required for applications such as video conferencing, security and continuous monitoring among others.
    Type: Application
    Filed: February 12, 2025
    Publication date: August 14, 2025
    Inventors: Xingyi Zhou, Anurag Arnab, Shyamal Deep Buch, Shen Yan, Austin Oliver Myers, Xuehan Xiong, Arsha Nagrani, Cordelia Luise Schmid
  • Publication number: 20250124708
    Abstract: Provided is an efficient approach to establish a foundational video-text model for tasks including open-vocabulary video classification, text-to-video retrieval, video captioning and video question-answering. Some example implementations include a model which can be referred to as VideoCoCa. Example implementations reuse a pretrained image-text contrastive captioner (CoCa) model and adapt it to video-text tasks with little or minimal extra training. While previous works adapt image-text models with various cross-frame fusion modules (for example, cross-frame attention layer or perceiver resampler) and finetune the modified architecture on video-text data, aspects of the present disclosure leverage findings that the generative attentional pooling and contrastive attentional pooling layers in the image-text CoCa design are instantly adaptable to “flattened frame embeddings”, yielding a strong zero-shot transfer baseline for many video-text tasks.
    Type: Application
    Filed: December 8, 2023
    Publication date: April 17, 2025
    Inventors: Shen Yan, Tao Zhu, Zirui Wang, Yuan Cao, Jiahui Yu
  • Publication number: 20250054306
    Abstract: Aspects of the disclosure are directed to methods and systems for short form previews of long form media items. A server can provide, to an artificial intelligence (AI) model, a long form media item to be shared with users. The server can receive, from the AI model, one or more frames that are predicted to contain content that is of interest to the users. The server can extract a segment of the long form media item that corresponds to the one or more frames, where the extracted segment corresponds to a short form media item preview. The short form media item preview can be provided for presentation to the users.
    Type: Application
    Filed: August 7, 2024
    Publication date: February 13, 2025
    Inventors: Daniel S. Cohen, Christopher R. Conover, Emily Rose Smith, Anoop Menon, Benjamin Lehn, Sudheendra Vijayanarasimhan, Bo Hu, Shen Yan, Xuehan Xiong, David Alexander Ross
  • Publication number: 20240371164
    Abstract: Methods and systems for video localization using artificial intelligence are provided herein. A set of video embeddings representing features of one or more video frames of a media it em and a set of textual embeddings corresponding to an event associated with the media item are obtained. Fused video-textual data is generated based on the set of video embeddings and the set of textual embeddings. The fused video-textual data indicates features of the video frames of the media item and textual data pertaining to the media item. The fused video-textual data is provided as an input to an artificial intelligence (AI) model trained to perform multiple video localization tasks with respect to media items of a platform. One or move outputs of the AI model are obtained. A segment of the media item that depicts the event is determined based on the one or move outputs of the AI model.
    Type: Application
    Filed: May 1, 2024
    Publication date: November 7, 2024
    Inventors: Shen Yan, Xuehan Xiong, Arsha Nagrani, Anurag Arnab, David Alexander Ross, Cordelia Schmid
  • Patent number: 11537901
    Abstract: A system and method for domain adaptation involves a first domain and a second domain. A machine learning system is trained with first sensor data and first label data of the first domain. Second sensor data of a second domain is obtained. Second label data is generated via the machine learning system based on the second sensor data. Inter-domain sensor data is generated by interpolating the first sensor data of the first domain with respect to the second sensor data of the second domain. Inter-domain label data is generated by interpolating first label data of the first domain with respect to second label data of the second domain. The machine learning system is operable to generate inter-domain output data in response to the inter-domain sensor data. Inter-domain loss data is generated based on the inter-domain output data with respect to the inter-domain label data. Parameters of the machine learning system are updated upon optimizing final loss data that includes at least the inter-domain loss data.
    Type: Grant
    Filed: December 31, 2019
    Date of Patent: December 27, 2022
    Assignee: Robert Bosch GmbH
    Inventors: Huan Song, Shen Yan, Nanxiang Li, Lincan Zou, Liu Ren
  • Patent number: 11134129
    Abstract: Embodiments of the present invention provide a packet processing method, including: receiving, by a first node, a first packet, where the first packet carries a first bit string, the first bit string includes M bit sets, each bit set corresponds to one node group, a value of the bit set is used to indicate whether one or more target nodes of the first packet include the corresponding node group; and determining, by the first node based on the first bit string and a second bit string, whether to send the first packet to a second node, where the second bit string includes N bit sets, each bit set corresponds to one node group, a value of the bit set is used to indicate whether a node belonging to the corresponding node group exists in one or more related nodes of the first node.
    Type: Grant
    Filed: June 24, 2020
    Date of Patent: September 28, 2021
    Assignee: Huawei Technologies Co., Ltd.
    Inventors: Shoushou Ren, Delei Yu, Shen Yan, Chuang Wang, Zongxin Dou, Wanhong Wang
  • Publication number: 20210201159
    Abstract: A system and method for domain adaptation involves a first domain and a second domain. A machine learning system is trained with first sensor data and first label data of the first domain. Second sensor data of a second domain is obtained. Second label data is generated via the machine learning system based on the second sensor data. Inter-domain sensor data is generated by interpolating the first sensor data of the first domain with respect to the second sensor data of the second domain. Inter-domain label data is generated by interpolating first label data of the first domain with respect to second label data of the second domain. The machine learning system is operable to generate inter-domain output data in response to the inter-domain sensor data. Inter-domain loss data is generated based on the inter-domain output data with respect to the inter-domain label data. Parameters of the machine learning system are updated upon optimizing final loss data that includes at least the inter-domain loss data.
    Type: Application
    Filed: December 31, 2019
    Publication date: July 1, 2021
    Inventors: Huan Song, Shen Yan, Nanxiang Li, Lincan Zou, Liu Ren
  • Publication number: 20200329111
    Abstract: Embodiments of the present invention provide a packet processing method, including: receiving, by a first node, a first packet, where the first packet carries a first bit string, the first bit string includes M bit sets, each bit set corresponds to one node group, a value of the bit set is used to indicate whether one or more target nodes of the first packet include the corresponding node group; and determining, by the first node based on the first bit string and a second bit string, whether to send the first packet to a second node, where the second bit string includes N bit sets, each bit set corresponds to one node group, a value of the bit set is used to indicate whether a node belonging to the corresponding node group exists in one or more related nodes of the first node.
    Type: Application
    Filed: June 24, 2020
    Publication date: October 15, 2020
    Inventors: Shoushou REN, Delei YU, Shen YAN, Chuang WANG, Zongxin DOU, Wanhong WANG