Patents by Inventor Shoufa CHEN

Shoufa CHEN has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Publication number: 20260179181
    Abstract: Embodiments of the disclosure provide a method, an apparatus, a device, a storage medium, and a program product for visual generation. The method includes: generating first visual content matching the text information at a first resolution by using a trained first content generation model and based on text information; performing up-sampling for the first visual content having the first resolution to obtain up-sampled first visual content having a second resolution, where the first resolution is lower than the second resolution; and generating second visual content matching the text information at the second resolution by using a trained second content generation model and based on the up-sampled first visual content and the text information.
    Type: Application
    Filed: August 22, 2025
    Publication date: June 25, 2026
    Inventors: Shilong ZHANG, Shoufa Chen, Chongjian Ge, Yi Jiang, Zehuan Yuan, Bingyue Peng
  • Publication number: 20260148433
    Abstract: A method includes: determining, in response to obtaining description information related to visual content generation, position information of a respective text unit in the description information based on a specified visual category, the specified visual category indicating a video category or an image category, the position information including at least one of spatial position information or temporal position information; generating a visual feature map matching the description information by using a trained content generation model and based on text encoding representation and the position information of the respective text unit in the description information; and generating visual content matching the specified visual category by using a trained decoder model and based on the visual feature map, the decoder model being trained to decode an image from a visual feature map corresponding to the image and to decode a video from a visual feature map corresponding to the video.
    Type: Application
    Filed: September 15, 2025
    Publication date: May 28, 2026
    Inventors: Shoufa CHEN, Chongjian GE, Fengda ZHU, Yuqi ZHANG, Yi JIANG, Zehuan YUAN, Bingyue PENG
  • Publication number: 20250336102
    Abstract: According to embodiments of the disclosure, a method, apparatus, a device, a medium, and a product for image generation are provided. The method includes: receiving a text sequence indicating condition information of image generation; inputting the text sequence into a trained image generation model; and generating, through the image generation model, a target image matching the condition information based on at least the text sequence. The target resolution of the target image is determined based on the text sequence. The image generation model is obtained through training based on a sample image set and a sample text sequence set. A sample image in the sample image set matches a sample text sequence in the sample text sequence set. Sample images in the sample image set have different resolutions, and sample text sequences have different text lengths.
    Type: Application
    Filed: April 29, 2025
    Publication date: October 30, 2025
    Inventors: Yi Jiang, Shoufa Chen, Peize Sun, Yuqi Zhang, Shilong Zhang, Zehuan Yuan
  • Publication number: 20250336185
    Abstract: A method for image generation includes: processing an input text sequence by using a trained language model to obtain an output sequence output by the language model, the output sequence including a plurality of indices in a language dictionary associated with the language model, the language model being trained on the language dictionary, the language dictionary including at least an index set corresponding to text encodings in a natural language and an index set corresponding to image encodings; constructing image encodings corresponding to the plurality of indices in the output sequence into a target feature map; and determining, by using a trained image decoder, a target image matching the text sequence from the target feature map, the image decoder being trained on a visual dictionary including the index set corresponding to the image encodings.
    Type: Application
    Filed: April 28, 2025
    Publication date: October 30, 2025
    Inventors: Yi Jiang, Peize Sun, Shoufa Chen, Shilong Zhang, Zehuan Yuan
  • Publication number: 20250140021
    Abstract: The present invention provides a video action detection method and electronic equipment based on an end-to-end framework, which includes a backbone network, a positioning module and a classification module. The method includes: feature extraction of the video clip to be tested by the backbone network, obtaining the video feature map of the video clip, which includes the feature maps of all frames in the video clip; the backbone network extracts the feature map of the key frame from the video feature map and obtains the actor's location features from the feature map of the key frame, and the action category features are obtained from the video feature map; the positioning and classification modules determine the actor's location and action category respectively from the features extracted by the backbone network. This method provided by the present invention has low complexity while achieving better detection performance at the same time.
    Type: Application
    Filed: August 19, 2022
    Publication date: May 1, 2025
    Inventors: Ping LUO, Shoufa CHEN, Jiajun SHEN