Patents by Inventor Yifan DING

Yifan DING has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Publication number: 20260214286
    Abstract: The embodiments of the disclosure provide a method, apparatus, device, storage medium and program product for real-time interaction. The method includes: determining, during a chat of a user with a digital assistant, whether to reply to a first user input with a content stream based on a user requirement indicated by the first user input; generating, in response to determining to reply to the first user input with the content stream, a first portion of streaming media content based on reply key information for the first user input; and presenting the streaming media content in the chat.
    Type: Application
    Filed: January 13, 2026
    Publication date: July 23, 2026
    Inventors: Ziyang ZHENG, Yifan DING, Siyu LIU
  • Publication number: 20260212865
    Abstract: The present disclosure relates to an interaction method and apparatus, an electronic device, and a computer-readable storage medium. The interaction method includes: entering a first voice mode in response to a first operation of a user, where maintaining the first voice mode does not require manual operation of the user; making an instant call with the user in the first voice mode; receiving visual information and voice information input by the user in the first voice mode; and generating reply information based on the visual information and the voice information input by the user.
    Type: Application
    Filed: January 12, 2026
    Publication date: July 23, 2026
    Inventors: Yifan DING, Jiayi ZHAO, Ziyang ZHENG
  • Publication number: 20260204258
    Abstract: Embodiments of the disclosure provide a method, an apparatus, a device, and a storage medium for interaction. The method includes: receiving, during a voice call between a user and a digital assistant, a first user input including at least a first voice input by the user to the digital assistant. First auxiliary content of a first modality associated with the first user input is obtained based on the first user input, the first modality being determined based on a user requirement indicated by the first user input. A voice reply for the first user input and a first preview view for the first auxiliary content are presented. In this way, the auxiliary content is matched to respond to the user while presenting the voice reply. Therefore, more intuitive and comprehensive information is provided for the user, thus improving the response efficiency of the digital assistant.
    Type: Application
    Filed: January 9, 2026
    Publication date: July 16, 2026
    Inventors: Ziyang ZHENG, Yifan DING, Siyu LIU
  • Publication number: 20260195585
    Abstract: Neural network architectures and machine learning techniques for world foundation models (WFMs), e.g., a WFM suitable for training Physical AI. In at least one embodiment, a system comprises processing circuitry to perform training and/or inferencing using a neural network configured to receive, as input, text and/or an image/video and to generate, as output, a temporally coherent and 3D-consistent video simulation. In at least one embodiment, the neural networks include both a self-attention mechanism configured to incorporate positional embeddings and a cross-attention mechanism configured to incorporate text embeddings. In at least one embodiment, the neural network is trained in a multi-stage training process during which cross-attention mechanisms are incorporated following a prior stage and before a subsequent stage.
    Type: Application
    Filed: April 16, 2025
    Publication date: July 9, 2026
    Inventors: Haoxiang Wang, Yifan Ding, Xian Liu, Jiaojiao Fan, Xiaohui Zeng, Yogesh Balaji, Ming-Yu Liu
  • Publication number: 20260195355
    Abstract: The disclosure provides a method, an apparatus, a device, a storage medium, and a program product for query. In one method, first media data is obtained by a first acquisition device of a terminal device. Second media data is obtained by a second acquisition device of the terminal device, in response to determining that the second media data needs to be obtained, the determination is based on the first media data and a query. A reply for the query is provided based on the first media data and the second media data.
    Type: Application
    Filed: January 8, 2026
    Publication date: July 9, 2026
    Inventors: Yifan DING, Ziyang ZHENG, Siyu LIU
  • Publication number: 20260179601
    Abstract: A method for residual adapters for few-shot text-to-speech speaker adaptation includes obtaining a text-to-speech (TTS) model configured to convert text into representations of synthetic speech, the TTS model pre-trained on an initial training data set. The method further includes augmenting the TTS model with a stack of residual adapters. The method includes receiving an adaption training data set including one or more spoken utterances spoken by a target speaker, each spoken utterance in the adaptation training data set paired with corresponding input text associated with a transcription of the spoken utterance. The method also includes adapting, using the adaption training data set, the TTS model augmented with the stack of residual adapters to learn how to synthesize speech in a voice of the target speaker by optimizing the stack of residual adapters while parameters of the TTS model are frozen.
    Type: Application
    Filed: February 13, 2026
    Publication date: June 25, 2026
    Applicant: Google LLC
    Inventors: Nobuyuki Morioka, Byungha Chun, Nanxin Chen, Yu Zhang, Yifan Ding
  • Patent number: 12656938
    Abstract: The present disclosure relates to a content posting method and device and a computer-readable storage medium, and relates to the field of computer technology. The content posting method includes: generating multimedia content to be posted by an agent according to obtained network information, wherein the multimedia content comprises related content generated according to the network information and key information in the network information; posting the multimedia content; displaying a human-machine interaction interface of the agent in response to an interactive request initiated by a user for the multimedia content; and realizing an interaction between the user and the agent based on the multimedia content in the human-computer interaction interface.
    Type: Grant
    Filed: February 25, 2025
    Date of Patent: June 16, 2026
    Assignee: Beijing Zitiao Network Technology Co., Ltd.
    Inventors: Jiayi Zhao, Daoyu Wang, Yifan Ding, Hui Sun, Ziyang Zheng
  • Publication number: 20260141611
    Abstract: Approaches presented herein may be used to generate three-dimensional (3D) scenes using one or more 3D assets obtained from a layout image. The layout image may be produced from an input prompt, such as a text prompt, and objects at least partially depicted by the layout image may be identified and then used to generate the one or more 3D assets. One or more trained models may be used as part of a pipeline to receive an input, generate the one or more layout images, identify objects at least partially depicted by the one or more layout images, produce one or more 3D assets based on a description of the objects and/or a cropped image of the objects, and then to produce a 3D scene representing the layout image.
    Type: Application
    Filed: November 15, 2024
    Publication date: May 21, 2026
    Inventors: Yifan Ding, Yin Cui, Yunhao Ge, Tsung-Yi Lin, Ming-Yu Liu, Chen-Hsuan Lin, Zekun Hao, Zhaoshuo Li, Xiaohui Zeng, Zeqi Gu
  • Publication number: 20260101090
    Abstract: The present disclosure relates to a content posting method and apparatus, an electronic device, a storage medium, and a program product, and relates to the field of computer technologies. The content posting method includes: displaying content posted by a first agent on a first interface, wherein the content posted by the first agent is generated according to first attribute information of the first agent; and posting comment content on a second interface, wherein the comment content includes at least one of a first comment content made by a user on the content posted by the first agent or a second comment content made by a second agent on the content posted by the first target agent, and the second comment content is generated according to second attribute information of the second agent.
    Type: Application
    Filed: December 10, 2025
    Publication date: April 9, 2026
    Inventors: Jiayi ZHAO, Daoyu WANG, Yifan DING, Hui SUN, Ziyang ZHENG
  • Patent number: 12573368
    Abstract: A method for residual adapters for few-shot text-to-speech speaker adaptation includes obtaining a text-to-speech (TTS) model configured to convert text into representations of synthetic speech, the TTS model pre-trained on an initial training data set. The method further includes augmenting the TTS model with a stack of residual adapters. The method includes receiving an adaption training data set including one or more spoken utterances spoken by a target speaker, each spoken utterance in the adaptation training data set paired with corresponding input text associated with a transcription of the spoken utterance. The method also includes adapting, using the adaption training data set, the TTS model augmented with the stack of residual adapters to learn how to synthesize speech in a voice of the target speaker by optimizing the stack of residual adapters while parameters of the TTS model are frozen.
    Type: Grant
    Filed: October 24, 2023
    Date of Patent: March 10, 2026
    Assignee: Google LLC
    Inventors: Nobuyuki Morioka, Byungha Chun, Nanxin Chen, Yu Zhang, Yifan Ding
  • Publication number: 20260038190
    Abstract: In various examples, techniques for performing data filtering for AI systems and applications is described herein. Systems and methods described herein may use a pipeline that is configured to filter shapes, such as three-dimensional shapes, in order to identify high-quality shapes for a final dataset. In some examples, to identify the high-quality shapes, training shapes along with ground truth scores associated with the training shapes may be used to train a machine learning model. During and/or after the training, the machine learning model may then be used to process data associated with additional shapes in order to determine quality scores associated with the additional shapes. These quality scores may then be used to select the high-quality shapes, such as shapes that satisfy a threshold quality score. In some examples, additional filtering may be performed by the pipeline, such as by using one or more rules for removing low-quality shapes.
    Type: Application
    Filed: July 31, 2024
    Publication date: February 5, 2026
    Inventors: Zekun Hao, Yin Cui, Fangyin Wei, Yunhao Ge, Yifan Ding, Xiaohui Zeng, Tsung-Yi Lin, Chen-Hsuan Lin, Ming-Yu Liu
  • Publication number: 20260038191
    Abstract: In various examples, techniques for automatic annotation of shapes for AI systems and applications is described herein. Systems and methods described herein may use a pipeline that is configured to generate annotations for shapes, such as three-dimensional shapes, using various types of captions. For instance, image data representing images of the shapes, data representing description of the shapes, and/or data representing a format for the annotations may be input into one or more multimodal language models. The multimodal language model(s) may then be configured to process the data and, based at least on the processing, generate short captions and long captions associated with the shapes. These captions may then be stored in association with the shapes and/or the images. In some examples, embeddings may initially be generated for the captions, where the embeddings are then stored in association with the shapes and/or the images.
    Type: Application
    Filed: July 31, 2024
    Publication date: February 5, 2026
    Inventors: Zekun Hao, Yin Cui, Fangyin Wei, Yunhao Ge, Yifan Ding, Xiaohui Zeng, Tsung-Yi Lin, Chen-Hsuan Lin, Ming-Yu Liu
  • Publication number: 20260038213
    Abstract: In various examples, techniques for aligning shapes using pose information for AI systems and applications is described herein. Systems and methods described herein may use one or more pipelines that are configured to align shapes, such as three-dimensional shapes, using poses of the shapes. In some examples, to identify the poses, training images along with ground truth poses associated with the training images may be used to train a machine learning model. During and/or after the training, the machine learning model may then be used to process additional images in order to identify poses of additional shapes. In some examples, a pose associated with a shape may include a gravity orientation and/or an azimuth orientation of the shape as represented by an image. These poses may then be used to align the additional shapes, such as with respect to a canonical pose.
    Type: Application
    Filed: July 31, 2024
    Publication date: February 5, 2026
    Inventors: Zekun Hao, Yin Cui, Fangyin Wei, Yunhao Ge, Yifan Ding, Xiaohui Zeng, Tsung-Yi Lin, Chen-Hsuan Lin, Ming-Yu Liu
  • Patent number: 12518453
    Abstract: The present disclosure relates to a content distribution method and apparatus, an electronic device, a storage medium, and a program product, and relates to the field of computer technologies. The content distribution method includes: generating dynamic content of a target agent according to timeliness information, and attribute information of the target agent; and distributing the dynamic content of the target agent in a network community, wherein the network community includes the dynamic content of the target agent and dynamic content of a user.
    Type: Grant
    Filed: February 13, 2025
    Date of Patent: January 6, 2026
    Assignee: Beijing Zitiao Network Technology Co., Ltd.
    Inventors: Jiayi Zhao, Daoyu Wang, Yifan Ding, Hui Sun, Ziyang Zheng
  • Publication number: 20250378603
    Abstract: Apparatuses, systems, and techniques to identify a location in which to place objects within a graphically rendered scene. In at least one embodiment, a location in which to place objects is identified using one or more neural networks, based, at least in part, on text or speech input to the one or more neural networks.
    Type: Application
    Filed: June 7, 2024
    Publication date: December 11, 2025
    Inventors: Yunhao Ge, Yifan Ding, Yin Cui, Tsung-Yi Lin, Chen-Hsuan Lin, Zekun Hao, Zhaoshuo Li, Xiaohui Zeng, Hanzi Mao, Ming-Yu Liu, John Peter Lewis, Donglai Xiang, Qianli Ma, Jiashu Xu Xu
  • Publication number: 20250348692
    Abstract: Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for performing speech-to-speech translation, including real-time speech-to-speech translation.
    Type: Application
    Filed: May 8, 2025
    Publication date: November 13, 2025
    Inventors: Alex Agranovich, Eliya Nachmani, Oleg Rybakov, Yifan Ding, Ye Jia, Nadav Bar, Byungha Chun, Michelle Dana Tadmor
  • Publication number: 20250335073
    Abstract: The present disclosure relates to an interaction method, device, electronic apparatus, storage medium and program product, and involves the technical field of artificial intelligence. The interaction method of the present disclosure comprises: displaying, in response to a user triggering an instruction creation function, an instruction creation interface; determining prompt information corresponding to a target instruction according to information for creating the target instruction input by the user in the instruction creation interface, wherein the target instruction is configured to indicate demand information of the user for an interaction function of an Agent; and generating and displaying an operation control of the target instruction according to the prompt information.
    Type: Application
    Filed: March 21, 2025
    Publication date: October 30, 2025
    Inventors: Hui SUN, Yifan DING, Geng HUANG
  • Publication number: 20250328223
    Abstract: The present disclosure relates to a content posting method and device and a computer- readable storage medium, and relates to the field of computer technology. The content posting method includes: generating multimedia content to be posted by an agent according to obtained network information, wherein the multimedia content comprises related content generated according to the network information and key information in the network information; posting the multimedia content; displaying a human-machine interaction interface of the agent in response to an interactive request initiated by a user for the multimedia content; and realizing an interaction between the user and the agent based on the multimedia content in the human-computer interaction interface.
    Type: Application
    Filed: February 25, 2025
    Publication date: October 23, 2025
    Inventors: Jiayi ZHAO, Daoyu WANG, Yifan DING, Hui SUN, Ziyang ZHENG
  • Publication number: 20250329087
    Abstract: The present disclosure relates to a content distribution method and apparatus, an electronic device, a storage medium, and a program product, and relates to the field of computer technologies. The content distribution method includes: generating dynamic content of a target agent according to timeliness information, and attribute information of the target agent; and distributing the dynamic content of the target agent in a network community, wherein the network community includes the dynamic content of the target agent and dynamic content of a user.
    Type: Application
    Filed: February 13, 2025
    Publication date: October 23, 2025
    Inventors: Jiayi ZHAO, Daoyu WANG, Yifan DING, Hui SUN, Ziyang ZHENG
  • Publication number: 20250289629
    Abstract: A removable valve and lid assembly can include a removable valve and a lid. The removable valve can have a pressure valve and a spout. The lid can have a closing structure operable to control access to an opening of the lid and a securing mechanism disposed adjacent to the opening, wherein the securing mechanism is configured to receive and secure the removable valve, and wherein the closing structure includes a protrusion configured to press against the spout of the removable valve when the removable valve is disposed within the opening of the lid and when the closing structure is pressed against the spout. The removable valve and lid assembly can be used to seal a container or bottle.
    Type: Application
    Filed: March 12, 2025
    Publication date: September 18, 2025
    Applicant: Munchkin Incorporated
    Inventors: FangYi Gao, Agnes Yena Lee, Yifan Ding