Patents by Inventor Doo Soon Kim
Doo Soon Kim has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).
-
Patent number: 12645959Abstract: This disclosure describes methods, non-transitory computer readable storage media, and systems that provide a platform for on-demand selection of machine-learning models and on-demand learning of parameters for the selected machine-learning models via cloud-based systems. For instance, the disclosed system receives a request indicating a selection of a machine-learning model to perform a machine-learning task (e.g., a natural language task) utilizing a specific dataset (e.g., a user-defined dataset). The disclosed system utilizes a scheduler to monitor available computing devices on cloud-based storage systems for instantiating the selected machine-learning model. Using the indicated dataset at a determined cloud-based computing device, the disclosed system automatically trains the machine-learning model.Type: GrantFiled: May 26, 2021Date of Patent: June 2, 2026Assignee: Adobe Inc.Inventors: Nham Van Le, Tuan Manh Lai, Trung Bui, Doo Soon Kim
-
Publication number: 20260082098Abstract: On a television or media device, it is painfully slow for users to click through a settings tree to perform device tasks. To address this issue, a voice-based application can be implemented to help users perform and access device tasks by voice. A user can make an utterance to reach a specific page in the setting tree where the user can then complete the device task. It is not trivial to implement the application. It can be a challenge to determine the precise device task intent from the utterance when there are hundreds of device tasks. The type of device task intent and the context of the user device may impact the way the user interface is to be updated. Some device tasks may be unsupported by the user device. Voice hints to help users learn to use their voice can follow a unique logic for suppressing voice hints.Type: ApplicationFiled: July 7, 2025Publication date: March 19, 2026Applicant: Roku, Inc.Inventors: Amit Vishvanath Desai, Siddhant Dinesh Shah, Tess Harty, Elizabeth Owen Bratt, Valeria Faria de Sá, I-Tsun Cheng, Doo Soon Kim, Arnaldo Carreno
-
Patent number: 12547616Abstract: Systems and methods for natural language processing are described. One or more embodiments of the present disclosure receive a query related to information in a table, compute an operation selector by combining the query with an operation embedding representing a plurality of table operations, compute a column selector by combining the query with a weighted operation embedding, compute a row selector based on the operation selector and the column selector, compute a probability value for a cell in the table based on the row selector and the column selector, where the probability value represents a probability that the cell provides an answer to the query, and transmit contents of the cell based on the probability value.Type: GrantFiled: May 11, 2021Date of Patent: February 10, 2026Assignee: ADOBE INC.Inventors: Dung Thai, Doo Soon Kim, Franck Dernoncourt, Trung Bui
-
Publication number: 20250391400Abstract: Some embodiments include a transcription knowledge graph that can resolve automatic speech recognition (ASR) engine output errors. In some embodiments, a transcription knowledge graph can utilize data from past sessions of the ASR engine to form a voice graph that can be analyzed to determine a correlation between a mis-transcription (error text) and the correct transcription (correct text). Thus, ASR engine outputs, even if they include a mis-transcription, can be adjusted to the correct transcription. Further, the correct transcriptions and the voice graph can be used to train machine learning (ML) algorithms to generate numerical representations of an entity. The ML algorithms can be applied to a transcription to correctly identify a corresponding entity label, even if the transcription was not utilized in the voice graph to train the ML algorithm.Type: ApplicationFiled: August 28, 2025Publication date: December 25, 2025Applicant: Roku, Inc.Inventors: Doo Soon Kim, Minsuk Heo, Zei-Chan Yeh, Behnam Asefisaray, Praful Chandra Mangalath
-
Publication number: 20250372090Abstract: Dialogue state tracking for voice assistants involves correctly tracking intent and entities of a task that a user is performing. A dialogue state, having a tracked intent and one or more tracked entities, can then be used to perform the task. Building a dialogue state tracking system within a voice assistant is not trivial. In some embodiments, a dialogue state tracking system involving one or more large language models can be implemented downstream of a natural language understanding system to produce the tracked intent and the one or more tracked entities. In some embodiments, a dialogue state tracking system involving one or more large language models can be implemented upstream of a natural language understanding system to produce rephrased natural language text, which is in turn processed by the natural language understanding system to produce the tracked intent and the one or more tracked entities.Type: ApplicationFiled: May 30, 2024Publication date: December 4, 2025Applicant: Roku, Inc.Inventors: I-Tsun Cheng, Zei-Chan Yeh, Doo Soon Kim, Praful Chandra Mangalath
-
Patent number: 12455911Abstract: Systems and methods for natural language processing are described. Embodiments of the present disclosure identify a task set including a plurality of pseudo tasks, wherein each of the plurality of pseudo tasks includes a support set corresponding to a first natural language processing (NLP) task and a query set corresponding to a second NLP task; update a machine learning model in an inner loop based on the support set; update the machine learning model in an outer loop based on the query set; and perform the second NLP task using the machine learning model.Type: GrantFiled: March 18, 2022Date of Patent: October 28, 2025Assignee: ADOBE INC.Inventors: Meryem M'hamdi, Doo Soon Kim, Franck Dernoncourt, Trung Huu Bui
-
Publication number: 20250310584Abstract: Disclosed herein are system, apparatus, article of manufacture, method and/or computer program product embodiments, and/or combinations and sub-combinations thereof, for tailoring and censoring content based on audience detected. An example embodiment operates by detecting an audience within a vicinity of a media device based on identifying information received by the media device, determining a category of the audience with a user identification system based on the identifying information, identifying a content tailoring rule for the audience based on the category of the audience, retrieving a content to be played by the media device, and modifying the content based on the content tailoring rule and a category label of the content.Type: ApplicationFiled: June 17, 2025Publication date: October 2, 2025Applicant: Roku, Inc.Inventors: Rakesh RAVURU, Bao NGUYEN, Behnam ASEFISARAY, Doo Soon KIM, Praful MANGALATH
-
Patent number: 12431123Abstract: Some embodiments include a transcription knowledge graph that can resolve automatic speech recognition (ASR) engine output errors. In some embodiments, a transcription knowledge graph can utilize data from past sessions of the ASR engine to form a voice graph that can be analyzed to determine a correlation between a mis-transcription (error text) and the correct transcription (correct text). Thus, ASR engine outputs, even if they include a mis-transcription, can be adjusted to the correct transcription. Further, the correct transcriptions and the voice graph can be used to train machine learning (ML) algorithms to generate numerical representations of an entity. The ML algorithms can be applied to a transcription to correctly identify a corresponding entity label, even if the transcription was not utilized in the voice graph to train the ML algorithm.Type: GrantFiled: June 9, 2023Date of Patent: September 30, 2025Assignee: Roku, Inc.Inventors: Doo Soon Kim, Minsuk Heo, Zei-Chan Yeh, Behnam Asefisaray, Praful Chandra Mangalath
-
Publication number: 20250259437Abstract: Embodiments of the disclosure provide a machine learning model for generating a predicted executable command for an image. The learning model includes an interface configured to obtain an utterance indicating a request associated with the image, an utterance sub-model, a visual sub-model, an attention network, and a selection gate. The machine learning model generates a segment of the predicted executable command from weighted probabilities of each candidate token in a predetermined vocabulary determined based on the visual features, the concept features, current command features, and the utterance features extracted from the utterance or the image.Type: ApplicationFiled: April 2, 2025Publication date: August 14, 2025Inventors: Seunghyun Yoon, Trung Huu Bui, Franck Dernoncourt, Hyounghun Kim, Doo Soon Kim
-
Patent number: 12363367Abstract: Disclosed herein are system, apparatus, article of manufacture, method and/or computer program product embodiments, and/or combinations and sub-combinations thereof, for tailoring and censoring content based on audience detected. An example embodiment operates by detecting an audience within a vicinity of a media device based on identifying information received by the media device, determining a category of the audience with a user identification system based on the identifying information, identifying a content tailoring rule for the audience based on the category of the audience, retrieving a content to be played by the media device, and modifying the content based on the content tailoring rule and a category label of the content.Type: GrantFiled: October 3, 2022Date of Patent: July 15, 2025Assignee: Roku, Inc.Inventors: Rakesh Ravuru, Bao Nguyen, Behnam Asefisaray, Doo Soon Kim, Praful Mangalath
-
Patent number: 12323262Abstract: Systems and methods for coreference resolution are provided. One aspect of the systems and methods includes inserting a speaker tag into a transcript, wherein the speaker tag indicates that a name in the transcript corresponds to a speaker of a portion of the transcript; encoding a plurality of candidate spans from the transcript based at least in part on the speaker tag to obtain a plurality of span vectors; extracting a plurality of entity mentions from the transcript based on the plurality of span vectors, wherein each of the plurality of entity mentions corresponds to one of the plurality of candidate spans; and generating coreference information for the transcript based on the plurality of entity mentions, wherein the coreference information indicates that a pair of candidate spans of the plurality of candidate spans corresponds to a pair of entity mentions that refer to a same entity.Type: GrantFiled: June 14, 2022Date of Patent: June 3, 2025Assignee: ADOBE INC.Inventors: Tuan Manh Lai, Trung Huu Bui, Doo Soon Kim
-
Patent number: 12293577Abstract: Embodiments of the disclosure provide a machine learning model for generating a predicted executable command for an image. The learning model includes an interface configured to obtain an utterance indicating a request associated with the image, an utterance sub-model, a visual sub-model, an attention network, and a selection gate. The machine learning model generates a segment of the predicted executable command from weighted probabilities of each candidate token in a predetermined vocabulary determined based on the visual features, the concept features, current command features, and the utterance features extracted from the utterance or the image.Type: GrantFiled: February 18, 2022Date of Patent: May 6, 2025Assignee: Adobe Inc.Inventors: Seunghyun Yoon, Trung Huu Bui, Franck Dernoncourt, Hyounghun Kim, Doo Soon Kim
-
Publication number: 20240412723Abstract: Some embodiments include a transcription knowledge graph that can resolve automatic speech recognition (ASR) engine output errors. In some embodiments, a transcription knowledge graph can utilize data from past sessions of the ASR engine to form a voice graph that can be analyzed to determine a correlation between a mis-transcription (error text) and the correct transcription (correct text). Thus, ASR engine outputs, even if they include a mis-transcription, can be adjusted to the correct transcription. Further, the correct transcriptions and the voice graph can be used to train machine learning (ML) algorithms to generate numerical representations of an entity. The ML algorithms can be applied to a transcription to correctly identify a corresponding entity label, even if the transcription was not utilized in the voice graph to train the ML algorithm.Type: ApplicationFiled: June 9, 2023Publication date: December 12, 2024Applicant: ROKU, INC.Inventors: Doo Soon Kim, Minsuk Heo, Zei-Chan Yeh, Behnam Asefisaray, Praful Chandra Mangalath
-
Patent number: 11893345Abstract: Systems and methods for natural language processing are described. One or more embodiments of the present disclosure receive a document comprising a plurality of words organized into a plurality of sentences, the words comprising an event trigger word and an argument candidate word, generate word representation vectors for the words, generate a plurality of document structures including a semantic structure for the document based on the word representation vectors, a syntax structure representing dependency relationships between the words, and a discourse structure representing discourse information of the document based on the plurality of sentences, generate a relationship representation vector based on the document structures, and predict a relationship between the event trigger word and the argument candidate word based on the relationship representation vector.Type: GrantFiled: April 6, 2021Date of Patent: February 6, 2024Assignee: ADOBE, INC.Inventors: Amir Pouran Ben Veyseh, Franck Dernoncourt, Quan Tran, Varun Manjunatha, Lidan Wang, Rajiv Jain, Doo Soon Kim, Walter Chang
-
Publication number: 20230403175Abstract: Systems and methods for coreference resolution are provided. One aspect of the systems and methods includes inserting a speaker tag into a transcript, wherein the speaker tag indicates that a name in the transcript corresponds to a speaker of a portion of the transcript; encoding a plurality of candidate spans from the transcript based at least in part on the speaker tag to obtain a plurality of span vectors; extracting a plurality of entity mentions from the transcript based on the plurality of span vectors, wherein each of the plurality of entity mentions corresponds to one of the plurality of candidate spans; and generating coreference information for the transcript based on the plurality of entity mentions, wherein the coreference information indicates that a pair of candidate spans of the plurality of candidate spans corresponds to a pair of entity mentions that refer to a same entity.Type: ApplicationFiled: June 14, 2022Publication date: December 14, 2023Inventors: Tuan Manh Lai, Trung Huu Bui, Doo Soon Kim
-
Publication number: 20230297603Abstract: Systems and methods for natural language processing are described. Embodiments of the present disclosure identify a task set including a plurality of pseudo tasks, wherein each of the plurality of pseudo tasks includes a support set corresponding to a first natural language processing (NLP) task and a query set corresponding to a second NLP task; update a machine learning model in an inner loop based on the support set; update the machine learning model in an outer loop based on the query set; and perform the second NLP task using the machine learning model.Type: ApplicationFiled: March 18, 2022Publication date: September 21, 2023Inventors: Meryem M'hamdi, Doo Soon Kim, Franck Dernoncourt, Trung Huu Bui
-
Publication number: 20230267726Abstract: Embodiments of the disclosure provide a machine learning model for generating a predicted executable command for an image. The learning model includes an interface configured to obtain an utterance indicating a request associated with the image, an utterance sub-model, a visual sub-model, an attention network, and a selection gate. The machine learning model generates a segment of the predicted executable command from weighted probabilities of each candidate token in a predetermined vocabulary determined based on the visual features, the concept features, current command features, and the utterance features extracted from the utterance or the image.Type: ApplicationFiled: February 18, 2022Publication date: August 24, 2023Inventors: Seunghyun Yoon, Trung Huu Bui, Franck Dernoncourt, Hyounghun Kim, Doo Soon Kim
-
Patent number: 11709690Abstract: The present disclosure relates to systems, methods, and non-transitory computer readable media for generating coachmarks and concise instructions based on operation descriptions for performing application operations. For example, the disclosed systems can utilize a multi-task summarization neural network to analyze an operation description and generate a coachmark and a concise instruction corresponding to the operation description. In addition, the disclosed systems can provide a coachmark and a concise instruction for display within a user interface to, directly within a client application, guide a user to perform an operation by interacting with a particular user interface element.Type: GrantFiled: March 9, 2020Date of Patent: July 25, 2023Assignee: Adobe Inc.Inventors: Nedim Lipka, Doo Soon Kim
-
Patent number: 11630952Abstract: This disclosure relates to methods, non-transitory computer readable media, and systems that can classify term sequences within a source text based on textual features analyzed by both an implicit-class-recognition model and an explicit-class-recognition model. For example, by applying machine-learning models for both implicit and explicit class recognition, the disclosed systems can determine a class corresponding to a particular term sequence within a source text and identify the particular term sequence reflecting the class. The dual-model architecture can equip the disclosed systems to apply (i) the implicit-class-recognition model to recognize implicit references to a class in source texts and (ii) the explicit-class-recognition model to recognize explicit references to the same class in source texts.Type: GrantFiled: July 22, 2019Date of Patent: April 18, 2023Assignee: Adobe Inc.Inventors: Sean MacAvaney, Franck Dernoncourt, Walter Chang, Seokhwan Kim, Doo Soon Kim, Chen Fang
-
Patent number: 11620457Abstract: Systems and methods for sentence fusion are described. Embodiments receive coreference information for a first sentence and a second sentence, wherein the coreference information identifies entities associated with both a term of the first sentence and a term of the second sentence, apply an entity constraint to an attention head of a sentence fusion network, wherein the entity constraint limits attention weights of the attention head to terms that correspond to a same entity of the coreference information, and predict a fused sentence using the sentence fusion network based on the entity constraint, wherein the fused sentence combines information from the first sentence and the second sentence.Type: GrantFiled: February 17, 2021Date of Patent: April 4, 2023Assignee: ADOBE INC.Inventors: Logan Lebanoff, Franck Dernoncourt, Doo Soon Kim, Lidan Wang, Walter Chang