Patents by Inventor Michael Busha

Michael Busha has filed for patents to protect the following inventions. This listing includes patent applications that are pending as well as patents that have already been granted by the United States Patent and Trademark Office (USPTO).

  • Publication number: 20250342177
    Abstract: Embodiments of the present disclosure relate to a system and method for classification and reclassification of structured and unstructured data using similarity-based signatures. Entities within a text document of structured and unstructured data are detected by a pre-trained artificial intelligence model. Multi-level embeddings are generated for each entity to capture contextual relationships, enabling calculation of similarity metrics and generation of similarity-based signatures. The entities are clustered based on the embeddings for purposes including visualization and batch classification. Clustering is performed in a first mode based on header information and data types, and in a second mode based on semantic meaning and format characteristics of column data. A user interface enables users to provide feedback on the clustering results, identifying cluster assignments as true positives or false positives.
    Type: Application
    Filed: May 2, 2025
    Publication date: November 6, 2025
    Inventors: Michael Busha, Michael Rinehart, Syed Muhammad Raza Abbas
  • Publication number: 20250077714
    Abstract: A system to analyse impact of data breaches on sensitive data is disclosed. The system includes a hardware processor and memory with program instructions for executing various modules. The data collection module retrieves and enriches impacted data from multiple repositories. The data identification module uses data loss prevention (DLP) and named entity recognition (NER) techniques, enhanced by large language models (LLMs), to accurately identify personal information. The identity deduplication module consolidates individual references using deterministic and probabilistic techniques, while the residency inference module applies machine learning and heuristic methods to determine residency based on various data sources. The analysis module assesses impacted data to identify relevant laws and estimate fines. The automation module streamlines response actions, including generating notifications and ensuring compliance.
    Type: Application
    Filed: August 30, 2024
    Publication date: March 6, 2025
    Inventors: Rehan Jalil, Michael Busha, Michael Rinehart
  • Publication number: 20250061129
    Abstract: A computer-implemented system for inferred lineage in data transformations is disclosed. A transformation module of the system includes a variation generation module to generate a plurality of variants by applying transformations to f source tables and corresponding target tables, a comparison module computes a table similarity and an inverse document frequency term, a sorting module to sort the columns of the source table, revise a plurality of estimates, and prune the column of the source table for a low revised estimates of coverage. A similar descriptors transformation module calculates a histogram-based similarity of two columns and variants, maps entries of the columns, calculate a score between the two histograms, and compute an artificial intelligence embedding. A mutual information module builds a supervised regression model, ranks importance of each feature, iteratively remove one feature, and stops the iterative removal of the feature at an occurrence of a jump in loss value.
    Type: Application
    Filed: August 2, 2024
    Publication date: February 20, 2025
    Inventors: Michael Rinehart, Michael Busha, Xiaolin Wang, Venkata Sunil Yarram
  • Patent number: 12105845
    Abstract: A system and a method for entity resolution is disclosed. An entity reference parsing subsystem to parse one or more entity references of a corresponding seed set of entity into corresponding one or more personal data properties and property values. A property value standardization subsystem performs one or more standardization operations for standardization of the corresponding one or more property values. A property value anonymization subsystem secures the one or more property values by performing one or more anonymization procedures. A property strength quantification subsystem identifies at least one additional property suspected to belong to the seed set of the entity, assigns a property strength score to the at least one additional property, adds the at least one additional property to the corresponding seed set of entity. A local entity resolution and a global entity resolution subsystem performs a first and a second entity resolution process respectively.
    Type: Grant
    Filed: November 6, 2020
    Date of Patent: October 1, 2024
    Assignee: SECURITI, Inc.
    Inventors: Michael Busha, Jiachen Mao, Michael Rinehart, Rehan Jalil
  • Publication number: 20230134223
    Abstract: A system and method for large scale categorization of website cookies is disclosed. The method includes gathering information about cookies from a first and second source. The cookies include complex and discrete features. The method includes populating the cookies into a first and second table. The method includes subjecting the first and second table to a machine learning technique to recognize and determine the features. The machine learning technique is operable to convert the complex features into discrete features, wherein the discrete features are set by using at least one of external datasets and embedding the complex features; embed the cookies, wherein a classifier is built as an output of embedding of the cookies; and create a model by using ensembling learning. The method includes categorizing the cookies into a third table and a fourth table. The method includes merging the third and fourth table.
    Type: Application
    Filed: November 1, 2022
    Publication date: May 4, 2023
    Inventor: Michael Busha
  • Publication number: 20220043934
    Abstract: A system and a method for entity resolution is disclosed. An entity reference parsing subsystem to parse one or more entity references of a corresponding seed set of entity into corresponding one or more personal data properties and property values. A property value standardization subsystem performs one or more standardization operations for standardization of the corresponding one or more property values. A property value anonymization subsystem secures the one or more property values by performing one or more anonymization procedures. A property strength quantification subsystem identifies at least one additional property suspected to belong to the seed set of the entity, assigns a property strength score to the at least one additional property, adds the at least one additional property to the corresponding seed set of entity. A local entity resolution and a global entity resolution subsystem performs a first and a second entity resolution process respectively.
    Type: Application
    Filed: November 6, 2020
    Publication date: February 10, 2022
    Inventors: Michael Busha, Jiachen Mao, Michael Rinehart, Rehan Jalil