Abstract: In a computer-implemented method for extracting information from a plurality of electronic documents, a plurality of electronic documents is accessed. Each electronic document of the plurality of electronic documents is segmented into segments comprising at least one word. The segments are converted into content-sensitive vectorizations. The content-sensitive vectorizations are compared to identify the content-sensitive vectorizations that are within a similarity threshold. Segments having content-sensitive vectorizations that are within the similarity threshold are grouped into a plurality of segment groups. Information is extracted from the plurality of segment groups for performing analysis of the plurality of electronic documents.
Type:
Grant
Filed:
November 9, 2023
Date of Patent:
October 21, 2025
Assignee:
CATYLEX, INC.
Inventors:
Andrew Downes, David Rosen, Dhruv Sharma, Jamie Wodetzki