Abstract: The method automates the creation of explainer videos from text-based data by using a processor to analyse and identify components of an input document, such as text and images. The processor determines the reading order based on visual analysis and extracts structured information, including text-based and image-based elements. The text-based elements include paragraphs, lines, tables, complex formulas and the like. The image-based elements include schematics, graphs, illustrations and the like. The processor identifies the optimal amount of content displayable per screen, and the associated timing selects a suitable layout from predefined layouts. It generates an explainer video for at least one topic in the input document, ensuring a coherent and visually organised presentation.
Abstract: A method for improved collection and annotation of training datasets for training machine learning models is described. An image is captured using an image capture device and established as a reference image frame. The object of interest in the reference frame is annotated using a geometrical shape and established as seed data along with a pre-defined quantity target. The dimension of the annotated object is scaled to establish an annotation guideline. The annotation guideline is shifted in position in subsequent live images and displayed along with a live image of the object of interest. The user is prompted to adjust the live image so that the object of interest is fully encompassed by annotation guideline and capture a second live image. The captured image is classified in real time as accepted based on a benchmarking threshold and stored. The process is repeated until the pre-defined quantity target is satisfied.