Abstract: The invention provides a method for identifying similarity between two audio files or tracks. The method comprises receiving a processed audio file and an original audio file, uncompressing the processed audio file, applying global loudness normalization and short-term loudness normalization on the processed audio file and the original audio file, converting the processed audio file and the original audio file into processed spectral image by time-frequency mapping, scaling, using linear interpolation, the processed spectral image, dividing the scaled-up processed spectral image into slices, searching for minimum Sum of Absolute Difference (SAD), using original spectral image as reference, for each slice.
Abstract: The invention provides a method and a system for hierarchical audio classification. For getting high accuracy prediction with high resolution and low predictor complexity, the disclosed method uses a hierarchical classification approach with stateful prediction per frame aided by parallel AI transient detector for resetting the states of all stages at class transitions. To improve accuracy perfectly tagged database by innovative techniques of labeling are utilized. Further data augmentation is also done using signal processing techniques like audio mixing, blending of different type of data. The disclosed method applies short term audio normalization on database for normalized training and prediction of AI based Long Short-Term Memory (LSTM) networks. The disclosed method then uses a novel hierarchical classification approach with stateful LSTM prediction per frame aided by a parallel transient detector for resetting the states of all stages of hierarchical LSTM classifiers at class transitions.
Abstract: The invention provides a method for identifying similarity between two audio files or tracks. The method comprises receiving a processed audio file and an original audio file, uncompressing the processed audio file, applying global loudness normalization and short-term loudness normalization on the processed audio file and the original audio file, converting the processed audio file and the original audio file into processed spectral image by time-frequency mapping, scaling, using linear interpolation, the processed spectral image, dividing the scaled-up processed spectral image into slices, searching for minimum Sum of Absolute Difference (SAD), using original spectral image as reference, for each slice.