Abstract: A technique for detecting similarities in large sets of binary code files, e.g., bytecode files, without requiring access or knowledge of the actual source code itself. In accordance with the technique, bytecode files are disassembled and preprocessed using positional encoding to prepare the disassembled bytecode files for use in conjunction with similarity detection tools.