Abstract: The present disclosure discloses a VAD method and system based on a joint DNN. The method includes: acquiring original audio data and a first frame-level label based on an open-source audio data set, adding noise on the original audio data to obtain first audio data, recording an environment sound to obtain second audio data, and performing clip-level label on the second audio data to obtain a clip-level label; inputting the first audio data, the second audio data and the corresponding labels into a first neural network to obtain a first-stage network model; obtaining a second frame-level label through the first-stage network model, and inputting the second audio data and the second frame-level label into a second neural network to obtain a second-stage network model; and performing VAD on an audio signal based on the first-stage network model and the second-stage network model.