Abstract: A technique for determining annotation items associated with video information is described. During this annotation technique, a content item that includes audio information and the video information is received. For example, a file may be downloaded from a uniform resource locator. Then, the audio information is extracted from the content item, and the audio information is analyzed to determine features or descriptors that characterize the audio information. Note that the features may be determined solely by analyzing the audio information or may be determined by subsequent further analysis of at least some of the video information based on the analysis of the audio information (i.e., sequential or cascaded analysis). Next, annotation items or tags associated with the video information are determined based on the features.