Skip to content
Advertisement
AudioMultimodalText

VoxLingua107

VoxLingua107

VoxLingua107 VoxLingua107 is a speech dataset for training spoken language identification models. The dataset consists of short speech segments automatically extracted from YouTube videos and labeled according the language of the video title and description, with some post-processing steps to filter out false positives. VoxLingua107 contains data for 107 languages. The total amount of speech in the training set is 6628 hours. The average amount of data per language is 62 hours.… See the full description on the dataset page: wds.

Source: Hugging Face Hub (TalTechNLP/voxlingua107_wds). Metadata imported from the dataset’s Hub tags.

Advertisement