Skip to content
Advertisement
AudioMultimodalText

Speech Recognition Alignment Dataset

Speech Recognition Alignment Dataset

Speech Recognition Alignment Dataset This dataset is a variation of several widely-used ASR datasets, encompassing Librispeech, MuST-C, TED-LIUM, VoxPopuli, Common Voice, and GigaSpeech. The difference is this dataset includes: Precise alignment between audio and text. Text that has been punctuated and made case-sensitive. Identification of named entities in the text. Usage First, install the latest version of the 🤗 Datasets package: pip install –upgrade pip pip… See the full description on the dataset page:

Source: Hugging Face Hub (nguyenvulebinh/asr-alignment). Metadata imported from the dataset’s Hub tags.

Advertisement