Skip to content
Advertisement
AudioMultimodalText

Audio-NTREX-4L

Audio-NTREX-4L

Audio-NTREX-4L Dataset Description Audio-NTREX-4L is a long-form multilingual speech translation dataset from 🇫🇷 French, 🇪🇸 Spanish, 🇵🇹 Portuguese and 🇩🇪 German to 🇬🇧 English designed to evaluate speech translation models on multi-sentence utterances. It is built from the text translation dataset NTREX by aggregating multiple sentences from a same context to create new source texts and their reference translation. We then use 3 different state-of-the-art… See the full description on the dataset page:

Source: Hugging Face Hub (kyutai/Audio-NTREX-4L). Metadata imported from the dataset’s Hub tags.

Advertisement