Audio
Marco_Longspeech
Marco-LongSpeech Dataset Marco-LongSpeech is a multi-task long speech understanding dataset containing 8 different speech understanding tasks…
Marco-LongSpeech Dataset Marco-LongSpeech is a multi-task long speech understanding dataset containing 8 different speech understanding tasks designed to benchmark Large Language Models on lengthy audio inputs. 📊 Dataset Statistics Task Statistics Task Train Val Test Total Unique Audios ASR 71,275 15,273 15,274 101,822 101,822 Temporal Relative QA 5,886 1,261 1,262 8,409 8,409 summary 4,366 935 937 6,238 6,238… See the full description on the dataset page: Longspeech.
Source: Hugging Face Hub (ATH-MaaS/Marco_Longspeech). Metadata imported from the dataset’s Hub tags.