Skip to content
Advertisement
AudioMultimodalTabularText

StreamAudio-2M

StreamAudio-2M

StreamAudio-2M Large-scale streaming-audio dataset for audio-LLM / audio-agent training. Each row is a stream: a sequence of audio turns sharing one unified schema. ~2.28M unique audio clips are organised into six task subsets. Subsets Subset Rows Description Stream Audio Understanding 90,738 Montages of audio-understanding clips (AudioSet / FMA): captions, choice & open QA Real time ASR 28,109 Streams of ASR clips (CommonVoice / GigaSpeech /… See the full description on the dataset page:

Source: Hugging Face Hub (zhifeixie/StreamAudio-2M). Metadata imported from the dataset’s Hub tags.

Advertisement