Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
StreamAudio-2M Large-scale streaming-audio dataset for audio-LLM / audio-agent training. Each row is a stream: a sequence of audio turns sharing one unified schema. ~2.28M unique audio clips are organised into six task subsets. Subsets Subset Rows Description Stream Audio Understanding 90,738 Montages of audio-understanding clips (AudioSet / FMA): captions, choice & open QA Real time ASR 28,109 Streams of ASR clips (CommonVoice / GigaSpeech /… See the full description on the dataset page:
Source: Hugging Face Hub (zhifeixie/StreamAudio-2M). Metadata imported from the dataset’s Hub tags.