Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
AudioSet AudioSet 1 is a large-scale dataset comprising approximately 2 million 10-second YouTube audio clips, categorised into 527 sound classes. We…
AudioSet AudioSet 1 is a large-scale dataset comprising approximately 2 million 10-second YouTube audio clips, categorised into 527 sound classes. We have pre-processed all audio files to a 16 kHz sampling rate and stored them in the WebDataset format for efficient large-scale training and retrieval. Download We recommend using the following commands to download the confit/audioset-16khz-wds dataset from HuggingFace. The dataset is available in two versions: train:… See the full description on the dataset page:
Source: Hugging Face Hub (confit/audioset-16khz-wds). Metadata imported from the dataset’s Hub tags.