Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
CineAudioSynth Synthetic cinematic audio for source separation. 453 scenes, ~22.8 h, 48 kHz / 16-bit / stereo WAV. Each data/scene NNNN/ contains:…
CineAudioSynth Synthetic cinematic audio for source separation. 453 scenes, ~22.8 h, 48 kHz / 16-bit / stereo WAV. Each data/scene NNNN/ contains: linear/ — additive render: mix.wav is the BIT-EXACT 16-bit sum of the four stems (speech, music, ambience, sfx) — max mix − Σstems = 0, verified per scene. Includes gain envelope.json (sidechain ducking envelopes). release/ — mastered render of the same scene (compression/limiting/loudness on the mix bus; intentionally… See the full description on the dataset page:
Source: Hugging Face Hub (disco-eth/cineaudiosynth). Metadata imported from the dataset’s Hub tags.