Skip to content
Advertisement
AudioMultimodalText

vctk

VCTK This is a processed clone of the VCTK dataset with leading and trailing silence removed using Silero VAD. A fixed 25 ms of padding has been…

VCTK This is a processed clone of the VCTK dataset with leading and trailing silence removed using Silero VAD. A fixed 25 ms of padding has been added to both ends of each audio clip to (hopefully) imrprove training and finetuning. The original dataset is available at: Reproducing This repository notably lacks a requirements.txt file. There’s likely a missing dependency or two, but roughly: pydub tqdm torch torchaudio… See the full description on the dataset page:

Source: Hugging Face Hub (jspaulsen/vctk). Metadata imported from the dataset’s Hub tags.

Advertisement