Skip to content
Advertisement
AudioMultimodalText

VOX-DUB

VOX-DUB is a human-based benchmark for evaluating AI dubbing systems.It includes: Audio fragments with original speech from real videos and their…

VOX-DUB is a human-based benchmark for evaluating AI dubbing systems.It includes: Audio fragments with original speech from real videos and their corresponding translated texts. Generated audio recordings produced by multiple dubbing/TTS systems. Human annotation results with pairwise A/B (+ SAME) evaluations across five aspects (pronunciation, naturalness, sound quality, emotion similarity, and voice similarity). Detailed annotation guidelines with examples for pairwise A/B… See the full description on the dataset page:

Source: Hugging Face Hub (toloka/VOX-DUB). Metadata imported from the dataset’s Hub tags.

Advertisement