Skip to content
Advertisement
AudioMultimodalText

irodori-refs-10k-v2

Irodori TTS Reference Voices v2 (10K) 10,000 reference voices generated with Aratako/Irodori-TTS-500M-v2-VoiceDesign (no ref=True) using a richer…

Irodori TTS Reference Voices v2 (10K) 10,000 reference voices generated with Aratako/Irodori-TTS-500M-v2-VoiceDesign (no ref=True) using a richer caption space than v1: 8 axes (gender × age × pitch × tone × speed × distance × emotion × quality) with per-voice unique caption combinations, plus gender alternation, an incompatibility filter (no contradictory “whisper + speak loudly” combos), and a similarity-rejection window so consecutive voices stay distinct. Each ref’s text… See the full description on the dataset page:

Source: Hugging Face Hub (SynDataLab-JA-Refs/irodori-refs-10k-v2). Metadata imported from the dataset’s Hub tags.

Advertisement