Skip to content
Advertisement
MultimodalTabularText

MintTTS Pre-tokenized Audio (somu9/hindi-hq)

MintTTS Pre-tokenized Audio (somu9/hindi-hq)

MintTTS Pre-tokenized Audio Tokens Pre-extracted audio codec tokens for TTS training. Source Dataset: somu9/hindi-hq Codec: MOSS-Audio-Tokenizer-Nano Codec sample rate: 48,000 Hz (stereo) Frame rate: 12.5 Hz (1 frame = 80ms) Stats Metric Value Total samples 439,507 Total audio hours 811.7h Codebooks 16 Avg frames/sample 83.1 Avg duration 6.6s Format JSONL file (manifest.jsonl) where each line is: { “text”:… See the full description on the dataset page:

Source: Hugging Face Hub (somu9/hindi-hq-tokens). Metadata imported from the dataset’s Hub tags.

Advertisement