Core-VIIRS-Nighttime-Light
Major TOM Core VIIRS Nighttime Light Annual radiance composites from the Visible Infrared Imaging Radiometer Suite (VIIRS) Day/Night…
Emolia · Filtered · NanoCodec (FSQ) Tokens A cleaned, pre-tokenized version of laion/Emolia prepared for text-to-speech (TTS) training. The pipeline is two steps: Quality filtering with the open-source audio filter tool — this removes the dirtiest recordings (noise, clipping, band-limiting, robotic artifacts, overlapping speakers), which matters a lot for TTS quality. Discrete audio tokenization with NVIDIA nvidia/nemo-nano-codec-22khz-1.89kbps-21.5fps (an FSQ neural audio… See the full description on the dataset page: data.
Source: Hugging Face Hub (chukypedro/english_data). Metadata imported from the dataset’s Hub tags.