Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
Multi-lingual TTS Data (leeoxiang/multi lingo data) Large-scale cross-lingual + same-lingual TTS corpus for training expressive multi-lingual…
Multi-lingual TTS Data (leeoxiang/multi lingo data) Large-scale cross-lingual + same-lingual TTS corpus for training expressive multi-lingual voice-clone models. Covers 4 target languages (en / ja / ko / zh), synthesized by bosonai/higgs-tts-3-4b via sglang-omni. 🚧 Upload in progress — the main samelang expressive corpus (≈ 4.78 M rows / ≈ 3.3 TB) is being pushed as parquet shards. Expect the shard count to grow over the next hours/days until each language reaches… See the full description on the dataset page: lingo data.
Source: Hugging Face Hub (leeoxiang/multi_lingo_data). Metadata imported from the dataset’s Hub tags.