Skip to content
Advertisement
AudioMultimodalText

alconost-multilingual-speech-gold

Multilingual Speech & Translation Dataset — EN↔JA/AR-EG/PL/RU (10 phrases, dual-take) Description 10 English source phrases with expert human…

Multilingual Speech & Translation Dataset — EN↔JA/AR-EG/PL/RU (10 phrases, dual-take) Description 10 English source phrases with expert human translations into Japanese, Egyptian Arabic (ar-EG), and Polish. Each target phrase is recorded by native speakers (two takes each). Audio files are WAV 48 kHz mono, 16‑bit PCM format. Translations are produced and QA’d by professional linguists; recordings follow consistent orthography/style (AR-EG: Egyptian dialect; JA/PL: standard). All… See the full description on the dataset page:

Source: Hugging Face Hub (alconost/alconost-multilingual-speech-gold). Metadata imported from the dataset’s Hub tags.

Advertisement