Skip to content
Advertisement

Japanese

25 results

MultimodalTabularText

JMedBench

Maintainers Junfeng Jiang@Aizawa Lab: jiangjf (at) is.s.u-tokyo.ac.jp Jiahao Huang@Aizawa Lab: jiahao-huang (at) g.ecc.u-tokyo.ac.jp If you find any…

100K–1M·JSON
ImageMultimodalText

Cauldron-JA

Dataset Card for The Cauldron-JA Dataset description The Cauldron-JA is a Vision Language Model dataset that translates 'The…

1M–10M·CC-BY·Parquet
AudioMultimodalText

Irodori-Ja-Spk2-10k

SynDataLab/Irodori-Ja-Spk2-10k 10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker…

10K–100K·CC-BY-NC-SA·Parquet
AudioMultimodalText

irodori-refs-10k-v2

Irodori TTS Reference Voices v2 (10K) 10,000 reference voices generated with Aratako/Irodori-TTS-500M-v2-VoiceDesign (no ref=True) using a richer…

10K–100K·Apache-2.0·Parquet
AudioMultimodalText

Irodori-Ja-Spk1-10k

SynDataLab/Irodori-Ja-Spk1-10k 10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker…

10K–100K·CC-BY-NC-SA·Parquet
AudioMultimodalText

MiscSpeech-ja

MiscSpeech-ja This dataset comprises audio and corresponding transcripts collected from a diverse range of YouTube videos and Podcasts.

10K–100K·CC-BY-SA·Parquet
AudioMultimodalText

spoken-magpie-ja

Spoken-magpie LLMの日本語Instruction Tuning用データllm-jp/magpie-sft-v1.0をCosyVoice2 TTSを使用して音声化した商用利用可能な日本語の音声言語モデルのSFT用データセットです。 ある程度の話者多様性を持つように生成されています。…

100K–1M·Apache-2.0·Parquet
AudioMultimodalText

sayoko-tts-corpus

サヨ子 音声コーパス ダウンロード方法 データセットを圧縮したzipファイルを、gdriveに置いています。 import gdown url = " RcBAaAyRwEIOWuTQFetVaMUU" gdown.download( url,…

<1K·CC-BY·Text (raw)
AudioMultimodalText

SaSLaW

This repository contains the data of SaSLaW corpus. You can download it via the following command: huggingface-cli download…

1K–10K·CC-BY-NC·Audio (folder)
AudioMultimodalText

irodori-tts-refs-12k

Irodori TTS Reference Voices (12,160) 12,160 Japanese reference voices generated with Aratako/Irodori-TTS-500M-v2-VoiceDesign. field type description…

10K–100K·Apache-2.0·Parquet
MultimodalTabularText

fineweb-2-edu-japanese

🍷 FineWeb2 Edu Japanese: High-Quality Educational Japanese Dataset This dataset consists of 120 million texts (approximately 89.3B tokens)…

100M–1B·ODC-BY·Parquet
AudioMultimodalText

Galgame-VisualNovel-Reupload

Galgame VisualNovel Reupload This repository is a reupload of the visual novel dataset OOPPEENN/56697375616C4E6F76656C5F44617461736574. The goal of…

1M–10M·Custom / Research-only·Parquet
AudioMultimodalText

SoulTide-AudioData-Dataset

目录结构 character/ └── char / ├── resource/ │ ├── audio/ 原始音频 │ ├── srt/ 原始srt │ └── processed/…

1K–10K·CC0·Audio (folder)
Advertisement