Skip to content
Advertisement

Arrow

15 results

Text

croissant_dataset_no_web_data

CroissantLLM: A Truly Bilingual French-English Language Model Dataset Ressources are currently being uploaded ! Licenses Data redistributed here…

10M–100M·Arrow
Text

Translate_datasets

Datasets From convert to RWKV datasets format It is recommended to shuffle before use Chinese - English &…

10M–100M·MIT·Arrow
Text

croissant_dataset

CroissantLLM: A Truly Bilingual French-English Language Model Dataset Licenses Data redistributed here is subject to the original license…

>1B·Arrow
ImageMultimodalText

XLRS-Bench-lite

🐙GitHub Information or evaluatation on this dataset can be found in this repo: 📜Dataset License Annotations of this…

1K–10K·CC-BY-NC-SA·Arrow
AudioMultimodalText

VieNeu-TTS-140h

pnnbao-ump/VieNeu-TTS-140h Mô tả Dataset A high-quality Vietnamese Text-to-Speech (TTS) dataset containing 74,858 audio samples with phonemized…

10K–100K·Apache-2.0·Arrow
AudioMultimodalText

vieneu-tts-140h-dataset

pnnbao-ump/VieNeu-TTS-140h Mô tả Dataset Dataset tiếng Việt chất lượng cao cho Text-to-Speech (TTS) với 74,858 mẫu audio và…

10K–100K·Apache-2.0·Arrow
MultimodalTabularText

smollm-chunked

FAISS Indices and Chunked Datasets for SmolLM and SmolLM2 corpora This repository contains part of the FAISS indices…

100M–1B·ODC-BY·Arrow
AudioMultimodalText

BigAudioDataset

AstraMindAI/BigAudioDataset Dataset Description AstraMindAI/BigAudioDataset is a large-scale, multilingual dataset designed for a wide range of audio…

1M–10M·Apache-2.0·Arrow
ImageMultimodalText

kb-books

open-rdl-books Dataset Description Language dan, dansk, Danish License Public Domain, cc0-1.0 Dataset Summary Documents from the Royal Danish…

1M–10M·CC0·Arrow
Advertisement