croissant_dataset_no_web_data
CroissantLLM: A Truly Bilingual French-English Language Model Dataset Ressources are currently being uploaded ! Licenses Data redistributed here…
15 results
CroissantLLM: A Truly Bilingual French-English Language Model Dataset Ressources are currently being uploaded ! Licenses Data redistributed here…
Datasets From convert to RWKV datasets format It is recommended to shuffle before use Chinese - English &…
CroissantLLM: A Truly Bilingual French-English Language Model Dataset Licenses Data redistributed here is subject to the original license…
🐙GitHub Information or evaluatation on this dataset can be found in this repo: 📜Dataset License Annotations of this…
Bambara-ASR-All Audio Dataset
pnnbao-ump/VieNeu-TTS-140h Mô tả Dataset A high-quality Vietnamese Text-to-Speech (TTS) dataset containing 74,858 audio samples with phonemized…
Dataset Card for Dataset Name This dataset is collected from youtube.
pnnbao-ump/VieNeu-TTS-140h Mô tả Dataset Dataset tiếng Việt chất lượng cao cho Text-to-Speech (TTS) với 74,858 mẫu audio và…
🐙GitHub Information or evaluatation on this dataset can be found in this repo: 📜Dataset License Annotations of this…
FAISS Indices and Chunked Datasets for SmolLM and SmolLM2 corpora This repository contains part of the FAISS indices…
AstraMindAI/BigAudioDataset Dataset Description AstraMindAI/BigAudioDataset is a large-scale, multilingual dataset designed for a wide range of audio…
LOTSA Data The Large-scale Open Time Series Archive (LOTSA) is a collection of open time series datasets for…
open-rdl-books Dataset Description Language dan, dansk, Danish License Public Domain, cc0-1.0 Dataset Summary Documents from the Royal Danish…