1M–10M
329 results
TAVGBench_1m
Installation Download this repo to a local folder, and unzip these .zip files under the TAVGBench 1m/data/. Then,…
open-web-math
Keiran Paster , Marco Dos Santos , Zhangir Azerbayev, Jimmy Ba GitHub ArXiv PDF OpenWebMath is a dataset…
lotsa_data
LOTSA Data The Large-scale Open Time Series Archive (LOTSA) is a collection of open time series datasets for…
Argimi-Ardian-Finance-10k-text
The ArGiMI Ardian datasets : Text-only version The ArGiMi project is committed to open-source principles and data sharing.…
codeparrot-clean
CodeParrot 🦜 Dataset Cleaned What is it? A dataset of Python files from Github. This is the deduplicated…
covost2
This is a partial copy of CoVoST2 dataset. The main difference is that the audio data is included…
Vchitect_T2V_DataVerse
Vchitect-T2V-Dataverse Vchitect Team1 1Shanghai Artificial Intelligence Laboratory Paper Project Page Data Overview The Vchitect-T2V-Dataverse is the…
LLaVA-Video-178K
Dataset Card for LLaVA-Video-178K Uses This dataset is used for the training of the LLaVA-Video model. We only…
yodas2_sidon
YODAS2-Sidon Overview This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis…
Goodreads-Books
Dataset Card for "BrightData/Goodreads-Books" Dataset Summary Explore a collection of millions of books with the Goodreads dataset, comprising…
InsViE
InsViE-1M: Effective Instruction-based Video Editing with Elaborate Dataset Construction Citation If you find this work helpful, please consider…