Skip to content
Advertisement

100K–1M

595 results

Audio

emolia-thinking-balanced-buckets

Emolia-Thinking — Balanced Per-Dimension Bucket Subset A balanced, per-dimension bucket subset of VoiceNet/emolia-thinking, derived from that…

100K–1M·CC-BY
AudioMultimodalText

BibleMMS

The Dataset associated with the Paper "Meta Learning Text-to-Speech Synthesis in over 7000 Languages" by Florian Lux, Sarina…

100K–1M·MIT·Parquet
AudioMultimodalText

parczech4speech-segmented

ParCzech4Speech (Sentence-Segmented Variant) Dataset Summary ParCzech4Speech (Sentence-Segmented Variant) is a large-scale Czech speech dataset based…

100K–1M·CC-BY·WebDataset
AudioMultimodalText

INTP

INTP: Intelligibility Preference Speech Dataset We establish a synthetic Intelligibility Preference Speech Dataset (INTP), including about 250K…

100K–1M·CC-BY-NC·Parquet
MultimodalTabularText

Talker-T2AV-Data

Talker-T2AV-Data Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling Paper (arXiv 2604.23586) · Code (GitHub) · Model ·…

100K–1M·Apache-2.0·CSV
ImageMultimodalText

hescape-pyarrow

HESCAPE • PyArrow Format HESCAPE (H&E + Spatial Contrastive Pretraining Benchmark) is a large-scale benchmark for multimodal learning…

100K–1M·CC-BY-NC-SA·Parquet
MultimodalTabularText

CanadaFireSat

Dataset Card for CanadaFireSat 🔥🛰️ In this benchmark, we investigate the potential of deep learning with multiple modalities…

100K–1M·MIT·Parquet
ImageMultimodalText

CommonForms

CommonForms: A Large, Diverse Dataset for Form Field Detection This repository hosts the CommonForms dataset, a web-scale dataset…

100K–1M·Apache-2.0·Parquet
ImageMultimodalText

deep-scores-v2

DeepScoresV2 — Complete A HuggingFace-formatted mirror of the complete version of the DeepScoresV2 dataset for music object detection.…

100K–1M·CC-BY·Parquet
Advertisement