Skip to content
Advertisement

JSON

248 results

Text

TinyStories-Multilingual

Novelist: TinyStories Multilingual Edition Dataset Summary The TinyStories Multilingual Edition is a high-fidelity synthetic dataset of short,…

10K–100K·Apache-2.0·JSON
Text

tatoeba-bitext-mining

Tatoeba An MTEB dataset Massive Text Embedding Benchmark 1,000 English-aligned sentence pairs for each language based on the…

100K–1M·CC-BY·JSON
Text

wmt24pp

WMT24++ This repository contains the human translation and post-edit data for the 55 en- xx language pairs released…

10K–100K·Apache-2.0·JSON
Text

multi30k

Multi30k This dataset contains the "multi30k" dataset, which is the "task 1" dataset from here. Each example consists…

10K–100K·JSON
MultimodalTextVideo

V1-33K-Old

V1: Toward Multimodal Reasoning by Designing Auxiliary Tasks 🚀 Toward Multimodal Reasoning via Unsupervised Task -- Future Prediction…

10K–100K·Apache-2.0·JSON
ImageMultimodalText

SDG-30K

SDG-30K — Structured Defect Grounding Dataset A 30,000-image dataset for structured defect grounding in text-to-image generations. Each image…

10K–100K·CC-BY-NC·JSON
Advertisement