Skip to content
Advertisement

Summarization

34 results

MultimodalTabularText

M4LE

Introduction M4LE is a Multi-ability, Multi-range, Multi-task, bilingual benchmark for long-context evaluation. We categorize long-context…

10K–100K·MIT·JSON
Text

UltraLink

multi-lingual, knowledge-grounded, multi-round dialogue dataset and model Summary • Construction Process • Paper • UltraLink-LM • Github Dataset…

1M–10M·MIT·JSON
Image

chemrXiv-pdf

ChemrXiv Pdf Introducing ChemrXiv Pdf, a dataset that offers access to all PDFs published until September 15, 2024.…

1K–10K·Custom / Research-only
Text

GH_text2code

Docstring to code data Dataset Summary This dataset contains pairs of English text and code from multiple programming…

10M–100M·Parquet
MultimodalTabularText

bhasha-sft

Bhasha SFT Bhasha SFT is a massive collection of multiple open sourced Supervised Fine-Tuning datasets for training Multilingual…

10M–100M·Mixed·Parquet
AudioMultimodalText

dowis

Do What I Say (DOWIS): A Spoken Prompt Dataset for Instruction-Following NEW DOWIS now also contains spoken and…

1K–10K·CC-BY·Parquet
Advertisement