Skip to content
Advertisement

10K–100K

748 results

AudioMultimodalText

malaysian-youtube

Malaysian Youtube Malaysian and Singaporean youtube channels, total up to 60k audio files with total 18.7k hours. URLs…

10K–100K·Parquet
Audio

NOTSOFAR

Introduction Welcome to the "NOTSOFAR-1: Distant Meeting Transcription with a Single Device" Challenge. This repo contains the baseline…

10K–100K·Audio (folder)
Audio

Marco_Longspeech

Marco-LongSpeech Dataset Marco-LongSpeech is a multi-task long speech understanding dataset containing 8 different speech understanding tasks…

10K–100K·Apache-2.0
Text

BLiMP

BLiMP

10K–100K·CC-BY·Parquet
Text

pile-10k

The first 10K elements of The Pile, useful for debugging models trained on it. See the HuggingFace page…

10K–100K·Other·Parquet
Text

databricks-dolly-15k

Summary databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several…

10K–100K·CC-BY-SA·JSON
Text

webdev-arena-preference-10k

WebDev Arena Preference Dataset This dataset contains 10K real-world Webdev Arena battle with 10 state-of-the-art LLMs. More details…

10K–100K·Custom / Research-only·JSON
Text

SWE-bench

Dataset Summary SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects…

10K–100K·Parquet
Advertisement