Skip to content
Advertisement

Text

2,175 results

MultimodalTabularText

KodCode-V1

🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding KodCode is the largest fully-synthetic open-source dataset…

100K–1M·CC-BY-NC·Parquet
MultimodalTabularText

dojo_quote

Languages: 简体中文 · English dojo quote — Latest Quote Snapshot Overview Point-in-time snapshot for the most recent trading…

10K–100K·Apache-2.0·Parquet
MultimodalTabularText

uci-har-federated

UCI Human Activity Recognition (HAR) Dataset Dataset Description The UCI Human Activity Recognition dataset is a widely-used benchmark…

10K–100K·MIT·Parquet
MultimodalTabularText

pg-en

Overview Property Value Source Project Gutenberg (English catalog) Snapshot 2026-07-02-18-47-04 Total files 50871 Total Tokens (BPE) ~7.14 billion…

10K–100K·MIT·Parquet
MultimodalTabularText

binance-btcusdt

BTCUSDT Perpetual Futures — 5-Minute Feature Dataset Complete historical dataset for Binance BTCUSDT USDT-Margined Perpetual Futures, covering…

>1B·MIT·Parquet
MultimodalTabularText

Stocks-Daily-Price

Dataset Information This dataset includes daily price data for various stocks. Instruments Included 7000+ US Stocks Dataset Columns…

10M–100M·Custom / Research-only·Parquet
MultimodalTabularText

the-heap

The Heap Dataset We develop The Heap, a new contamination-free multilingual code dataset comprising 57 languages, which facilitates…

10M–100M·GPL-3.0·Parquet
Advertisement