Skip to content
Advertisement

100M–1B

93 results

Text

OmniCorpus-CC

🐳 OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text ⭐️ NOTE: Several parquet files…

100M–1B·CC-BY·Parquet
Text

LEMAS-Dataset-train

Overview This dataset is part of LEMAS-Project (lemas-project.github.io/LEMAS-Project). It contains a large-scale training set (150k+ hours) and a…

100M–1B·CC-BY·JSON
Text

OpenScan

OpenScan

100M–1B·Custom / Research-only·Text (raw)
MultimodalSensor / Time-seriesTabular

binance-futures-ohlcv-2018-2026

🐱 币安期货 Main4 数据集 (BTC/ETH/BNB/SOL) 本项目由交易猫基金会支持(交易猫基金会 CA:0x8a99b8d53eff6bc331af529af74ad267f3167777)。 精简版币安期货历史数据 - 只包含 4 个主要币种,适合快速下载和研究使用。 📊 数据概览…

100M–1B·MIT·CSV
3D / Point Cloud

Imaginarium-Dataset

Imaginarium: Vision-guided High-Quality 3D Scene Layout Generation SIGGRAPH ASIA 2025 & ACM Transactions on Graphics (TOG) 📖 Introduction…

100M–1B
MultimodalTabularText

IndustryCorpus2

Industry models play a vital role in promoting the intelligent transformation and innovative development of enterprises. High-quality industry…

100M–1B·Apache-2.0·Parquet
MultimodalTabularText

stack-edu

💻 Stack-Edu Stack-Edu is a 125B token dataset of educational code filtered from The Stack v2, precisely the…

100M–1B·Parquet
Advertisement