Skip to content
Advertisement

Text Generation

217 results

Text

Danbooruwildcards

This is a set of wildcards for danbooru tags. Artist:Prompts for random artist styles, covering approximately 0.6M different…

10M–100M·Custom / Research-only·Text (raw)
Text

nemotron-cc-translated

Helsinki-NLP/nemotron-cc-translated nemotron-cc-tanslated is a collection of automatically translated documents from nemotron-cc taken out of the…

>1B·CC0·Parquet
Text

MegaMath

MegaMath: Pushing the Limits of Open Math Copora Megamath is part of TxT360, curated by LLM360 Team. We…

100M–1B·ODC-BY·Parquet
Text

fineweb-edu-translated

Helsinki-NLP/fineweb-edu-translated fineweb-edu-tanslated is a collection of automatically translated documents from fineweb-edu. Translations are…

>1B·ODC-BY·Parquet
Text

demo_data

1,000 examples from gpt4 en 1,000 examples from gpt4 zh 300 examples from toolcall en 300 examples from…

1K–10K·Apache-2.0
MultimodalTabularText

Goodreads-Books

Dataset Card for "BrightData/Goodreads-Books" Dataset Summary Explore a collection of millions of books with the Goodreads dataset, comprising…

1M–10M·Custom / Research-only·CSV
MultimodalTabularText

AutoMathText-V2

🚀 AutoMathText-V2: A 2.46 Trillion Token AI-Curated STEM Pretraining Dataset   🎉 AutoMathText-v2 has surpassed 1.5 million downloads!…

100M–1B
Text

TinyStories

Dataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary. Described in…

1M–10M·Other·Parquet
Text

Alpaca

Alpaca

10K–100K·CC-BY-NC·Parquet
Advertisement