Skip to content
Advertisement
MultimodalTabularText

Open Markdown

Open Markdown

Open Markdown Clean markdown from the web, ready for training and retrieval What is it? Open Markdown is a large-scale web text dataset built from Common Crawl. Common Crawl is a non-profit that crawls the web and freely provides its archives to the public. Every page goes through a pipeline that extracts the main content from raw HTML, converts it to clean Markdown, and packages the result into Parquet files with WARC metadata for traceability. The dataset… See the full description on the dataset page:

Source: Hugging Face Hub (open-index/open-markdown-v2). Metadata imported from the dataset’s Hub tags.

Advertisement