Skip to content
Advertisement

Common Crawl is a 501(c)(3) nonprofit, founded in 2007 by Gil Elbaz, that has crawled the web on a roughly monthly cadence since 2008 and published the archives at no cost through Amazon Web Services’ Open Data program. Each crawl covers on the order of billions of pages, distributed as raw WARC, metadata WAT and plain-text WET files.

Text
Advertisement