Web / Real-time Data
This category covers organisations whose supply comes from the public internet: crawlers, scraping infrastructure, proxy networks, extraction APIs, and the structured databases built on top of them. It includes the raw collection layer of residential and datacentre proxy pools, headless-browser and anti-bot services, and managed crawling operations, along with the derived layer where continuously collected pages become search, e-commerce, news, company, contact, or technology-profile datasets sold as feeds and APIs. What unites them is the collection mechanism rather than the subject matter: the data is publicly reachable, and the vendor’s engineering problem is acquiring it at scale and keeping it fresh.Buyers are usually trading control against effort. A proxy network or scraping API leaves parsing, scheduling, and quality assurance with the buyer but imposes no schema; a managed extraction service or ready-made dataset removes that work but constrains what can be collected. Questions that separate vendors include success rate and latency against the specific target sites that matter rather than aggregate benchmarks; how the IP pool is sourced and how consent is obtained from residential peers; what position the vendor takes on robots directives, terms of service, personal data, and copyright, and whether that position is contractual or merely stated; whether a historical archive exists or collection starts at signup; and how breakages caused by source-site changes are detected and repaired.This category is often confused with data marketplaces, which resell third-party supply, and with proprietary data providers, which originate data through their own sensors, panels, or contractual access rather than from the open web.
53 results
Outscraper
Web / Real-time DataCloud-based scraping platform whose Google Maps scraper extracts business listings, reviews and location data via API or no-code dashboard.
TabularScrapeHero
Web / Real-time DataWeb-scraping service provider offering pre-built and custom data-extraction projects plus alternative-data products for research and investment use.
TextPromptCloud
Web / Real-time DataManaged web-scraping / Data-as-a-Service provider that crawls specified sources and delivers structured data feeds on a schedule.
TextScraperAPI
Web / Real-time DataWeb-scraping API that manages proxy rotation, headless browsers and CAPTCHA solving behind a single request endpoint.
TextDecodo (Smartproxy)
Web / Real-time DataResidential, datacenter and mobile proxy network with scraping APIs, rebranded from Smartproxy to Decodo in April 2025.
TextCommon Crawl
Web / Real-time DataNonprofit that crawls the web monthly and freely publishes petabyte-scale WARC/WAT/WET archives, a widely cited pretraining source for language models.
TextScrapingBee
Web / Real-time DataAPI-based web scraping service offering JavaScript rendering, proxy rotation and a Google-search-results endpoint; acquired by Oxylabs in 2025.
TextCoresignal
Web / Real-time DataPublic web data provider selling structured, API- or file-delivered datasets on companies, professionals and job postings.
Tabular