Skip to content
Advertisement

Web / Real-time Data

This category covers organisations whose supply comes from the public internet: crawlers, scraping infrastructure, proxy networks, extraction APIs, and the structured databases built on top of them. It includes the raw collection layer of residential and datacentre proxy pools, headless-browser and anti-bot services, and managed crawling operations, along with the derived layer where continuously collected pages become search, e-commerce, news, company, contact, or technology-profile datasets sold as feeds and APIs. What unites them is the collection mechanism rather than the subject matter: the data is publicly reachable, and the vendor’s engineering problem is acquiring it at scale and keeping it fresh.Buyers are usually trading control against effort. A proxy network or scraping API leaves parsing, scheduling, and quality assurance with the buyer but imposes no schema; a managed extraction service or ready-made dataset removes that work but constrains what can be collected. Questions that separate vendors include success rate and latency against the specific target sites that matter rather than aggregate benchmarks; how the IP pool is sourced and how consent is obtained from residential peers; what position the vendor takes on robots directives, terms of service, personal data, and copyright, and whether that position is contractual or merely stated; whether a historical archive exists or collection starts at signup; and how breakages caused by source-site changes are detected and repaired.This category is often confused with data marketplaces, which resell third-party supply, and with proprietary data providers, which originate data through their own sensors, panels, or contractual access rather than from the open web.

53 results

Explorium

Web / Real-time Data

Data platform that automatically matches a company's records to relevant external, largely web-sourced data signals to enrich analytics and ML…

Tabular

Apollo.io

Web / Real-time Data

B2B contact and company database paired with a sales-engagement platform, built on data aggregated and enriched from public web sources.

Tabular

Lusha

Web / Real-time Data

B2B contact-data platform providing verified emails, phone numbers and company profiles sourced from public web data and user contributions.

Tabular

Grepsr

Web / Real-time Data

Managed web-scraping service where an in-house team builds and maintains custom crawlers and delivers structured data.

Text

Outscraper

Web / Real-time Data

Cloud-based scraping platform whose Google Maps scraper extracts business listings, reviews and location data via API or no-code dashboard.

Tabular

ScrapeHero

Web / Real-time Data

Web-scraping service provider offering pre-built and custom data-extraction projects plus alternative-data products for research and investment use.

Text

PromptCloud

Web / Real-time Data

Managed web-scraping / Data-as-a-Service provider that crawls specified sources and delivers structured data feeds on a schedule.

Text

Zyte

Web / Real-time Data

Web-scraping platform and managed data-extraction service, maintainer of the open-source Scrapy framework.

Text

SerpApi

Web / Real-time Data

Real-time API that returns structured JSON results from Google, Bing, YouTube, Amazon and other search engines.

Tabular

ScraperAPI

Web / Real-time Data

Web-scraping API that manages proxy rotation, headless browsers and CAPTCHA solving behind a single request endpoint.

Text

Decodo (Smartproxy)

Web / Real-time Data

Residential, datacenter and mobile proxy network with scraping APIs, rebranded from Smartproxy to Decodo in April 2025.

Text

Diffbot

Web / Real-time Data

Computer-vision-based web-page extraction APIs and a knowledge graph of entities assembled by continuously crawling the web.

Text

Common Crawl

Web / Real-time Data

Nonprofit that crawls the web monthly and freely publishes petabyte-scale WARC/WAT/WET archives, a widely cited pretraining source for language models.

Text

ScrapingBee

Web / Real-time Data

API-based web scraping service offering JavaScript rendering, proxy rotation and a Google-search-results endpoint; acquired by Oxylabs in 2025.

Text

Sequentum

Web / Real-time Data

Low-code enterprise web-scraping platform, formerly sold as Content Grabber, for building and operating large scraper fleets.

Text

Coresignal

Web / Real-time Data

Public web data provider selling structured, API- or file-delivered datasets on companies, professionals and job postings.

Tabular

NetNut

Web / Real-time Data

Proxy network offering residential, ISP, mobile and datacenter IPs for web scraping and data collection.

Text

SOAX

Web / Real-time Data

Proxy network and scraper APIs built on a residential IP pool for public web data collection.

Text

IPRoyal

Web / Real-time Data

Proxy provider offering residential, ISP, datacenter and mobile IPs plus web-unblocking and scraping tools.

Text

Rayobyte

Web / Real-time Data

Proxy network (residential, datacenter, ISP, mobile) and scraping API operator, rebranded from Blazing SEO in 2022.

Text

Infatica

Web / Real-time Data

Proxy network provider (residential, datacenter, mobile) with scraper APIs for web-data collection.

Text

Webshare

Web / Real-time Data

Proxy provider (datacenter, residential, static residential) with a self-serve dashboard and API aimed at developers.

Text

Proxyrack

Web / Real-time Data

Rotating residential and datacenter proxy network used for web scraping, price monitoring and automation.

Text

GeoSurf

Web / Real-time Data

Residential proxy network used for geo-targeted web scraping, price monitoring and ad verification.

Text
Advertisement