Skip to content
Advertisement

Web / Real-time Data

This category covers organisations whose supply comes from the public internet: crawlers, scraping infrastructure, proxy networks, extraction APIs, and the structured databases built on top of them. It includes the raw collection layer of residential and datacentre proxy pools, headless-browser and anti-bot services, and managed crawling operations, along with the derived layer where continuously collected pages become search, e-commerce, news, company, contact, or technology-profile datasets sold as feeds and APIs. What unites them is the collection mechanism rather than the subject matter: the data is publicly reachable, and the vendor’s engineering problem is acquiring it at scale and keeping it fresh.Buyers are usually trading control against effort. A proxy network or scraping API leaves parsing, scheduling, and quality assurance with the buyer but imposes no schema; a managed extraction service or ready-made dataset removes that work but constrains what can be collected. Questions that separate vendors include success rate and latency against the specific target sites that matter rather than aggregate benchmarks; how the IP pool is sourced and how consent is obtained from residential peers; what position the vendor takes on robots directives, terms of service, personal data, and copyright, and whether that position is contractual or merely stated; whether a historical archive exists or collection starts at signup; and how breakages caused by source-site changes are detected and repaired.This category is often confused with data marketplaces, which resell third-party supply, and with proprietary data providers, which originate data through their own sensors, panels, or contractual access rather than from the open web.

53 results

Bright Data

Web / Real-time Data
Featured

Web-data platform offering a large proxy network, scraping APIs and ready-made datasets used to source and structure public web data…

ImageMultimodalText

Similarweb

Web / Real-time Data

Digital data and intelligence platform for web and app traffic.

Tabular

Import.io

Web / Real-time Data

No-code, point-and-click platform for converting web pages into structured, exportable data, in operation since 2012 and now owned by Neuralogics.

Text

Cognism

Web / Real-time Data

Sales-intelligence platform providing verified B2B contact and company data plus buying-intent signals, compiled from public and licensed sources.

Tabular

DataHen

Web / Real-time Data

Managed data-crawling and web-scraping service for custom, enterprise-scale extraction projects.

Text

Datahut

Web / Real-time Data

Fully managed web-scraping service delivering structured data across e-commerce, real estate and other verticals.

Text

Scrapingdog

Web / Real-time Data

Web-scraping API handling proxy rotation, browser rendering and CAPTCHA solving for search engines, e-commerce sites and social profiles.

Text

ScrapingAnt

Web / Real-time Data

Web-scraping API offering headless-browser rendering, rotating proxies and an AI-based data-extraction endpoint.

Text

Web Scraper (webscraper.io)

Web / Real-time Data

Free Chrome/Firefox extension for building scrapers visually, paired with a paid cloud service for scheduled, large-scale runs.

Text

ZenRows

Web / Real-time Data

Web-scraping toolkit combining a scraper API, anti-bot bypass and a hosted Puppeteer/Playwright scraping browser.

Text

Scrapfly

Web / Real-time Data

Web scraping API offering anti-bot bypass, JavaScript rendering and screenshot capture, self-funded since 2020.

Text

Crawlbase

Web / Real-time Data

Web scraping and crawling API platform (proxy, scraper and storage products), rebranded from ProxyCrawl in 2022.

Text

Octoparse

Web / Real-time Data

No-code point-and-click web scraper (desktop and cloud) for building and scheduling data-extraction workflows without programming.

Text

ParseHub

Web / Real-time Data

Desktop web-scraping application with a point-and-click interface designed to handle JavaScript-heavy and infinite-scroll pages.

Text

Browse AI

Web / Real-time Data

No-code platform for extracting data from websites and monitoring pages for changes, using trained 'robots' instead of manual selectors.

Text

BuiltWith

Web / Real-time Data

Technology-profiling service that scans websites to report the CMS, analytics, advertising and hosting technologies they run.

Tabular

Wappalyzer

Web / Real-time Data

Technology-detection tool (browser extension and API) that inspects a site's code to identify its CMS, frameworks and other software.

Tabular

DomainTools

Web / Real-time Data

DNS, WHOIS and domain-intelligence platform built from continuously collected domain registration and infrastructure data.

Tabular

Majestic

Web / Real-time Data

SEO backlink-analysis platform built on an independently operated web crawler that has indexed links since 2004; publishes the open Majestic…

Tabular

Webz.io

Web / Real-time Data

Web-data API provider, formerly Webhose.io, delivering structured feeds from open, deep and dark web sources for news, cybersecurity and research…

Text

NewsCatcher

Web / Real-time Data

API that continuously crawls and structures news articles from over 100,000 web sources for monitoring and dataset-building use cases.

Text

The GDELT Project

Web / Real-time Data

Open, freely downloadable database that monitors global news media and encodes people, locations, organizations and events into structured records, updated…

Text

People Data Labs

Web / Real-time Data

Person- and company-data API used to enrich, search and match records against a database built from public web and licensed…

Tabular
Advertisement