Web-data platform offering a large proxy network, scraping APIs and ready-made datasets used to source and structure public web data…
ImageMultimodalTextWeb / Real-time Data
This category covers organisations whose supply comes from the public internet: crawlers, scraping infrastructure, proxy networks, extraction APIs, and the structured databases built on top of them. It includes the raw collection layer of residential and datacentre proxy pools, headless-browser and anti-bot services, and managed crawling operations, along with the derived layer where continuously collected pages become search, e-commerce, news, company, contact, or technology-profile datasets sold as feeds and APIs. What unites them is the collection mechanism rather than the subject matter: the data is publicly reachable, and the vendor’s engineering problem is acquiring it at scale and keeping it fresh.Buyers are usually trading control against effort. A proxy network or scraping API leaves parsing, scheduling, and quality assurance with the buyer but imposes no schema; a managed extraction service or ready-made dataset removes that work but constrains what can be collected. Questions that separate vendors include success rate and latency against the specific target sites that matter rather than aggregate benchmarks; how the IP pool is sourced and how consent is obtained from residential peers; what position the vendor takes on robots directives, terms of service, personal data, and copyright, and whether that position is contractual or merely stated; whether a historical archive exists or collection starts at signup; and how breakages caused by source-site changes are detected and repaired.This category is often confused with data marketplaces, which resell third-party supply, and with proprietary data providers, which originate data through their own sensors, panels, or contractual access rather than from the open web.
53 results
Similarweb
Web / Real-time DataDigital data and intelligence platform for web and app traffic.
TabularScrapingdog
Web / Real-time DataWeb-scraping API handling proxy rotation, browser rendering and CAPTCHA solving for search engines, e-commerce sites and social profiles.
TextScrapingAnt
Web / Real-time DataWeb-scraping API offering headless-browser rendering, rotating proxies and an AI-based data-extraction endpoint.
TextWeb Scraper (webscraper.io)
Web / Real-time DataFree Chrome/Firefox extension for building scrapers visually, paired with a paid cloud service for scheduled, large-scale runs.
TextWappalyzer
Web / Real-time DataTechnology-detection tool (browser extension and API) that inspects a site's code to identify its CMS, frameworks and other software.
TabularDomainTools
Web / Real-time DataDNS, WHOIS and domain-intelligence platform built from continuously collected domain registration and infrastructure data.
TabularNewsCatcher
Web / Real-time DataAPI that continuously crawls and structures news articles from over 100,000 web sources for monitoring and dataset-building use cases.
TextThe GDELT Project
Web / Real-time DataOpen, freely downloadable database that monitors global news media and encodes people, locations, organizations and events into structured records, updated…
TextPeople Data Labs
Web / Real-time DataPerson- and company-data API used to enrich, search and match records against a database built from public web and licensed…
Tabular