Skip to content
Advertisement

Data Processing & Curation

This category covers the tooling and platforms that sit between raw data and a training run: ingestion and connector platforms, streaming and change-data-capture pipelines, workflow orchestration, document parsing and extraction, dataset curation and deduplication, quality validation and observability, versioning and lineage for large binary datasets, feature stores, and multimodal exploration and visualisation. The unifying trait is that the customer supplies the data and the vendor supplies the machinery. Domain-specific processing platforms belong here too — medical imaging data management, robotics sensor-log platforms, industrial time-series analytics, geospatial pipeline builders — because they perform the same function inside a single vertical.Vendors here are easy to mis-compare, because the category spans open-source projects with commercial cloud editions, managed services, and enterprise platforms with overlapping feature lists. Questions that separate them include where computation actually happens and whether data must leave the customer’s account; how the tool behaves on unstructured and multimodal payloads rather than on rows and columns, which is where most claims break down; what the failure mode is when a source schema changes; whether pricing scales with data volume, compute, connector count, or seats; how much of the product is genuinely open source and what is withheld for the commercial edition; and whether it integrates with the training and storage stack the team already runs.The nearest adjacent categories are governance and compliance, which concerns who may use data and under what policy rather than moving and reshaping it, and labeling and annotation, where the work is performed by people rather than by pipelines.

61 results

Bespoke Labs

Data Processing & Curation

Open-source data curation tooling for LLM training and evaluation

Text

Graphlit

Data Processing & Curation

Managed ingestion platform for unstructured content in AI agent pipelines

Multimodal

Extend

Data Processing & Curation

Document parsing and extraction API platform for AI workflows

ImageText

Matillion

Data Processing & Curation

Cloud-native ETL platform with AI-driven pipeline automation

Tabular

Bigeye

Data Processing & Curation

Data observability platform for quality monitoring and lineage

Tabular

Chunkr

Data Processing & Curation

Document layout analysis and chunking API for RAG pipelines

ImageText

Prefect

Data Processing & Curation

Workflow orchestration platform for data and ML pipelines

Tabular

Evidently AI

Data Processing & Curation

Open-source evaluation and monitoring framework for data and LLM apps

TabularText

Deepchecks

Data Processing & Curation

AI/data testing and observability platform, part of Check Point

TabularText

Datafold

Data Processing & Curation

Data diffing and quality testing platform for data pipelines

Tabular

Soda

Data Processing & Curation

Data quality monitoring platform with code and no-code interfaces

Tabular

Monte Carlo

Data Processing & Curation

Data and AI observability platform for pipeline and agent monitoring

Tabular

Fivetran

Data Processing & Curation

Automated, managed ELT connectors for cloud data warehouses

Tabular

Cleanlab

Data Processing & Curation

AI reliability platform for detecting and fixing data and agent-response errors

Text

Lightly AI

Data Processing & Curation

Data curation and self-supervised pretraining for computer vision

ImageVideo

Activeloop

Data Processing & Curation

Data engine for versioning and streaming multimodal AI training data

Multimodal

Coactive AI

Data Processing & Curation

Contextual intelligence platform for image and video content libraries

ImageVideo

Unstructured

Data Processing & Curation

ETL platform for turning unstructured documents into LLM-ready data

ImageText

LlamaIndex

Data Processing & Curation

Data framework and hosted parsing platform for LLM document retrieval

Text

Unstract

Data Processing & Curation

LLM-based document data extraction, open source and managed

Text

Datavolo

Data Processing & Curation

Multimodal data pipelines for GenAI, built on Apache NiFi

Multimodal

Airbyte

Data Processing & Curation

Open-source data integration platform with a large connector catalog

Tabular

dltHub

Data Processing & Curation

Open-source Python library and managed runtime for code-first data pipelines

Tabular
Advertisement