Bespoke Labs
Data Processing & CurationOpen-source data curation tooling for LLM training and evaluation
TextThis category covers the tooling and platforms that sit between raw data and a training run: ingestion and connector platforms, streaming and change-data-capture pipelines, workflow orchestration, document parsing and extraction, dataset curation and deduplication, quality validation and observability, versioning and lineage for large binary datasets, feature stores, and multimodal exploration and visualisation. The unifying trait is that the customer supplies the data and the vendor supplies the machinery. Domain-specific processing platforms belong here too — medical imaging data management, robotics sensor-log platforms, industrial time-series analytics, geospatial pipeline builders — because they perform the same function inside a single vertical.Vendors here are easy to mis-compare, because the category spans open-source projects with commercial cloud editions, managed services, and enterprise platforms with overlapping feature lists. Questions that separate them include where computation actually happens and whether data must leave the customer’s account; how the tool behaves on unstructured and multimodal payloads rather than on rows and columns, which is where most claims break down; what the failure mode is when a source schema changes; whether pricing scales with data volume, compute, connector count, or seats; how much of the product is genuinely open source and what is withheld for the commercial edition; and whether it integrates with the training and storage stack the team already runs.The nearest adjacent categories are governance and compliance, which concerns who may use data and under what policy rather than moving and reshaping it, and labeling and annotation, where the work is performed by people rather than by pipelines.
61 results
Open-source data curation tooling for LLM training and evaluation
TextManaged ingestion platform for unstructured content in AI agent pipelines
MultimodalOpen-source evaluation and monitoring framework for data and LLM apps
TabularTextAI/data testing and observability platform, part of Check Point
TabularTextData and AI observability platform for pipeline and agent monitoring
TabularData curation and self-supervised pretraining for computer vision
ImageVideoData engine for versioning and streaming multimodal AI training data
MultimodalContextual intelligence platform for image and video content libraries
ImageVideoETL platform for turning unstructured documents into LLM-ready data
ImageTextData framework and hosted parsing platform for LLM document retrieval
TextMultimodal data pipelines for GenAI, built on Apache NiFi
MultimodalDistributed dataframe engine for multimodal AI data processing
Multimodal