Oxen.ai
Data Processing & CurationGit-like version control for large AI datasets and model files
MultimodalThis category covers the tooling and platforms that sit between raw data and a training run: ingestion and connector platforms, streaming and change-data-capture pipelines, workflow orchestration, document parsing and extraction, dataset curation and deduplication, quality validation and observability, versioning and lineage for large binary datasets, feature stores, and multimodal exploration and visualisation. The unifying trait is that the customer supplies the data and the vendor supplies the machinery. Domain-specific processing platforms belong here too — medical imaging data management, robotics sensor-log platforms, industrial time-series analytics, geospatial pipeline builders — because they perform the same function inside a single vertical.Vendors here are easy to mis-compare, because the category spans open-source projects with commercial cloud editions, managed services, and enterprise platforms with overlapping feature lists. Questions that separate them include where computation actually happens and whether data must leave the customer’s account; how the tool behaves on unstructured and multimodal payloads rather than on rows and columns, which is where most claims break down; what the failure mode is when a source schema changes; whether pricing scales with data volume, compute, connector count, or seats; how much of the product is genuinely open source and what is withheld for the commercial edition; and whether it integrates with the training and storage stack the team already runs.The nearest adjacent categories are governance and compliance, which concerns who may use data and under what policy rather than moving and reshaping it, and labeling and annotation, where the work is performed by people rather than by pipelines.
61 results
Git-like version control for large AI datasets and model files
MultimodalReal-time feature computation and serving platform for ML
Sensor / Time-seriesTabularGit-like version control for data lakes and AI-ready data
MultimodalOpen-source data quality validation framework for data pipelines
TabularHealthcare NLP models, de-identification tooling, and clinical data for AI teams.
TextFederated computing platform for training AI on healthcare data without moving it.
Multimodal