Scale AI supplies human-labeled training data, RLHF, red-teaming and model evaluation to frontier AI labs, enterprises and government programs.
MultimodalLabeling & Annotation
This category covers organisations that supply human judgement and the software that organises it: annotation of images, video, LiDAR and sensor data, documents, speech, and text; reinforcement learning from human feedback and preference ranking; expert review in specialist domains such as medicine, law, and code; transcription; and commissioned data collection, where a workforce or contributor network is deployed to record speech, capture video, or teleoperate robots to a customer’s specification. The language-services and speech-data firms and the physical-AI collection companies belong here alongside the general annotation vendors, because their supply-chain role is the same: labour and process applied to data, delivered to order.The category splits along a line worth understanding before shortlisting. Some vendors sell software a customer staffs itself, some sell a managed workforce, and many sell both, with sharply different economics. Questions that separate them include who employs the annotators and under what pay and working conditions; how quality is measured and what is contractually guaranteed, since consensus scoring, gold-standard tasks, and blind rework are not equivalent; whether genuine domain experts are available for specialist work or only generalists; what turnaround looks like at volume rather than in a pilot; which annotation types, sensor fusions, and file formats are natively supported; whether data can remain inside the customer’s environment; and who owns the resulting labels and any models trained during the engagement.Vendors that produce labelled data by simulation rather than human effort are catalogued under synthetic data, and platforms that automate dataset selection without human labelling sit under data processing and curation.
78 results
Pangeanic
Labeling & AnnotationValencia-based AI-data and language-technology company providing multilingual datasets, human feedback, and secure machine translation for enterprise AI.
MultimodalTextArgos Multilingual
Labeling & AnnotationKraków-founded localization company whose AI division supplies generative-AI and RLHF training data alongside custom machine translation.
TextLionbridge
Labeling & AnnotationWaltham, MA-founded language-services company whose AI Data Services division provides text, vision, and speech annotation for AI model training.
AudioImageTextGlobose Technology Solutions
Labeling & AnnotationRefine Your AI Vision: Premium Data Collection with a Human Touch
AudioTextTranscribeMe
Labeling & AnnotationTranscription company offering human-verified speech-to-text data and custom AI dataset creation for model training.
AudioSelectStar
Labeling & AnnotationKorean data-construction platform providing large-scale data collection and annotation, including professional voice-actor speech recordings for AI training.
AudioTextDatabaker Technology
Labeling & AnnotationIntelligent speech interaction and AI data services expert
AudioTextXDOF
Labeling & AnnotationBerkeley robot-training-data infrastructure startup building teleoperation data pipelines and annotation systems for physical-AI labs; released the ABC-130K bimanual manipulation dataset…
Sensor / Time-seriesVideo