Skip to content
Advertisement

Labeling & Annotation

This category covers organisations that supply human judgement and the software that organises it: annotation of images, video, LiDAR and sensor data, documents, speech, and text; reinforcement learning from human feedback and preference ranking; expert review in specialist domains such as medicine, law, and code; transcription; and commissioned data collection, where a workforce or contributor network is deployed to record speech, capture video, or teleoperate robots to a customer’s specification. The language-services and speech-data firms and the physical-AI collection companies belong here alongside the general annotation vendors, because their supply-chain role is the same: labour and process applied to data, delivered to order.The category splits along a line worth understanding before shortlisting. Some vendors sell software a customer staffs itself, some sell a managed workforce, and many sell both, with sharply different economics. Questions that separate them include who employs the annotators and under what pay and working conditions; how quality is measured and what is contractually guaranteed, since consensus scoring, gold-standard tasks, and blind rework are not equivalent; whether genuine domain experts are available for specialist work or only generalists; what turnaround looks like at volume rather than in a pilot; which annotation types, sensor fusions, and file formats are natively supported; whether data can remain inside the customer’s environment; and who owns the resulting labels and any models trained during the engagement.Vendors that produce labelled data by simulation rather than human effort are catalogued under synthetic data, and platforms that automate dataset selection without human labelling sit under data processing and curation.

78 results

Scale AI

Labeling & Annotation
Featured

Scale AI supplies human-labeled training data, RLHF, red-teaming and model evaluation to frontier AI labs, enterprises and government programs.

Multimodal

Andovar

Labeling & Annotation

Multilingual AI-training-data company providing text, audio, image, and video datasets built from studio-based data collection.

AudioImageTextVideo

Pangeanic

Labeling & Annotation

Valencia-based AI-data and language-technology company providing multilingual datasets, human feedback, and secure machine translation for enterprise AI.

MultimodalText

Welo Data

Labeling & Annotation

AI training-data division of Welocalize, a New York-founded localization company operating since 1997.

Text

Argos Multilingual

Labeling & Annotation

Kraków-founded localization company whose AI division supplies generative-AI and RLHF training data alongside custom machine translation.

Text

Blend

Labeling & Annotation

Tel Aviv-founded localization company (formerly One Hour Translation) that merged with data-annotation platform Tasq.ai to add AI training-data services.

AudioText

Lionbridge

Labeling & Annotation

Waltham, MA-founded language-services company whose AI Data Services division provides text, vision, and speech annotation for AI model training.

AudioImageText

Vistatec

Labeling & Annotation

Dublin-founded localization company whose Data & AI division provides data management, LLM data preparation, and human-centered language-AI services.

AudioText

Tarjama

Labeling & Annotation

MENA's Leading Enterprise Language Ecosystem

AudioText

Karya

Labeling & Annotation

Building AI for real-world complexity

AudioText

Aya Data

Labeling & Annotation

Your Trusted Global AI Data & Annotation Company

AudioText

Datamundi

Labeling & Annotation

We Create High Quality Human Data To Fuel Your AI

AudioText

TranscribeMe

Labeling & Annotation

Transcription company offering human-verified speech-to-text data and custom AI dataset creation for model training.

Audio

SelectStar

Labeling & Annotation

Korean data-construction platform providing large-scale data collection and annotation, including professional voice-actor speech recordings for AI training.

AudioText

AISHELL

Labeling & Annotation

Beijing speech-data company that publishes the open AISHELL Mandarin speech corpora and sells custom speech datasets.

Audio

Datatang

Labeling & Annotation

Beijing AI training-data provider with pre-built and custom speech-recognition and text-to-speech datasets, serving over 5,000 clients.

AudioText

Flitto

Labeling & Annotation

Multilingual Data for AI & Real-time AI Interpretation Solution Company

AudioText

XDOF

Labeling & Annotation

Berkeley robot-training-data infrastructure startup building teleoperation data pipelines and annotation systems for physical-AI labs; released the ABC-130K bimanual manipulation dataset…

Sensor / Time-seriesVideo
Advertisement