Skip to content
Advertisement

Datagen was a synthetic data company that generated simulated visual data for computer vision. Using a data-as-code approach, it produced photorealistic, human-centric 3D imagery, complete with automatic labels, so teams could train and test perception models without collecting and annotating real-world photographs. The company ceased operations in 2024 and is included here for reference within the AI data supply chain.

What they provided

Datagen offered a self-serve platform and API for designing and generating labeled synthetic image datasets with granular control over people, faces, objects, and indoor environments. Because scenes were simulated, datasets came with precise ground-truth annotations and could be balanced across attributes and edge cases that are hard to capture in real footage.

  • Segment: Synthetic data for computer vision (historical)
  • Modalities: Photorealistic 3D images of humans, faces, and indoor scenes
  • Delivery: Self-serve platform and API, with automatic labeling
  • Common uses: Face-related tasks, in-cabin automotive, AR/VR, and security

Where it fit in the AI data supply chain

Datagen sat at the synthetic-data layer for the visual modality, positioned as an alternative to real image collection and manual labeling for human-centric perception. It raised significant venture funding, including a Series B led by Scale Venture Partners, before winding down amid market competition and a difficult pivot. Its trajectory is a useful reference point for how single-modality synthetic-imagery businesses fared as generative AI reshaped the data landscape.

3D / Point CloudImage
Advertisement