Skip to content
Advertisement
3D / Point CloudAudioImageMultimodalText

Emergence-Text-Image-Audio-3D

Emergence: The Four Forms of Intelligence Summary A multimodal dataset that unifies Text, Image, Audio, and 3D modalities with quad-modality…

Emergence: The Four Forms of Intelligence Summary A multimodal dataset that unifies Text, Image, Audio, and 3D modalities with quad-modality alignment for every sample, ensuring that each record contains semantically consistent representations of the same concept. This dataset is curated by using 3D assets from Objaverse as anchors and aligning them with semantically corresponding images and audio clips from various sources using a embedding search… See the full description on the dataset page:

Source: Hugging Face Hub (VINAY-UMRETHE/Emergence-Text-Image-Audio-3D). Metadata imported from the dataset’s Hub tags.

Advertisement