Skip to content
Advertisement
ImageMultimodalText

emova-alignment-7m

EMOVA-Alignment-7M 🤗 EMOVA-Models 🤗 EMOVA-Datasets 🤗 EMOVA-Demo 📄 Paper 🌐 Project-Page 💻 Github 💻 EMOVA-Speech-Tokenizer-Github Overview…

EMOVA-Alignment-7M 🤗 EMOVA-Models 🤗 EMOVA-Datasets 🤗 EMOVA-Demo 📄 Paper 🌐 Project-Page 💻 Github 💻 EMOVA-Speech-Tokenizer-Github Overview EMOVA-Alignment-7M is a comprehensive dataset curated for omni-modal pre-training, including vision-language and speech-language alignment. This dataset is created using open-sourced image-text pre-training datasets, OCR datasets, and 2,000 hours of ASR and TTS data. This dataset is part of the EMOVA-Datasets… See the full description on the dataset page:

Source: Hugging Face Hub (Emova-ollm/emova-alignment-7m). Metadata imported from the dataset’s Hub tags.

Advertisement