Skip to content
Advertisement

Multimodal

2,368 results

AudioMultimodalText

IndicVoices

IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages Updates 23 December 2025 We now have…

1M–10M·CC-BY·Parquet
AudioMultimodalText

synthetic-asr-hi

Synthetic ASR data — hi Generated by Valsea-ASR/synthetic-data-pipeline. Audio is synthetic (TTS), targeted as training data for downstream…

10K–100K·JSON
AudioMultimodalVideo

JavisInst-Omni

JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation HomePage Paper GitHub TL;DR We introduce JavisGPT, a…

10K–100K·Apache-2.0
AudioMultimodalText

synthetic-asr-zh

Synthetic ASR data — zh Generated by Valsea-ASR/synthetic-data-pipeline. Audio is synthetic (TTS), targeted as training data for downstream…

10K–100K·JSON
AudioMultimodalText

malaysian-youtube

Malaysian Youtube Malaysian and Singaporean youtube channels, total up to 60k audio files with total 18.7k hours. URLs…

10K–100K·Parquet
AudioMultimodalVideo

TAVGBench_1m

Installation Download this repo to a local folder, and unzip these .zip files under the TAVGBench 1m/data/. Then,…

1M–10M·MIT
Advertisement