Skip to content
Advertisement
AudioMultimodalVideo

JavisInst-Omni

JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation HomePage Paper GitHub TL;DR We introduce JavisGPT, a multimodal…

JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation HomePage Paper GitHub TL;DR We introduce JavisGPT, a multimodal LLM that can understand audiovisual inputs and simultaneously generate synchronized sounding videos in a unified model. We also curate the JavisInst-Omni dataset to facilitate instruction-tuning for comprehension and generation on sounding videos. 📰 News 2025.12.30 🚀 We release the training… See the full description on the dataset page:

Source: Hugging Face Hub (JavisVerse/JavisInst-Omni). Metadata imported from the dataset’s Hub tags.

Advertisement