Skip to content
Advertisement

Open

2,922 results

Audio

WearVox

WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables Paper: WearVox: An Egocentric Multichannel Voice Assistant Benchmark for…

1K–10K·CC-BY-NC·Audio (folder)
AudioMultimodalVideo

JavisInst-Omni

JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation HomePage Paper GitHub TL;DR We introduce JavisGPT, a…

10K–100K·Apache-2.0
Audio

GTSinger

GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks Yu Zhang , Changhao…

<1K·Audio (folder)
AudioMultimodalText

malaysian-youtube

Malaysian Youtube Malaysian and Singaporean youtube channels, total up to 60k audio files with total 18.7k hours. URLs…

10K–100K·Parquet
Audio

NOTSOFAR

Introduction Welcome to the "NOTSOFAR-1: Distant Meeting Transcription with a Single Device" Challenge. This repo contains the baseline…

10K–100K·Audio (folder)
Audio

AIR-Bench-Dataset

AIR-Bench Arxiv: is the AIR-Bench dataset download page.AIR-Bench encompasses two dimensions: foundation and chat benchmarks. The former consists…

<1K·CC-BY-NC·Audio (folder)
AudioMultimodalVideo

TAVGBench_1m

Installation Download this repo to a local folder, and unzip these .zip files under the TAVGBench 1m/data/. Then,…

1M–10M·MIT
Audio

VCB-Bench

VCB-Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents Introduction Voice Chat Bot Bench (VCB Bench)…

AudioMultimodalText

MRSAudio

MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations Humans rely on multisensory integration to perceive…

100K–1M·CC-BY·CSV
Advertisement