Skip to content
Advertisement
AudioMultimodalText

avspeech-visual-audio

AVSpeech Video + Audio A restructured subset of the AVSpeech dataset with separated media streams and derived identifiers. clip id: unique identifier…

AVSpeech Video + Audio A restructured subset of the AVSpeech dataset with separated media streams and derived identifiers. clip id: unique identifier per clip, derived as {youtube id} {start:.3f} {end:.3f} from the original AVSpeech CSV columns. avspeech metadata: JSON string with the original AVSpeech row fields (youtube id, start sec, end sec, x center, y center). video: video-only stream (no audio), stream-copied without re-encoding. audio: audio-only stream, stream-copied… See the full description on the dataset page:

Source: Hugging Face Hub (ProgramComputer/avspeech-visual-audio). Metadata imported from the dataset’s Hub tags.

Advertisement