Skip to content
Advertisement

Dataset Description, Collection, and Source MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark based on scientific talks that is designed to evaluate instruction-following in crosslingual, multimodal settings over both short- and long-form inputs. MCIF spans three core modalities — speech, vision, and text — and four diverse languages (English, German, Italian, and Chinese), enabling a comprehensive evaluation of MLLMs’… See the full description on the dataset page:

Source: Hugging Face Hub (Rendra86318/MCIF). Metadata imported from the dataset’s Hub tags.

Advertisement