Skip to content
Advertisement

TaiwanVQA: Benchmarking and Enhancing Cultural Understanding in Vision-Language Models Dataset Summary TaiwanVQA is a visual question answering (VQA) benchmark designed to evaluate the capability of vision-language models (VLMs) in recognizing and reasoning about culturally specific content related to Taiwan. This dataset contains 2,736 images captured by our team, paired with 5,472 manually designed questions that cover diverse topics from daily life in Taiwan… See the full description on the dataset page:

Source: Hugging Face Hub (hhhuang/TaiwanVQA). Metadata imported from the dataset’s Hub tags.

Advertisement