Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
The VCR-Wiki Dataset for Visual Caption Restoration (VCR) 🏠Paper 👩🏻‍💻 GitHub 🤗 Huggingface Datasets 📏 Evaluation with lmms-eval This is the official Hugging Face dataset for VCR-Wiki, a dataset for the Visual Caption Restoration (VCR) task. VCR is designed to measure vision-language models’ capability to accurately restore partially obscured texts using pixel-level hints within images. text-based processing becomes ineffective in VCR as accurate text restoration depends… See the full description on the dataset page:
Source: Hugging Face Hub (vcr-org/VCR-wiki-zh-hard). Metadata imported from the dataset’s Hub tags.