Skip to content
Advertisement
ImageMultimodalText

Cauldron-JA

Dataset Card for The Cauldron-JA Dataset description The Cauldron-JA is a Vision Language Model dataset that translates 'The Cauldron' into Japanese…

Dataset Card for The Cauldron-JA Dataset description The Cauldron-JA is a Vision Language Model dataset that translates ‘The Cauldron’ into Japanese using the DeepL API. The Cauldron is a massive collection of 50 vision-language datasets (training sets only) that were used for the fine-tuning of the vision-language model Idefics2. To create a Japanese Vision Language Dataset, datasets related to OCR, coding, and graphs were excluded because translating them into Japanese… See the full description on the dataset page:

Source: Hugging Face Hub (turing-motors/Cauldron-JA). Metadata imported from the dataset’s Hub tags.

Advertisement