Skip to content
Advertisement
ImageMultimodalText

DenseFusion-1M for comprehensive image descriptions

DenseFusion-1M for comprehensive image descriptions

Paper GitHub Introduction An image is worth a thousand words”. Comprehensive image descriptions are essential for multi-modal perception, while images contains various visual elements of different granularities that are challenging to harness. We propose Perceptural Fusion to integrate the diverse visual perception experts for capturing visual elements and adopt a MLLM as a centric pivot for… See the full description on the dataset page:

Source: Hugging Face Hub (BAAI/DenseFusion-1M). Metadata imported from the dataset’s Hub tags.

Advertisement