Skip to content
Advertisement

Text Generation

217 results

ImageMultimodalText

llava-en-zh-300k

This dataset is composed by 150k examples of English Visual Instruction Data from LLaVA. 150k examples of English…

100K–1M·Apache-2.0·Parquet
Image

SandThink

SandThink Dataset (v1.0) SandThink 是一个专为具身智能 (Embodied AI) 任务设计的大规模指令微调与偏好对齐数据集。该数据集通过结构化的 Chain-of-Thought (CoT) 推理过程,显著提升了 Vision-Language-Action…

<1K·MIT·Images (folder)
ImageMultimodalText

ImgCode-8.6M

MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical Reasoning Repo: Paper: Introduction We introduce MathCoder-VL, a series…

1M–10M·Apache-2.0·Parquet
Image

C3B

English 简体中文 C³B: Comics Cross-Cultural Benchmark Culture In a Frame: C³B as a Comic-Based Benchmark for Multimodal Cultural…

1K–10K·CC-BY·Images (folder)
AudioMultimodalText

spoken-magpie-ja

Spoken-magpie LLMの日本語Instruction Tuning用データllm-jp/magpie-sft-v1.0をCosyVoice2 TTSを使用して音声化した商用利用可能な日本語の音声言語モデルのSFT用データセットです。 ある程度の話者多様性を持つように生成されています。…

100K–1M·Apache-2.0·Parquet
ImageMultimodalText

emova-alignment-7m

EMOVA-Alignment-7M 🤗 EMOVA-Models 🤗 EMOVA-Datasets 🤗 EMOVA-Demo 📄 Paper 🌐 Project-Page 💻 Github 💻 EMOVA-Speech-Tokenizer-Github Overview…

1M–10M·Apache-2.0·Parquet
Advertisement