Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
Bhasha SFT Bhasha SFT is a massive collection of multiple open sourced Supervised Fine-Tuning datasets for training Multilingual Large Language…
Bhasha SFT Bhasha SFT is a massive collection of multiple open sourced Supervised Fine-Tuning datasets for training Multilingual Large Language Models. The dataset contains collation of over 13 million instances of instruction-response data for 3 Indian languages (Hindi, Gujarati, Bengali) and English having both human annotated and synthetic data. Curated by: Soket AI Labs Language(s) (NLP): English, Hindi, Bengali, Gujarati License: cc-by-4.0, apache-2.0, mit … See the full description on the dataset page:
Source: Hugging Face Hub (soketlabs/bhasha-sft). Metadata imported from the dataset’s Hub tags.