Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
sTinyStories A spoken version of TinyStories Synthesized with LJ voice using FastSpeech2. The dataset was synthesized to boost the training of Speech Language Models as detailed in the paper “Slamming: Training a Speech Language Model on One GPU in a Day”. It was first suggested by Cuervo et. al 2024. We refer you to the SlamKit codebase to see how you can train a SpeechLM with this dataset. Usage from datasets importload dataset dataset =… See the full description on the dataset page:
Source: Hugging Face Hub (slprl/sTinyStories). Metadata imported from the dataset’s Hub tags.