Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
Preprocessed Common catalogue (CC-BY) DCAE
The images are resized and then encoded with the DC-AE f32 autoencoder. The resizing is done with a bucketmanager with base resolution 512×512, minimum side length 256, maximum side length 1024, all sides are divisible by 32 ofcourse as they needed to be encoded by the DCAEf32 encoder. The captions are generated with moondream2, encoded with siglip and bert. (Bert embeddings variance is very high, so use a norm layer). The text embeddings are padded to 64 tokens, but i have provided the… See the full description on the dataset page: commoncatalog-cc-by DCAE.
Source: Hugging Face Hub (SwayStar123/preprocessed_commoncatalog-cc-by_DCAE). Metadata imported from the dataset’s Hub tags.