Skip to content
Advertisement
MultimodalTabularText

embeddings-pre-training-curated

Embeddings pre-training curated data This dataset is the English subset of lightonai/embeddings-pre-training, assembled to reproduce the English data…

Embeddings pre-training curated data This dataset is the English subset of lightonai/embeddings-pre-training, assembled to reproduce the English data recipe described in the mGTE technical report (Zhang et al., 2024). The mGTE paper describes the data sources used to train the GTE family of multilingual text embedding and reranking models, but does not release the data itself. This dataset is our reconstruction of the English portion of that recipe, curated as part of a research… See the full description on the dataset page:

Source: Hugging Face Hub (lightonai/embeddings-pre-training-curated). Metadata imported from the dataset’s Hub tags.

Advertisement