Skip to content
Advertisement
Text

mmarco-contrastive

mMARCO-contrastive The dataset is a modification of mMARCO focusing on French and English parts. The aim is to train a bi-encoder model using all…

mMARCO-contrastive The dataset is a modification of mMARCO focusing on French and English parts. The aim is to train a bi-encoder model using all hard negatives from the database. Instead of having a query/positive/negative triplet, we pair all negatives with a query and a positive. However, it’s worth noting that there are many false negatives in the dataset. This isn’t a big issue with a triplet view because false negatives are much fewer in number, but it’s more significant with… See the full description on the dataset page:

Source: Hugging Face Hub (Cyrile/mmarco-contrastive). Metadata imported from the dataset’s Hub tags.

Advertisement