Skip to content
Advertisement
Text

NorwegianCourtsBitextMining

NorwegianCourtsBitextMining An MTEB dataset Massive Text Embedding Benchmark Nynorsk and Bokmål parallel corpus from Norwegian courts. Norwegian…

NorwegianCourtsBitextMining An MTEB dataset Massive Text Embedding Benchmark Nynorsk and Bokmål parallel corpus from Norwegian courts. Norwegian courts have two standardised written languages. Bokmål is a variant closer to Danish, while Nynorsk was created to resemble regional dialects of Norwegian. Task category t2t Domains Legal, Written Reference How to evaluate on this task You can evaluate an embedding model on this… See the full description on the dataset page:

Source: Hugging Face Hub (mteb/NorwegianCourtsBitextMining). Metadata imported from the dataset’s Hub tags.

Advertisement