Skip to content
Advertisement
MultimodalTabularText

fineweb-2-edu-japanese

🍷 FineWeb2 Edu Japanese: High-Quality Educational Japanese Dataset This dataset consists of 120 million texts (approximately 89.3B tokens) filtered…

🍷 FineWeb2 Edu Japanese: High-Quality Educational Japanese Dataset This dataset consists of 120 million texts (approximately 89.3B tokens) filtered from the 376 million Japanese texts in FineWeb2 that were deemed educational. The following subsets are also provided: default: Approximately 120M texts (120 million texts) totaling around 89.3B tokens sample 10BT: A random sample of about 10B tokens from the default dataset small tokens: Data composed solely of texts with 512 tokens… See the full description on the dataset page:

Source: Hugging Face Hub (hotchpotch/fineweb-2-edu-japanese). Metadata imported from the dataset’s Hub tags.

Advertisement