Skip to content
Advertisement
MultimodalTabularText

TheStack-Rust

TheStack-Rust

Dataset 1: TheStack – Rust – Cleaned Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Rust, a popular statically typed language. Target Language: Rust Dataset Size: Training: 900,000 files Validation: 50,000 files Test: 50,000 files Preprocessing: Selected Rust as the target language due to its… See the full description on the dataset page:

Source: Hugging Face Hub (ammarnasr/the-stack-rust-clean). Metadata imported from the dataset’s Hub tags.

Advertisement