Skip to content
Advertisement
MultimodalTabularText

the-heap

The Heap Dataset We develop The Heap, a new contamination-free multilingual code dataset comprising 57 languages, which facilitates LLM evaluation…

The Heap Dataset We develop The Heap, a new contamination-free multilingual code dataset comprising 57 languages, which facilitates LLM evaluation reproducibility. The reproduction packge can be found here. Is your code in The Heap? If you would like to have your data removed from the dataset, follow the instructions on GitHub. Citation If you use this dataset as part of your research please cite us: @INPROCEEDINGS {11052803, author = { Katzy, Jonathan and… See the full description on the dataset page:

Source: Hugging Face Hub (AISE-TUDelft/the-heap). Metadata imported from the dataset’s Hub tags.

Advertisement