Skip to content
Advertisement

Dataset Card for NatureBench NatureBench is a cross-discipline benchmark of 90 tasks distilled from peer-reviewed Nature-family publications, spanning 6 scientific domains. It is designed to evaluate whether AI coding agents can move beyond reproduction toward discovery: each task asks an agent to solve a real scientific machine-learning problem and is scored against the source paper’s reported state of the art. 📄 arXiv paper: 💻 GitHub code… See the full description on the dataset page:

Source: Hugging Face Hub (FrontisAI/NatureBench). Metadata imported from the dataset’s Hub tags.

Advertisement