Skip to content
Advertisement
MultimodalTabularText

ResearchClawBench

ResearchClawBench           Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery Quick Start…

ResearchClawBench           Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery Quick Start Submit Tasks How It Works Domains Leaderboard Add Your Agent ResearchClawBench is a benchmark that measures whether AI coding agents can independently conduct scientific research — from reading raw data to producing publication-quality reports — and then rigorously evaluates the results against real human-authored papers.… See the full description on the dataset page:

Source: Hugging Face Hub (InternScience/ResearchClawBench). Metadata imported from the dataset’s Hub tags.

Advertisement