Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
DRBench: A Realistic Benchmark for Enterprise Deep Research 📄 Paper 💻 GitHub 💬 Discord DRBench is the first of its kind benchmark designed to evaluate deep research agents on complex, open-ended enterprise deep research tasks. It tests an agent’s ability to conduct multi-hop, insight-driven research across public and private data sources, just like a real enterprise analyst. ✨ Key Features 🔎 Real Deep Research Tasks: Not simple fact lookups. Tasks… See the full description on the dataset page:
Source: Hugging Face Hub (ServiceNow/drbench). Metadata imported from the dataset’s Hub tags.