Skip to content
Advertisement
Text

WideSearch

WideSearch: Benchmarking Agentic Broad Info-Seeking Dataset Summary WideSearch is a benchmark designed to evaluate the capabilities of Large Language…

WideSearch: Benchmarking Agentic Broad Info-Seeking Dataset Summary WideSearch is a benchmark designed to evaluate the capabilities of Large Language Model (LLM) driven agents in broad information-seeking tasks. Unlike existing benchmarks that focus on finding a single, hard-to-find fact, WideSearch assesses an agent’s ability to handle tasks that require gathering a large amount of scattered, yet easy-to-find, information. The challenge in these tasks lies not in… See the full description on the dataset page:

Source: Hugging Face Hub (ByteDance-Seed/WideSearch). Metadata imported from the dataset’s Hub tags.

Advertisement