Perplexity AI Releases WANDR, Open Benchmark for Research Agents
Perplexity AI released WANDR (Wide ANd Deep Research), an open benchmark and evaluation harness designed to test research agents on realistic knowledge work tasks. The benchmark comprises 500 challenging data-collection tasks requiring agents to discover large sets of entities and investigate each with supporting evidence, totaling 170,495 source-backed records across lower, middle, and higher difficulty levels. WANDR grades agents by verifying each claim against cited evidence, checking whether pages are usable and excerpts truly support requirements.