Agents’ Last Exam
Can AI agents complete the kind of long-horizon work that creates real economic value?
A living benchmark for professional workflows with verifiable outcomes. It spans 55 sub-industries and more than 1,500 tasks contributed by over 300 domain experts, evaluating whether an agent can deliver the work rather than merely answer a question about it.
Highlights: OpenAI’s GPT‑5.6 release; 50+ media coverage; 100K+ downloads in one month.
