New benchmark finds models still struggle with long-horizon planning
Researchers tested a dozen frontier models on tasks requiring 20+ sequential steps.
··5 min read
Papers, benchmarks and findings worth knowing.
Researchers tested a dozen frontier models on tasks requiring 20+ sequential steps.
··5 min read
The Nakoda AI Brief
One email. Every Friday. No noise.