Skip to content
Research

New benchmark finds models still struggle with long-horizon planning

Researchers tested a dozen frontier models on tasks requiring 20+ sequential steps.

By Nakoda Newsroom

·5 min read

Prefer Nakoda AI News on Google

A new benchmark released this week measures how well AI models handle tasks that require planning more than twenty steps ahead.

Across a dozen frontier models tested, average success rates dropped sharply after the tenth step, with most models failing to recover from early mistakes.

The researchers say the results suggest current training methods reward short-term coherence over long-term plan correction.

Written by

Nakoda Newsroom

Independent journalism at the intersection of AI, business and society. Part of the Nakoda AI ecosystem.

Follow