AI & LLMs

Time Horizon (50%)

also: 50% time horizon · task time horizon

The length of task an AI agent can complete at 50% reliability — a single trendable measure of capability.

The time horizon measures capability by task length: how long a task, in human time, the agent finishes at a given success rate — usually 50%. Popularized by METR, it has grown roughly exponentially across frontier models. Because 50% is far from production-grade, a long horizon signals reach, not trustworthy autonomy.

Worked example: a way to measure agent capability by the LENGTH of task it can complete reliably — roughly how long the task would take a human — capturing that agents fail not on any single step but on maintaining coherence over many steps. Gotcha: it reframes progress from ‘can it do this benchmark’ to ‘how long a task can it sustain’, which tracks real usefulness better than one-shot accuracy — but ‘reliably’ hides a success-rate threshold, and horizon depends heavily on the harness (tools, memory, verification), not the model alone.