Agent Time-Horizon Explorer

An interactive projection of the METR “time horizon” trend — the length of task an AI agent can complete at 50% reliability, which has been doubling on a regular cadence. Set the current horizon and the doubling period, and the tool projects the calendar date when agents cross the one-hour, one-workday, one-week and beyond thresholds. Every assumption is a slider, so you can stress-test the optimistic and conservative cases yourself.

2.0hcurrent 50% horizon
×2 / 7modoubling cadence
1 hour already reached
1 workday (8h) +14.0 months
1 work-week (40h) +30.3 months
1 month (160h) +44.3 months

A structured extrapolation of the METR time-horizon trend, not a forecast — it assumes the exponential holds unbroken. And it projects the 50%-reliability line: the horizon at which an agent succeeds 99% of the time is far shorter, so a “1-week horizon” is not a week of unsupervised work. Move the sliders to compare the optimistic and conservative cases.

The most useful single number for tracking AI agents is not a benchmark score — it is a time horizon: the length of task, in human hours, that an agent can complete at 50% reliability. METR introduced this framing and observed something striking: across frontier models over several years, the horizon has grown roughly exponentially, doubling on a regular cadence. Exponential growth is a straight line on a logarithmic axis, which makes the trend easy to extend — and extending it is exactly what this tool does.

The math is a clean exponential. If the horizon doubles every T months, then reaching a target task length from today takes T × log₂(target ÷ current) months. Feed it a current horizon and a doubling period and it solves for the calendar date at which agents cross the thresholds that matter operationally: one hour, one workday, one week, one month. Because those inputs are debated — the doubling period especially, which some recent releases suggest is shrinking — every one of them is a slider, so you can run the optimistic and the conservative case side by side.

The critical caveat is baked into the framing: this projects the 50%-reliability line, and 50% is nowhere near production. The horizon at which an agent succeeds 99% of the time is far shorter, and the gap widens as stakes rise. So a projected “one-week horizon” does not mean a week of unsupervised autonomous work — it means the agent finishes week-length tasks half the time. Use the tool to compare scenarios and pressure-test assumptions, not to pin a date on the calendar; no exponential runs forever, and the honest value here is seeing how much the answer moves when you nudge the inputs.

How it works

  • Time horizon = task length an agent finishes at 50% reliability.
  • The horizon has been doubling on a regular cadence (METR).
  • Projects calendar dates for 1h / 1 workday / 1 week thresholds.
  • Every assumption — current horizon, doubling period — is a slider.

Frequently asked questions

What is a “time horizon”?

It is the length of task — measured by how long it takes a skilled human — that an AI agent can complete at a given reliability, usually 50%. METR popularized measuring capability this way: instead of a benchmark score, you ask “how long a task can the model finish half the time?” A model with a two-hour horizon reliably handles work a human would take up to two hours on, and struggles beyond that. It converts fuzzy “how good is it” into a single, trendable number.

Where does the doubling come from?

METR found this horizon has been growing exponentially — roughly doubling on a fixed cadence over several years of frontier models. Exponential growth plots as a straight line on a log axis, so projecting it forward is just extending that line. The tool defaults to a representative doubling period but lets you change it, because the exact figure is debated and has arguably been accelerating on recent releases.

How seriously should I take the projected dates?

As a structured extrapolation, not a prophecy. It assumes the past exponential continues unbroken, which no trend does forever — data limits, reliability walls at higher stakes, and the gap between 50% and the 99% you need in production all bend the curve. Treat the dates as “if the current pace held, this is when,” and use the sliders to see how sensitive the answer is to the doubling period. The honest use is comparing scenarios, not betting on one.

Why 50% reliability and not higher?

Because that is where the trend is measured, and it matters that it is only half. A 50% horizon means the agent finishes that task length half the time — nowhere near production-grade for high-stakes work. The horizon at 80% or 99% reliability is much shorter, so a “one-week horizon at 50%” does not mean you can hand an agent a week of unsupervised work. The tool projects the 50% line because that is the published trend; read every date with that asterisk.