What an Eight-Hour AI Score Leaves Out
A task-completion time horizon describes a model's performance against tasks timed for human experts. Turning that score into a promise of a working day saved requires another kind of evidence: what gets accepted, whose attention it consumes, and when the work becomes useful.
The hypothetical working day
Imagine a manager reading that an AI system has an eight-hour task-completion horizon. This is a hypothetical score, not a claim about any current model. The tempting inference is straightforward: hand over tomorrow's assignment and free a working day. But the assignment still has to become something the team can use. Someone may need to explain its boundaries, supply missing context, inspect the result, and settle questions the original request left open.
Our starting question is therefore practical: when does delegated output become accepted work? A patch awaiting review and a patch ready to maintain occupy different places in a team's schedule. A horizon can help decide which work to try delegating. The inference from that score to a day of labor saved needs its own accounting.
Four different measures
METR's task-completion horizon identifies the human task duration at which a fitted curve predicts a specified success rate, commonly 50%. It does not measure how long the agent runs unattended. The suite largely covers self-contained software work. The dashboard, marked updated May 8, 2026 when consulted October 2, warns that measurements beyond 16 hours are unreliable on this suite.
Thomas Kwa's January 2026 limitations note emphasizes uncertainty and differences across domains: a 50% horizon does not establish dependable delegation, while estimating 99% reliability would require much richer, cleaner evidence. It is a less-reviewed research note, not necessarily METR's institutional view. The dashboard also explains that an aggregate success rate mixes easier and harder tasks; it does not assign the same chance of success to every individual assignment.
Our explanatory framework separates the quantities a manager might otherwise collapse. The table is a way to organize questions, not a measured dataset.
| Measure | Question answered |
|---|---|
| Capability horizon | What human task duration corresponds to a specified benchmark success probability? |
| Elapsed machine time | How long did the agent's run take? |
| Human attention | How much active work did people spend directing, checking, or repairing it? |
| Accepted output | What usable result met the team's agreed requirements? |
A score needs a benchmark version
Time Horizon 1.1, released January 29, 2026, expanded the suite from 170 to 228 tasks and its eight-hour-plus tasks from 14 to 31. Only five of those 31 long tasks had measured human baselines; the others used estimates. The evaluation infrastructure also moved from Vivaria to Inspect. Changing the task mix changes what is estimated, and uncertainty intervals remain wide.
Our interpretation is that revision can improve a benchmark while complicating comparison. A longer horizon can be meaningful evidence of progress. To use that evidence responsibly, keep the suite version and uncertainty alongside the point estimate; a staffing forecast still needs workplace measurements.
Passing checks and finishing a contribution
An August 13, 2025 METR evaluation tested 18 tasks from two large repositories with Claude 3.7 Sonnet in a basic Inspect ReAct scaffold. Automated success was 38%, with a 95% confidence interval of plus or minus 19 percentage points. Of 15 pull requests manually reviewed, none was mergeable as submitted. This small, older experiment used limited elicitation; it establishes neither present-day performance nor an upper bound for better agents.
Consider a hypothetical patch that satisfies its tests but introduces an awkward dependency. The maintainer must decide whether to accept the dependency, redesign the patch, or ask for another attempt. The example locates work outside the automated checks: someone must reconcile a proposed solution with a codebase's future. A benchmark can test a specific requirement well while leaving that decision to another person.
Acceptance also depends on the assignment. An exploratory prototype can earn its keep despite rough edges. A contribution intended for long-term maintenance faces different obligations. The right boundary follows the intended use of the output.
Productivity evidence over time
In a randomized study published July 10, 2025, 16 experienced developers worked on 246 issues in familiar open-source repositories. Tasks assigned to allow early-2025 AI took 19% longer. That is evidence about a narrow setting with the tools then available. It is not a universal effect of AI assistance or a productivity estimate for October 2026.
METR's February 24, 2026 update explained why its later experiment gave an unreliable signal. Developers who disliked the prospect of working without AI, and tasks they especially wanted AI for, were missing from the sample; lower pay also contributed to selection, and concurrent agent use complicated time accounting. Raw estimates suggested speedups with intervals crossing zero. The authors considered larger gains likely, but the experiment did not establish their size or prove a reversal.
The May 11, 2026 survey provides positive evidence of perceived benefit. Among 349 technical workers surveyed from February through April, median estimates across three value measures ranged from 1.4 to 2 times; median self-reported speed was 3 times. These were a convenience sample's reports, subject to selection and counterfactual judgments, rather than measured causal gains. Reported value also concerns a potentially changed mix of work.
Our reading keeps the chronology intact. An earlier measured slowdown does not cancel later experiences of substantial benefit. Enthusiastic reports also cannot repair an experiment's missing comparison. Each informs a different question. A team adopting agents today should ask what improved in its own workflow, including whether it now undertakes worthwhile work it previously postponed. Faster completion is only one possible benefit.
Three ledgers for a local trial
The following is our proposed method, not a finding established by these studies. Keep separate records of accepted outcomes, total active human work, and elapsed turnaround. Agree on the acceptance boundary before starting. Otherwise, the trial can appear successful simply by quietly reducing the quality required to call something finished.
The outcome ledger records what was accepted and why. Include rejected attempts, abandoned assignments, and tasks routed back to people. A generated artifact is an intermediate event. Acceptance might mean a maintainer approves a change, a report supports the decision it was commissioned for, or a prototype answers the intended question. The criterion should fit the work rather than reward whatever the agent happens to produce easily.
The human-work ledger includes prompting, supplying context, inspecting results, repairing mistakes, and reviewing someone else's AI-assisted contribution. Record whose effort it was. Counting only the initiating developer's time can turn a transfer to maintainers into an apparent saving. Include work spent on failures too: the rejected patch may still have consumed a careful review.
The turnaround ledger tracks when a request began and when an accepted result became available. Waiting matters for deadlines, but it differs from active attention. An agent can run while its user works elsewhere; a completed artifact can wait in a review queue. Those intervals affect delivery without necessarily consuming human work continuously. Preserve both records instead of forcing them into a single duration.
These records can point in different directions without contradicting one another. A workflow might reduce active effort while increasing turnaround because review happens later. Another might deliver sooner while demanding more attention from a scarce specialist. The team's objective determines which trade is attractive. For an urgent request, delivery time may dominate; for a recurring backlog, sustainable review capacity may matter more.
Concurrent work needs particular care. If someone supervises several runs, record their actual attention without allocating the entire overlapping wall-clock interval to every run. If several people review one result, count each person's active contribution. This prevents both inflated costs from duplicated waiting and understated costs from invisible collaboration.
For a useful comparison, retain enough detail about task type and difficulty to see whether the assisted workflow received easier assignments. Note changes in tools and review rules. A local trial need not answer every research question, but it should reveal when an improvement depends on a different task mix or a more permissive definition of done.
These ledgers also change who receives credit. The person who produces a large volume of drafts may be visibly productive while the colleague absorbing review becomes a bottleneck. Shared accounting can recognize both contributions. It can also show a legitimate trade: additional review may be worthwhile if the team gains valuable output. The point is to make that trade visible to everyone doing the work.
A claim someone can use
A responsible delegation claim should describe its operating conditions. State the task class, model and tool setup, success and acceptance criteria, active human effort, failed attempts, review responsibilities, and update date. This is our recommendation for making a claim useful, rather than a reporting standard proven by the cited research.
The resulting statement might say that a team accepts agent-prepared drafts after a named kind of review, with a recorded amount of human work. It should let a reader understand where completion was judged and who still had to decide. A capability number alone cannot reveal those arrangements.
Review is part of the design of delegation. Where it is cheap and decisive, substantial assistance can be valuable even when some attempts fail. Where deciding correctness requires reconstructing the work, an impressive artifact may create a difficult inspection burden. The local question is whether the workflow delivers enough accepted value for the attention it asks of people.
Return to the hypothetical eight-hour score. It can justify trying a more ambitious assignment. The next claim needs evidence from the work itself. Keep the capability measurement, then add the account of acceptance and effort. That gives progress a practical meaning: more useful work, with the people responsible for it able to see what they gained and what they still carry.
Sources
- METR, Task-Completion Time Horizons of Frontier AI Models. Dashboard updated May 8, 2026; consulted October 2, 2026.
- Thomas Kwa, Clarifying limitations of time horizon. Research note, January 22, 2026.
- METR, Time Horizon 1.1. January 29, 2026.
- METR, Algorithmic vs. Holistic Evaluation. Page dated August 13, 2025.
- METR, Early-2025 AI and experienced open-source developer productivity. July 10, 2025.
- METR, Changing the developer productivity experiment design. February 24, 2026.
- METR, Self-reported impact of early-2026 AI on technical worker productivity. May 11, 2026.
Related reading
- The Reliability Scorecard Becomes the Agent Gate
- The Coding Agent Becomes the Maintainer
- The Evaluation Archive Becomes the Frontier Claim
Production: commissioned by the site operator; drafted by GPT-6.1 Sol; research, source checking, and editorial review by the coordinating AI assistant.