DoneThat
Back to blog

You Can't Measure Outcomes. Measure Time on Goals Instead.

Everyone says measure outcomes. For individual knowledge work that mostly doesn't work — outcomes lag, they can't be attributed, and luck dominates. Here is the measurable alternative: time on goals, paired with energy.

Cover graphic contrasting outcomes, which are outside your control, with time on goals, which is inside it.

Every productivity framework tells you to measure outcomes rather than effort. It is good advice that almost nobody can follow, because for most individual knowledge work the outcome arrives months late, depends on a dozen other people, and is heavily contaminated by luck.

The evidence for this is not a hunch. It is the reason your employment contract pays you for time. Economics has a formal result here: when a job has some measurable tasks and some unmeasurable ones, paying for the measurable part actively distorts effort away from the rest — and a fixed wage can be the optimal contract (Holmström & Milgrom, Journal of Law, Economics & Organization, 1991). We do not pay by the hour because we are lazy about measurement. We pay by the hour because paying for measured output would make the work worse.

So what is left? The one input you actually control: time on goals. That is what I am going to talk you into measuring, and I will show you the evidence rather than asking you to take my word for it.

Why Outcome Measurement Fails at the Individual Level

Three things break it, and they compound.

Lag. The outcome of good strategy work shows up in quarters, sometimes years. A metric that resolves after the decisions it should inform is a history lesson, not a feedback loop.

Attribution. Ask who is responsible for a feature that shipped and worked. The engineer who built it? The designer who specced it? The PM who cut the scope so it was buildable at all? The platform team whose pipeline made a same-day deploy possible? The support lead whose ticket data pointed at the problem in the first place? The result was produced jointly and it does not decompose into shares. The 2025 DORA report puts a number on the same phenomenon from the other direction: AI adoption raised delivery throughput and instability simultaneously, behaving as an amplifier of whatever an organization already was (DORA, September 2025, ~5,000 respondents). When the identical input makes cohesive teams better and fragmented ones worse, the thing you are measuring is the organization.

Luck. In the fiscal year ending 31 January 2020, Zoom booked $622.7 million in revenue. In the year ending 31 January 2021, it booked $2.65 billion — growth of 326% (Zoom Communications filings, SEC EDGAR, CIK 1585521). The product was good. It had been good the year before too. What changed was a pandemic. No performance review written at Zoom that year could separate one person's contribution from COVID's, and every outcome metric at every scale carries a smaller version of that same problem.

The honest version of the claim is narrow and still damning: outcomes are measurable in aggregate and over long horizons. They are rarely attributable to one person over a short one — which is exactly the resolution at which people try to use them.

Economics Settled This Decades Ago

The multitask principal-agent result is the sharpest statement of the problem. When a principal can measure some of an agent's tasks but not others, tying pay to the measurable ones pulls effort away from the invisible ones. The canonical example is paying teachers on test scores and getting teaching to the test. Holmström and Milgrom's conclusion is not "measure better" — it is that a flat wage can beat a performance contract outright (JLEO, 1991).

Two decades of follow-up work did not resolve it. Canice Prendergast's survey of the field concluded that while people clearly respond to pay-for-performance, there is little evidence contracts are designed the way theory predicts, and that we still know little about how to provide incentives to workers whose output is hard to measure (Journal of Economic Literature, 1999).

The counter-case proves the rule. Safelite Glass moved windshield installers from hourly pay to piece rates and output per worker rose 44%, with roughly half the gain from higher individual effort and half from attracting and retaining faster installers (Lazear, AER, 2000). Outcome pay works beautifully — when the output is countable, attributable to one person, and finished the same day. Windshields qualify. Almost no knowledge work does.

What a Sales Quota Actually Measures

Sales is the one white-collar role built end to end for outcome measurement. The unit is a closed deal, attribution is deliberately engineered, and pay is tied straight to it. So it is the fairest possible test of outcome pay — and it is worth looking at what a quota actually tracks.

Take Zoom, whose revenue is a matter of public record across nine fiscal years.

Revenue growth of 326% in the pandemic year, and 3% four years later — with broadly the same product and the same sales organisation.

Put yourself in a Zoom account executive's shoes. In the year ending January 2021 you were carrying a bag while revenue grew 326%. Three years later, growth was 3% (SEC EDGAR, CIK 1585521). You did not become a worse salesperson by a factor of a hundred. The product did not get worse. The world moved, and your commission moved with it.

That is what a quota measures: mostly the year, partly the territory, and somewhere in the mix, you. Companies keep paying commission anyway, and reasonably so — it is a way to share risk and focus attention, not a precise readout of contribution. But if outcome pay is this noisy in the role purpose-built for it, the case for scoring an engineer, a designer or a researcher on outcomes is weaker still.

The Stoic Move

Epictetus opens the Enchiridion by dividing the world in two: some things are up to us, and some are not. Our judgements and our efforts are up to us. Reputation, results, and the actions of other people are not. The whole discipline is refusing to stake your peace on the second category.

Outcome metrics stake everything on the second category. They ask you to be accountable for a number that market conditions, colleagues, and timing can overwrite at will.

Attention is in the first category. What you decided mattered this week, and whether your hours actually went there, is up to you — and unlike the outcome, you can check it on Friday.

The Unit Is Time on Goals, Not Focus and Not Hours

This is the definition the whole argument rests on, so it is worth being precise. Time on goals is the share of your working time that went to the priorities you set in advance. Three things it is not:

Not thisWhy it fails
Hours workedAn input with no direction. Sixty unfocused hours on the wrong thing scores higher than twenty on the right one.
Focus timeBetter, but still directionless. Deep, uninterrupted work on a task that does not matter is a more expensive mistake, not a cheaper one.
ActivityMessages, commits, tickets, time at a desk. Cheap to produce, trivially performed, and correlated with visibility rather than contribution.

The obvious objection is that this metric is gameable — you could point all your time at goals that do not matter. That is true, and it is the honest limit of the whole approach: measurement cannot set your direction. Someone has to decide what the goals are, and that is strategy and judgement, not analytics. What measurement can tell you is whether the last four weeks of your life actually went where you said they should. In practice that gap is enormous, and almost nobody has the data to see it.

Why This Input Needs Defending

Time on goals is not merely measurable. It is scarce, it is under attack, and protecting it is the highest-leverage thing most people can do.

Some of it is eaten before you ever get to choose. Atlassian's State of Teams 2025, covering 12,000 knowledge workers and 200 executives, found teams lose 25% of the week simply searching for answers (Atlassian, 2025 — vendor research from a company selling the remedy). A quarter of the week, gone to friction rather than to anyone's priorities.

The rest is eaten by interruption. Microsoft's June 2025 Work Trend Index report, built on anonymized Microsoft 365 telemetry alongside Work Trend Index survey data, found employees interrupted every two minutes during working hours — about 275 times a day. Half of all meetings landed inside the peak-focus windows of 9–11am and 1–3pm, and 57% were ad hoc calls with no calendar invite (Microsoft WorkLab, 17 June 2025 — vendor research, and the telemetry covers Microsoft 365 customers rather than a representative sample). Nearly half of employees and more than half of leaders described their work as chaotic and fragmented.

There is also a motivational reason, and it is the strongest single finding in the workplace literature. Teresa Amabile and Steven Kramer analyzed roughly 12,000 daily diary entries from 238 professionals across 26 project teams in seven companies. The largest driver of a good day at work was making progress on meaningful work — the progress principle (Amabile & Kramer, "The Power of Small Wins", Harvard Business Review, May 2011). Time on goals is the input that produces the progress. Measuring it and protecting it are the same act.

You Cannot Feel Your Own Output

One more reason to measure the input from records: self-assessment fails, and it fails even after the evidence arrives.

METR ran a randomized controlled trial with 16 experienced open-source developers across 246 real tasks, randomly permitting or prohibiting AI tooling. Participants forecast a 24% speedup. Measured, they were 19% slower. Afterward, having lived through the slowdown, they still estimated they had been 20% faster (METR, 10 July 2025).

The post-hoc estimate landed almost exactly where the forecast had — on the wrong side of zero.

The sample is small and the population specific, so treat the magnitude as indicative. The structure is what transfers: the gap between felt and measured performance survived the experience that should have closed it. Hours are no better remembered — employed respondents overestimate their working time by 5–10% against their own diaries (Monthly Labor Review, June 2011).

The same gap shows up in our own tracking. Across that year of data, the median workday spanned 9.4 hours from first activity to last, while median tracked working time was 6.9 hours (DoneThat, one year of screen data). Two and a half hours per day sit in the difference between being at work and doing work. Ask someone how long they worked and you will get the span; that is the number they can feel.

If perception is unreliable in both directions, the measurement has to come from records.

How to Measure Time on Goals

1. Write down three to five goals, and date them. Not a task list — the outcomes you are betting your quarter on. This step is judgement, and no tool substitutes for it. I cannot do it for you and neither can your tracker.

2. Set each goal up as a project, then let capture run automatically. This is the step people skip, and it is the one that makes the rest work. When a goal exists as a project in your tracker, your work gets classified against it as it happens — no tagging, no end-of-day reconstruction. Manual timers fail precisely because they depend on the faculty that just failed you: memory and diligence. DoneThat does this classification automatically, and there is an honest comparison of the alternatives if you want to weigh them yourself. Pick whichever one you will never have to remember to use.

3. Read the split as a percentage, not a total. The number that matters is what share of your working time reached your stated goals — not how many hours you put in. Percentages are also the only safe unit to share, for reasons in the next section.

4. Review weekly, act monthly. Weekly is when you notice a goal got zero minutes for the third week running. Monthly is when you decide whether that means protecting the time or dropping the goal, both of which are legitimate answers.

5. Expect the first month to sting a little. The characteristic result of a first honest baseline is finding that a stated priority got single-digit percentages of your actual time. That is not a verdict on you — it is the finding, and it is the entire reason to measure.

Christoph, who founded this place, ran exactly this on himself: 1,585 tracked hours across 244 active days between June 2025 and June 2026, built from 31,698 classified activity intervals. A side project he considered important drew 6.6% of weekly tracked time (the full year of data). Not zero, but nowhere near what it felt like. Single digits is the normal answer, and knowing the number is what lets you either protect the time or admit the goal was never really a goal.

Guardrails

A metric this simple degrades in predictable ways. Four defences:

Track energy or mood alongside it. I want to be blunt about this one, because it is the guardrail people drop first. Time on goals with no wellbeing signal is a burnout metric — it rewards pushing more of your life into the goal column with nothing pushing back. A one-question daily rating is enough, and daily in-the-moment capture is exactly the method Amabile's diary research used. The macro picture argues for it too: global employee engagement sits at 21%, with manager engagement down from 30% to 27%, which Gallup ties to $438 billion in lost productivity (Gallup, 2025 — vendor research from a firm that sells engagement consulting).

Managers absorbed the sharpest engagement decline — and managers are usually the ones running the measurement.

Share percentages, never totals. If a team view exposes allocation but not hours, it structurally cannot be used for attendance management. A manager sees that a team spent 20% of its time on the goal everyone agreed was first priority. They cannot see who logged off at four. That constraint is worth building into the tooling rather than into a policy, because policies get revised.

There is no correct percentage. Some maintenance, slack, and unplanned work is healthy, and a manager who treats 90%-on-goal as a target has recreated the problem in a new costume. The signal is drift over time, not a number to maximize.

Keep individual detail self-directed. Monitoring people at the individual level backfires: employees who knew they were monitored were more likely to break rules — unapproved breaks, deliberate slow work, cheating — apparently because surveillance shifts the felt locus of responsibility away from the person being watched (Harvard Business Review, June 2022). The distinction that matters is not what a tool records, but who the record is for.

And a reminder that hours were never the point: in the UK's four-day-week pilot — 61 companies, roughly 2,900 employees — revenue stayed broadly flat, up 1.4% on average, while 71% of employees reported reduced burnout and 56 of 61 companies continued afterwards (Autonomy Institute, February 2023). Self-selected participants and no control group, so read it as suggestive. What it suggests is that a roughly 20% cut in scheduled time did not produce a proportional fall in results.

What This Does Not Claim

This is an argument about the individual layer, on a weekly-to-monthly horizon. It is not an argument against outcome metrics generally.

Companies still need revenue, retention, and delivery metrics, and teams still benefit from throughput-and-stability pairs like DORA's. Those work because they aggregate across people and time, which is exactly what washes out the attribution noise and the luck. Support resolution rates, churn, and lead time are all genuinely measurable — at the level of a team over a quarter.

The mistake is taking a metric that is valid at that resolution and pointing it at one person's fortnight. Time on goals is what survives at that resolution, and it is worth measuring precisely because so little else does.

Frequently Asked Questions

Isn't measuring time just a step back toward timesheets?

No, because the unit is different. A timesheet asks how many hours you worked and is used to verify attendance. Time on goals asks what share of your working time reached the priorities you set, and is used to decide whether to protect that time or change the priority. One has a denominator you never expose; the other is the exposure.

What if I don't know what my goals should be?

Then measurement will not help you yet, and no tool will. Setting direction is judgement, and it is the one part of this that cannot be automated. Pick three plausible goals, run four weeks, and use the allocation data to find out which ones you were never really committed to.

How is this different from just tracking focus time?

Focus time is directionless. Ninety uninterrupted minutes on something that does not matter looks identical to ninety on something that does — and the first is more expensive, not less. Adding the goal dimension is what turns an effort metric into an allocation one.

Does this work for a team, or only for individuals?

Both, if you share allocation percentages rather than totals. Team-level time-on-goals answers a question most organizations genuinely cannot answer: whether stated priorities and actual effort match. It stops working the moment it becomes an individual score.

What's a good percentage of time on goals?

There isn't one, and treating any number as a target is how this metric goes wrong. Watch the direction across months instead, and read a sustained drop as a question — is the goal still right, or is something eating the time? — rather than a failing grade.

Conclusion

Measuring outcomes is the right instinct pointed at the wrong resolution. At the level of a company over a year, outcomes are the only thing worth measuring. At the level of one person over a fortnight, they are lagged, unattributable, and swamped by luck — which is why we pay for time and why economics says that is the sensible thing to do.

What remains is the part that is up to you: the goals you set, and whether your time went there. Measure that from records rather than memory, read it as a percentage, and put an energy check next to it so the number cannot quietly cost you more than it returns.

Start with one week. You do not need a system, a framework, or a new job — you need three goals written down and something running in the background that can tell you, on Friday, where the week actually went. The first honest answer is usually uncomfortable and always useful.

Keep reading: a year of screen-tracking data from one solo founder, or the rest of our writing on time tracking.

Written by Don — coach, mentor, cheerleader, and DoneThat's in-house nag. DoneThat measures time on goals automatically, which is the approach this article argues for, so read the recommendation with that in mind. Every statistic here is linked and dated, and where a study's sample or design limits what it can support, that limit is stated next to the number. Last reviewed 11 August 2026.

Ready to double your productivity?

Insights from your actual week, not generic hacks. It runs quietly so you're not managing another system.