When a Measure Becomes a Target, How Can We Tell Real Improvement from Numerical Improvement?

When Numbers Start Making Decisions · Season Two, “When the Measure Becomes the Target” · Article 12

1. Twelve better scores do not describe one kind of change

A higher four-hour emergency result can mean better patient flow or a different departure timestamp. Fewer overdue elective surgeries can come from greater capacity or a waiting-list clock that begins later. More bank sales can reflect the discovery of genuine needs or the displacement of customer interests by a scorecard.

An injury rate of zero can mean hazards were removed or reports were suppressed. Better train punctuality can result from genuine reliability, a slower timetable or the exclusion of cancellations—although New South Wales repaired one such boundary by treating cancelled and skipped-stop services as late.

A higher university rank can accompany better research and teaching or optimisation of citations and reputation. A five-star aged care home combines residents’ experience, compliance, staffing and quality measures but does not guarantee every shift. A platform’s satisfaction score aggregates customer reactions, while crossing an 85 per cent line can determine whether a courier keeps access to work.

Health Stars can encourage less sugar and sodium or lead a product team to optimise compensating variables. NABERS reads operational energy and therefore sits relatively close to a real outcome. The 2–3 per cent inflation target goes further still: when people believe and adapt to it, the measure participates in creating the stability it seeks.

If this season is compressed into the sentence “when a measure becomes a target, it ceases to be a good measure”, its central finding disappears. A target does not inevitably fail. It enters the measured structure. Real improvement, local optimisation, substitution, avoidance, reclassification and stabilising feedback can follow. The purpose of governance is not to prevent all adaptation, but to align the easiest path to a better score with the underlying purpose.

2. First distinguish four forms of numerical improvement

The first is outcome improvement. A building reduces energy while maintaining service; an emergency patient reaches appropriate care sooner; the critical controls at a worksite become more reliable. The measure and the world improve together.

The second is process improvement. Training, meetings, inspections and installation of equipment increase. These activities may produce outcomes, but do not prove them. Public organisations commonly mistake completion of an activity for final success.

The third is composition change. The people admitted by a university, treated by a hospital or served by a platform change, improving the average. Avoiding complex cases, for example, can reduce average waiting while excluding those most in need.

The fourth is recording change. Departure is redefined, a clock is paused, a category altered, an injury omitted or a denominator changed. The score improves without a corresponding change in service.

The four cannot be identified from intent alone. A short-stay unit may improve care or simply stop the emergency clock. Light duties may aid rehabilitation or conceal a lost-time injury. Evidence must reconnect the measure with its causal chain and omitted outcomes.

3. Test one: how far is the measure from the final purpose?

From resources to public purpose, we can sketch a chain:

resources → activities → outputs → service quality → individual outcomes → long-term social impact

Items towards the left are easier to count and complete. Those towards the right are closer to the purpose and more affected by external conditions. Training sessions are activities, completed operations are outputs, and a patient’s restored function is an outcome. Selling a product is an output; the customer’s long-term financial interest is the purpose.

NABERS is relatively strong because it rereads twelve months of measured energy rather than counting retrofit projects. LTIFR is hazardous when it records one class of realised injury but is described as risk itself.

In 2026, the Australian National Audit Office summarised 504 performance results across 21 Commonwealth entities. It warned that counting targets met or not met is insufficient to assess public performance: context, target difficulty, data classification and factors outside an entity’s control also matter. The audit work also found that efficiency and outcome impact remained undermeasured, with many entities reporting what they did rather than what changed. ANAO: “Performance Statements of Major Australian Government Entities—2024–25 Audit Program”

The first question is therefore not whether the figure is accurate. It is: even if perfectly accurate, how many steps remain between this measure and the purpose? The greater the distance, the less decision-making authority it should carry without corroboration.

4. Test two: what is the cheapest route to a better score?

For every KPI, perform an adversarial design exercise: if I cared only about the number and not the original purpose, what would be the cheapest way to improve it?

If the cheap route in emergency care is an earlier administrative transfer, audit physical occupancy. If the cheap university-ranking route is hiring highly cited researchers, examine teaching and disciplinary distribution. If the cheap aged-care route is increasing average minutes, inspect nights and the professional mix. If the cheap safety route is non-reporting, remove penalties attached to reporting.

This question does not presume bad faith. People respond to clear signals under scarcity, and managers attend to areas carrying the strongest accountability. An institution that does not know the cheapest route through its measure cannot distinguish intended response from side effect.

The ideal design makes the cheapest score improvement a real improvement. NABERS gets relatively close through operational data, annual reassessment and accredited verification. Rail punctuality closed an obvious shortcut by treating cancellation and skipped stops as late. Health Star reformulation is better aligned when reducing sodium is a more effective route than adding a convenient compensating ingredient.

5. Test three: who disappeared from the denominator?

Average results often improve because the population changes. Elective surgery time begins at formal listing, leaving people before the list outside the clock. Platform ratings include completed and rated orders. A graduate-employment result can omit non-respondents. Aged-care averages may obscure residents unable to express their views.

Every performance report should therefore include a “missing population” account. Who belongs to the underlying service purpose but did not enter the denominator? Who exited, transferred, was refused or was reclassified during the period? What happened to them?

Reporting only successful completers produces institutional survivor bias. When the most difficult cases leave, the remaining population looks better. If refusing high-risk cases improves the score, a measure can reward exclusion.

Denominator audit must include geography, disability, age, language and complexity. A better total does not show that every group benefited. Where the indicator distributes public resources or access to work, distributional harm cannot be offset by the average.

6. Test four: is there unnatural bunching near the boundary?

When 240 minutes, 90 days, 85 per cent or 4.50 stars matter, observations can cluster immediately before the line. The bunching can show managers appropriately rescuing cases that are about to become overdue. It can also show optimisation of timestamps, classifications or rounding.

Auditors should inspect continuous distributions instead of only the pass rate. Emergency departments should report intervals and long tails. Elective surgery should compare days 89, 90 and 91 with recategorisation. Aged care should retain the unrounded score. A platform should display sample size and uncertainty.

The smaller the real difference across the boundary and the larger the difference in consequence, the more important review becomes. Season Two moved from the people on either side of a threshold to behaviour around a target, but the boundary remains the place where numerical authority becomes easiest to see.

7. Test five: are there conditions that a high total cannot offset?

Composite scores are convenient and compensating. Sales can offset complaints, staffing minutes can offset poor experience, and punctuality can offset extreme delays. Some values should not be averaged.

Aged care provides a useful example. One-star compliance forces a one-star overall rating, and two-star compliance caps the overall result at two. Serious non-compliance cannot be washed away. Bank remuneration should similarly allow material customer harm and misconduct to reduce a bonus to zero instead of being compensated by profit.

A measurement system therefore needs three structures: combined indicators for orientation, hard constraints against severe harm, and qualitative judgement to interpret cases. None can replace the others.

Safety, legality, basic rights and data integrity usually belong among the non-compensable conditions. If a target is achieved by violating any of them, the result should not be reported as success.

8. Test six: can reality disprove and revise the number?

An authoritative measure must allow reality to return and correct it.

At the case level, a courier should be able to use location, messages and restaurant-wait evidence to challenge attribution of a low rating. A patient should be able to request review of a surgery category. Providers and the public should be able to correct aged-care data, while an appeal must not keep urgent risk hidden.

At the system level, overturned cases, anomalous distributions, research and user feedback must enter the next version of the algorithm and definition. Health Stars changed after review, aged-care staffing and compliance rules were revised in 2025–26, and NABERS separated renewable-energy sourcing from the efficiency star. These are signs of a measurement system that can learn.

Revision damages simple time-series continuity but can increase legitimacy. Versioning is the answer: retain the old method, mark the break and, where necessary, calculate both versions in parallel. A clean graph is not a reason to preserve a known defect.

ANAO similarly recommends that performance measures remain contemporary, receive regular fit-for-purpose review and move from outputs towards outcomes as programs mature. ANAO: “Reporting Meaningful Performance Information”

9. Test seven: who still bears judgement and consequence?

As an indicator matures, an organisation can call it an objective result. Every measure still contains decisions about purpose, data, weights, thresholds, exceptions and consequences. Responsibility can be distributed; it cannot disappear.

The 2–3 per cent target does not set a mortgage automatically. The Monetary Policy Board must judge persistence, employment, lags and distributional cost and explain its conclusions. An 85 per cent satisfaction rate does not deactivate a worker on its own. A platform chooses to connect feedback with account access and must provide a fair procedure and remedy.

Systems can aggregate, predict, warn and execute. A person or institution still has to say that the consequence is proportionate to the evidence, give reasons and accept responsibility for error. “Data-driven” must not become a phrase for expelling judgement.

10. An eight-question framework for practice

When an organisation claims that a metric improved, ask:

  1. Purpose: which real-world outcome are we trying to improve?
  2. Distance: is the measure an input, activity, output, proxy or final outcome?
  3. Shortcut: what is the cheapest way to raise the score without improving reality?
  4. Denominator: who is absent, exited, refused or reclassified?
  5. Distribution: does the average conceal harm to a group, time period or extreme tail?
  6. Constraints: can safety, legality or basic rights be offset by another high result?
  7. Disproof: can an individual see reasons, correct data and obtain review by someone able to change the outcome?
  8. Learning: can reversals, anomalies and side effects change the next version and comparable past decisions?

The framework does not promise one simple overall score. Its purpose is to prevent metric governance from becoming another KPI.

Conclusion: a measure cannot stand outside the structure and pretend only to observe

Reality is continually formed through relationships, constraints, differences and correction. Once a measure enters an institution, it becomes part of those relationships. It reallocates attention, resources, reputation and responsibility. The measured people and organisations learn, and their adaptation changes the meaning of the measure in return.

We cannot distinguish real improvement from numerical improvement by reading the number alone. We must inspect the causal chain, cheapest route to a higher score, missing denominator, distribution, non-compensable constraints and capacity for disproof and revision.

My final judgement is that modern institutions need measurable targets and a structure in which evidence beyond the target constrains them. A measure may guide action but should not monopolise the definition of success. It can delegate knowledge but cannot eliminate the person responsible for judgement.

Season One concluded that a number’s decision-making authority cannot exceed its measurement capacity. Season Two adds:

A measure’s power to reward cannot exceed the institution’s capacity to identify the behaviour that reward induces.

Only when a better score is confirmed by real outcomes, boundary audit and continuing correction do we have reason to say that the number improved and the world moved forward with it.

Principal sources


Discover more from Geoffrey Chen

Subscribe to get the latest posts sent to your email.