How Far Can a Group Average Be Used to Judge an Individual?

When Numbers Start Making Decisions · Season Three, “The People the Average Does Not See” · Article 12

1. Twelve cases, four different kinds of number

This season began with an area index. SEIFA can identify concentrations of educational and occupational disadvantage and guide university funding, but it cannot see every student inside an SA1. The Modified Monash Model distinguishes degrees of rurality, but an MM2–MM3 line cannot contain every clinic’s recruitment problem.

It then moved to measurements and equations. BMI simplifies height and weight into a screening ratio. Pulse oximetry estimates arterial oxygen through light, with uncertainty that can be distributed unequally by skin pigmentation. eGFR estimates kidney filtration from creatinine and population equations, so 59 and 60 are adjacent estimates before they are categories.

A 50th-percentile male crash dummy standardises a safety test without representing every body. Ten grams of alcohol standardises exposure without standardising its effect. A five-year cardiovascular risk score looks personal but predicts from outcomes among groups.

Finally, the season examined administrative and household models. JSCI predicts relative employment difficulty and can open support or impose control. Default Market Offer comparisons use representative electricity consumption to compare plans, not forecast every bill. NatHERS uses a standard household to compare building designs, not promise one family’s energy use.

These are not one error repeated twelve times. They are four structures: an area proxy, a standard comparison, an estimated measurement and a risk prediction. Each transfers group evidence to a decision. The ethical question is how far the transfer may go.

2. Begin with level fit: what does the number actually describe?

The first test is whether the level of measurement matches the level of decision.

SEIFA describes an area. It is well suited to place-based outreach and institutional resource allocation. It is less suited to deciding whether one student is in hardship. Modified Monash describes a location and can structure geographic incentives, while a specific practice may require additional workforce evidence.

The ecological fallacy occurs when an association at group level is attributed to each member. The ABS explicitly warns that SEIFA is assigned to areas rather than individuals. ABS: Using and interpreting SEIFA

Level mismatch is not always fatal. An area proxy may be the only practical way to distribute a large national fund. Its authority should weaken as the consequence moves towards an individual. A map can allocate outreach funding with no personal exclusion; it needs stronger safeguards if it blocks a particular application.

The rule is: evidence may travel between levels, but it must not silently change what it describes.

3. Then test use fit: which question was the number built to answer?

The same number can be valid for one purpose and invalid for another. BMI can support population surveillance and clinical screening without measuring individual body composition. Normalised eGFR can support CKD classification yet require adjustment or confirmation for a high-consequence dosing decision. A NatHERS thermal star compares design performance, not total household electricity cost.

Use fit asks whether the downstream decision resembles the task for which the model or measurement was validated. “Accurate” has no meaning apart from a purpose and acceptable error. An oxygen reading that is adequate for observing a trend may be insufficient to discharge a symptomatic patient. A standard drink count useful for weekly risk guidance is not a driving-time formula.

Institutions often lose this boundary through data reuse. Once a number exists in a record, it becomes tempting to apply it to eligibility, pricing, enforcement or performance management. The administrative cost of reuse is low; the validity cost can be high.

Every consequential use should therefore state the original measurement purpose, the new decision and the evidence connecting them.

4. Audit the distribution, not only the average error

A group average can conceal who receives the error. Pulse oximeters may meet overall performance requirements while overestimating low oxygen more often in darker-skinned patients. A standard crash dummy can support safer cars while leaving smaller or older occupants less well represented. A risk model can be calibrated overall and miscalibrated within a population.

Distributional audit asks:

  • Which groups were represented in design and validation?
  • What is the direction, not just the mean size, of error?
  • Does error increase near the threshold where action changes?
  • Who bears false reassurance, false exclusion or unnecessary intervention?
  • Can a strong aggregate outcome compensate for harm concentrated in one group?

The Australian AI Ethics Principles include fairness, transparency, contestability and accountability. Department of Industry: Australia’s AI Ethics Principles These principles are useful beyond artificial intelligence because the governance problem begins before automation. Any model that turns group patterns into individual consequences needs the same discipline.

Average accuracy is necessary evidence. It is never the complete fairness assessment.

5. Make consequences proportional to the evidence

Not every error has the same cost. An imprecise proxy used to offer optional tutoring is different from the same proxy used to deny financial assistance. A risk score used to invite a cardiovascular prevention discussion is different from one that automatically starts treatment. JSCI used to offer tailored help is different from JSCI used to intensify obligations.

The direction of action matters. Rough evidence is more defensible when it opens opportunity, triggers review or adds resources. It requires greater precision and contestability when it closes access, labels a person permanently or imposes a sanction.

This asymmetry can be expressed as a governance principle:

The more irreversible, coercive or individual the consequence, the less a group proxy may decide by itself.

Systems should design false positives and false negatives deliberately. A broad health screen may accept more false positives because assessment can correct them and missing a serious condition is worse. A public benefit program may need an exception path because false exclusion deprives someone who cannot absorb the loss.

6. Keep the threshold visible as an institutional choice

The lowest 25 per cent of SEIFA, MM3, BMI 35, eGFR 60 and five-year CVD risk of 10 per cent are useful lines. None is a tear in nature. Continuous variation has been turned into an administrable category.

Near a line, the measured difference is usually smaller than the difference in consequence. That demands reliable inputs, clear rounding, versioning and a path to consider adjacent evidence. A category should tell a worker what process begins, not absolve the worker from asking whether the process fits.

Thresholds also need outcome review. Are people just below BMI 35 with serious disease receiving appropriate alternatives? Do practices just outside an MM category show equal workforce difficulty? Do eGFR results fluctuate across 60 without meaningful clinical change? Bunching and reversals reveal where the structure is under pressure.

The aim is not endless exceptions. It is a standard rule plus an intelligible mechanism for cases in which the rule’s own purpose would be defeated.

7. Give explanation and correction to the affected person

Contestability has several layers. A person should be able to correct source data: address, height, smoking status, blood pressure, health history or employment circumstances. They should be able to understand the classification method sufficiently to identify misapplication. They should also be able to present evidence that the proxy omitted when the consequence warrants individual review.

These are different rights. Correcting a geocoded address does not change the SEIFA method. Correcting creatinine does not determine whether the eGFR equation fits an unusual body. Updating a JSCI answer does not answer whether the resulting service is appropriate.

The reviewer must have authority to change the outcome. A complaints mailbox that can only explain the score is not meaningful review. Reasons should identify the influential evidence, applicable rule, version and available next step without exposing protected model details or other people’s data.

Australia’s public-sector work on automated decision-making has similarly emphasised transparency and reporting around systems that affect people. OAIC: Automated decision-making and public reporting The deeper principle is institutional: delegated calculation does not remove responsibility for reasons.

8. Make lived outcomes capable of revising the model

Models are often evaluated before deployment and then treated as infrastructure. Season Three shows why post-deployment evidence is indispensable.

Actual crash injuries should broaden test occupants. Clinical discrepancies should improve pulse-oximeter validation and guidance. Repeated eGFR misclassification should refine reporting and confirmation. Employment outcomes and participant experience should reshape service assessment. Actual household energy and comfort should test NatHERS assumptions. Bill-based plan rankings should test whether a representative electricity comparison remains useful.

Learning requires versioning. When the method changes, institutions should record what changed, why, when it applies and whether affected past decisions need review. Preserving a clean historical series is not a reason to preserve a known inequity.

A trustworthy number is not one reality can never contradict. It is one embedded in a system capable of hearing contradiction and improving.

9. An eight-question framework for group evidence and individual judgement

Before allowing a group-derived number to influence a person, ask:

  1. Level: Does the number describe an area, product, population, body or individual?
  2. Purpose: What question was it designed and validated to answer?
  3. Fit: Does this person and decision fall within that scope?
  4. Distribution: Which groups receive larger or systematically directed errors?
  5. Boundary: How uncertain is the result near the action threshold?
  6. Consequence: Does the number open support or close opportunity, and is the consequence proportionate?
  7. Contestability: Can data, application and omitted individual evidence be reviewed by someone able to change the result?
  8. Learning: Can outcomes, reversals and subgroup harms revise future versions and affected past cases?

No single overall score can replace these questions. Their purpose is to keep the relationship between evidence and authority visible.

Conclusion: the individual is not an error term around the mean

Group evidence is one of modern institutions’ most powerful achievements. It reveals patterns no individual case can show, supports consistent standards and makes national allocation possible. Rejecting it would return decisions to opaque intuition and unequal discretion.

But a group measure does not contain every person within the group. My final judgement is that it may orient, screen, compare and allocate in proportion to its demonstrated fit. As consequences become more personal, coercive or irreversible, individual evidence, explanation, review and responsibility must become stronger.

The average is not the enemy of the individual. The danger begins when an institution forgets which one it measured.

Season Three’s conclusion is therefore:

Use group evidence to see structure; use individual evidence to judge the person; and design every boundary so reality can answer back.


Discover more from Geoffrey Chen

Subscribe to get the latest posts sent to your email.