When Numbers Start Making Decisions · Season One, “People on Either Side of a Threshold” · Article 12
1. Robodebt did not begin by asking whether the number could prove a debt
An Australian social-security recipient in casual or part-time work might earn a large amount in some fortnights and nothing in others. Entitlement was assessed from actual income in each fortnight, while taxation data ordinarily recorded a total across a longer period.
Robodebt averaged that total across the relevant fortnights and compared the result with income previously reported to Centrelink. If a person did not respond or supplied incomplete records, the system could use the average to fill the gaps and raise a debt. Work once undertaken by compliance officers was moved into an online process, while the person had to find old payslips or employer records to disprove a number already generated by government. Commonwealth Ombudsman: “Centrelink’s Automated Debt Raising and Recovery System”
The arithmetic could be perfectly correct without proving actual income in any fortnight. “Average income per period” and “income actually earned in this period” are different facts.
The income-averaging method was later acknowledged to lack a lawful basis, and the Robodebt Royal Commission recommended clear review paths, plain-language disclosure of automated processes, independent expert scrutiny of business rules and algorithms, and institutional oversight of technical operation, fairness, bias and user experience. Royal Commission into the Robodebt Scheme: Final Report
Robodebt was not a machine-learning risk score. It relied mainly on data matching, averaging and rule automation. That makes its lesson clearer: danger does not begin only with a mysterious black box. An ordinary calculation authorised to refuse, recover money or change rights also requires explanation, review and correction.
2. Was the person rejected by a number or by an institution?
“The algorithm rejected me” is understandable but incomplete. A score has no legal personality and does not choose its own threshold. A typical chain includes:
- an institution defines the policy purpose;
- data providers record or infer personal information;
- designers choose variables, formulas, training data or business rules;
- decision-makers set a threshold and attach a consequence;
- software executes at scale; and
- staff, courts or platforms accept, override or maintain the result.
Responsibility can be distributed without disappearing. Government cannot answer “the system calculated it”. A bank cannot end with “the model found the risk too high”. A university cannot treat a completed rank as an explanation.
Appeal first rebuilds an address for responsibility. The affected person needs to know which institution can change the decision, who must state reasons, who can correct inputs and who is responsible for a system-wide error.
3. Explanation is not appeal
An institution may call itself transparent because it says, “Your score was below the standard”, “Your income exceeded the limit” or “The system classified you as high risk”. These are outcome notices, not reasons.
An explanation capable of supporting challenge must address at least four layers:
- input: which information about me was used, where it came from and which period it covered;
- measurement: what score, class or inference was produced and what it actually measures;
- threshold: where the line is, why it is used, and which exceptions or discretion exist; and
- decision: what role the number played in this case and what other reasons led to the outcome.
Even a complete explanation is not yet an appeal. Explanation allows a person to understand. Appeal creates an opportunity to change an erroneous outcome.
The Commonwealth Ombudsman’s 2025 Automated Decision-making—Better Practice Guide recognises the difference. Automated systems should not prevent access to reasons or information held by an agency. Notices should identify challenge and appeal paths. Systems should preserve reasoning and audit records that internal and external reviewers can understand, and agencies should design remedies for large-scale failure before deployment. Commonwealth Ombudsman: “Automated Decision-making—Better Practice Guide”
4. Effective appeal must be able to examine six different failures
When people say “the number is wrong”, they can mean different things.
1. Are the inputs accurate?
A credit report may include another person’s account, a benefit system may use the wrong tax year, or a risk assessment may contain an incorrect criminal history. Review must permit inspection, correction and tracing of data sources.
2. Does the number measure the relevant fact?
The CPI can accurately measure average price movement without representing one household’s cost of living. Annual averaging can be arithmetically correct without proving fortnightly income. Here the inputs and arithmetic can be accurate while the measurement does not answer the legal question.
3. Is the threshold justified?
BAC 0.05 converts continuous risk into an enforceable boundary. ATAR 90.00 may manage entry to a limited pool. An income threshold makes entitlement administrable. An individual reviewer may lack power to rewrite the line, but the institution must disclose who selected it, for what purpose, and on what evidence and trade-offs.
4. Is the consequence proportionate to the number’s capacity?
Emergency triage can order care but is not a complete diagnosis. A group recidivism score may inform supervision planning but should not determine imprisonment. Relevance to a decision does not confer capacity to carry the entire consequence.
5. Could the person genuinely participate?
A notice sent to an old address, an inaccessible portal, unexplained technical language or an impossible burden of producing old records can make a formal review right useless. Procedural fairness concerns whether people can actually enter the process, not only whether an entrance exists.
6. Can correction repair the real loss?
Removing an incorrect record, restoring eligibility, stopping debt recovery, refunding money, restoring an opportunity and compensating appropriate loss can all form part of a remedy. Changing a database status to “resolved” is incomplete if the right and the real-world loss remain unrestored.
Separating these failures prevents an institution from using “the arithmetic was correct” to avoid an invalid measurement, “the threshold is lawful” to avoid a disproportionate consequence, or “complaints are available” to avoid the fact that the reviewer cannot change the outcome.
5. What counts as genuine human review?
“Human in the loop” is often offered as a guarantee, but human presence does not necessarily produce judgement.
If a reviewer sees only a red alert and not the source data, must complete a case in minutes, needs several approvals to disagree with a model, does not understand its limits, or merely resubmits the same inputs to the same process, the human is a rubber stamp.
Genuine review requires that the reviewer:
- can change, cancel or pause the original decision;
- can inspect source inputs, the calculation path, applicable rules and model limitations;
- can receive new information and explanation from the person;
- is not unreasonably penalised for overriding the system;
- gives independent reasons rather than repeating a label; and
- has appropriate independence in high-risk matters.
The value of human judgement is not that people are naturally accurate. People are biased, tired and inconsistent. Its distinctive value is that someone can confront facts the model did not encode, compare conflicting reasons and accept responsibility for the final choice.
6. An individual appeal must feed error back into the system
Robodebt exposed a deeper failure: an individual can win review while the institution fails to learn.
Between 2016 and 2022, first-tier Administrative Appeals Tribunal reviewers made hundreds of decisions questioning whether income averaging proved actual income, overpayment or debt. Some found the method unlawful. Those first-tier decisions were not public, and adverse results did not become an effective system-wide correction. A later Royal Commission recommendation addressed publication of significant social-security review decisions. Federal Court of Australia: Justice Kyrou, “Mechanisms in the ART Bill to thwart Robodebt-type maladministration”
Appeal has two functions. It corrects one case by removing a debt or restoring a right. It also enables institutional learning by identifying a rule that repeatedly generates the same failure.
If one hundred people win separately and the same erroneous decision is sent to the next person, review is only cleaning up damage produced by automation. Correction must feed into business rules, model versions, training, notices and earlier cases. Sometimes the agency must proactively identify all affected people rather than waiting for each to appeal.
7. Minimum conditions for a number authorised to refuse
Safeguards should respond to consequence. A mistaken music recommendation and a calculation that creates debt or affects liberty do not require the same process.
For a digital decision that materially affects eligibility, opportunity, property, health, safety or liberty, the minimum should include:
- a defined purpose;
- clear legal authority and accountable owners;
- verifiable measurement, data sources and error limits;
- a defensible threshold and proportionate consequence;
- case-specific reasons available to the person;
- review with power to change the outcome;
- timely, complete remedy and an appropriate pause against irreversible harm; and
- system learning through audits, outcome monitoring and correction of similar cases.
Australian policy now contains parts of this structure. The federal government’s responsible AI policy 2.0, effective from December 2025, requires applicable agencies to nominate accountable officials, maintain internal use-case registers, undertake impact assessment and publish transparency statements. Digital Transformation Agency: “Policy for the responsible use of AI in government” New Privacy Act transparency obligations concerning certain automated decisions are due to commence on 10 December 2026.
These measures matter, but a register, privacy policy or transparency statement does not automatically create an individual right of appeal. Knowing that an agency uses AI and being able to require it to change a refusal are different things.
8. Appeal has limits, but they must be explicit
Appeal does not guarantee the outcome a person wants. A university reviewer can correct marks and procedure but cannot create unlimited places. A benefit reviewer can correct income but cannot ignore a valid statutory threshold. A court can correct misuse of a risk score without ignoring relevant proved conduct.
What effective appeal guarantees is a decision based on correct facts and lawful rules, a real examination of reasons, and a route for error to change the result.
Different paths also do different work. Data correction addresses an input. Internal review reconsiders an agency decision. Merits review re-examines facts, law and policy to reach the correct or preferable decision. Judicial review primarily examines legality. Complaints and oversight can reveal service or system problems but may not replace a process empowered to alter the individual outcome.
For Centrelink decisions, a person ordinarily seeks formal review by a Services Australia Authorised Review Officer, and many decisions can then be reviewed by the Administrative Review Tribunal. The ART makes its own decision on the relevant facts, law and policy. Administrative Review Tribunal: “Applying for a review”
Conclusion: decision power and appeal must be designed together
Season One has moved from the one-hour employment rule through CPI, ATAR, NAPLAN, BAC 0.05, credit scores, emergency triage, fire danger, 1.5°C, benefit thresholds and judicial risk scores. These numbers do different work: description, coordination, ordering scarce resources, or changing individual rights.
Their common issue is not perfection. Every institutional tool compresses reality and has error. The critical question is how much decision power a number receives and how much explanation and correction power the affected person receives in return.
My answer is yes:
If a score, category, classification or calculation can reject a person, it must open an appeal path proportionate to the consequence and capable of genuinely changing an erroneous result.
Not every refusal is wrong. Appeal exists because an institution accepts that it can be wrong. It keeps responsibility identifiable and prevents error from spreading at the speed and scale of automation.
Modern institutions cannot operate without numbers. But a number’s decision power must not exceed its measurement capacity, and a decision’s consequence must not exceed the institution’s capacity to explain, review and correct it. Appeal is not customer service added after digital decision-making. It is part of what makes the decision legitimate.
Primary sources
- Royal Commission into the Robodebt Scheme: Final Report
- Commonwealth Ombudsman: Centrelink’s Automated Debt Raising and Recovery System
- Commonwealth Ombudsman: Automated Decision-making—Better Practice Guide
- Administrative Review Council: Best Practice Guide—Statements of Reasons
- Federal Court of Australia: Mechanisms in the ART Bill to thwart Robodebt-type maladministration
- Administrative Review Tribunal: Applying for a review
- Digital Transformation Agency: Policy for responsible use of AI in government
- OAIC: APP 1 automated decision-making obligations
Discover more from Geoffrey Chen
Subscribe to get the latest posts sent to your email.