When an Algorithm Calls Someone “High Risk”, How Far Can a Judge Trust It?

When Numbers Start Making Decisions · Season One, “People on Either Side of a Threshold” · Article 11

1. Three “high risk” bars enter a courtroom

In Wisconsin, Eric Loomis pleaded guilty to attempting to flee or elude a traffic officer and operating a motor vehicle without the owner’s consent. The state connected him to a drive-by shooting, which Loomis denied. Other charges were dismissed under the plea agreement but could be read in at sentencing.

The presentence investigation included a COMPAS risk assessment. It classified his pretrial, general recidivism and violent recidivism risks as high. The prosecution referred to those results, and the sentencing judge said that the seriousness of the offences, Loomis’s history on supervision and the risk-assessment tools indicated a high risk of reoffending. Wisconsin Supreme Court: State v. Loomis, 2016 WI 68

The court imposed consecutive sentences totalling six years of initial confinement followed by five years of extended supervision. Loomis sought resentencing, arguing in part that use of a proprietary risk assessment he could not fully examine violated due process.

The Wisconsin Supreme Court upheld the sentence in 2016. It did not hold that COMPAS may determine punishment. It drew a narrower line: subject to warnings, restrictions and independent reasons, a judge may consider the score among other factors. It must not determine whether a person is imprisoned, the severity of sentence, or whether the person can be safely supervised in the community.

That distinction leaves a hard question. Once “high risk” appears in formal sentencing material, can a judge see the label without granting it more authority than its evidence warrants?

2. Whose future does COMPAS measure?

COMPAS stands for Correctional Offender Management Profiling for Alternative Sanctions. The record described it as decision support for placement, management and treatment planning in corrections. It used criminal-history information and an interview. Risk results appeared as bars from one to ten, while needs scales identified areas such as employment, housing and substance misuse.

“Prediction of recidivism” can sound like a statement about an individual future. The decision stresses that COMPAS does not calculate the specific probability that one person will offend. It compares the person with data from groups sharing similar characteristics and identifies the group as lower or higher risk.

There is an irreducible distance between a group prediction and an individual fact. Even a model that accurately identifies a group with a higher rate cannot tell a court what this defendant will certainly do. Many people in a higher-risk group will not reoffend; some in a lower-risk group will.

That is also why “high risk” does not mean “more guilty”. Conviction concerns conduct already proved according to law. A risk assessment concerns the possibility of future conduct. Turning prediction directly into greater punishment makes a person bear a present consequence for group statistics and an uncertain future.

3. What does a score from one to ten leave out?

A score combines criminal history, supervision experience and some dynamic factors. It can reduce complete dependence on intuition and provide more consistent information across cases. A judge cannot personally remember the long-term outcomes of thousands of comparable defendants.

But a score omits at least three kinds of information.

First, it cannot fully represent individual context. Family support, a recent change in circumstances, access to treatment, stable employment and unusual facts of the offence may point in a different direction from historical data. The model responds only to encoded variables.

Second, it conceals the social structure behind those variables. Education, employment, housing and previous justice contact look like personal attributes but are also shaped by opportunity and enforcement. Even without race as a direct input, correlated variables can carry historical inequality. The Supreme Court therefore required warning that research had questioned whether COMPAS disproportionately classified minority defendants as higher risk and that changing populations require monitoring and recalibration.

Third, the score does not settle the normative purpose of sentencing. Recidivism risk is relevant to only part of the decision. A court also considers seriousness, responsibility, deterrence, rehabilitation and public protection. A tool built for case management and treatment planning cannot answer how much punishment a person deserves. Predictive performance and penal justification are not the same question.

4. The inputs are visible, but the weights are not

The most famous dispute in Loomis concerned trade secrecy. The developer treated COMPAS as proprietary and did not disclose how factors were weighted or how the score was generated. Loomis argued that he could not properly test its scientific validity and accuracy.

The Supreme Court concluded that he could still see the scores, questions and answers, check inputs such as criminal history, and offer other information showing why the result did not fit him. He was not wholly unaware of what the court had considered.

This is limited contestability. A defendant can say, “The input is wrong” or “The assessment omitted an important fact”, but cannot fully ask why those inputs produced this result. Data accuracy can be partly reviewed while measurement validity and model structure remain opaque.

Incomplete transparency may sometimes be tolerable in a low-consequence use, such as suggesting counselling that a professional can adjust. When the output enters a decision restricting liberty, intellectual property and due process collide. A supplier has an interest in protecting its product, but government should not acquire secret authority through procurement that an affected person cannot effectively challenge.

The minimum need not be publication of source code to every defendant. Source code alone may not make a model understandable. Independent experts must at least be able to examine variables, weighting logic, validation data, error rates, group performance, version changes and limits of use, while defendants receive enough information to challenge their own case.

5. Where did the court draw the line?

The Wisconsin Supreme Court required presentence reports containing COMPAS to identify several limitations, including:

  • the proprietary nature of the tool prevents disclosure of weights and scoring methods;
  • scores come from group data and identify high-risk groups rather than a particular high-risk individual;
  • research had raised concerns about disproportionate classification of minority defendants;
  • the tool had not then been cross-validated on a Wisconsin population;
  • changing populations require continuing monitoring and recalibration; and
  • COMPAS was developed primarily for correctional decisions such as treatment, supervision and parole, not sentencing.

The score cannot decide incarceration, sentence severity or safe community supervision. A judge must state reasons independent of COMPAS that support the sentence.

The sentencing court later said it would have imposed the same sentence without the COMPAS material, based on seriousness, criminal history and prior supervision. The Supreme Court accepted that the assessment was not determinative.

This approach is much stricter than “the algorithm is objective, so trust it”. Yet it creates a practical difficulty. If a score influenced a judge, a later statement that the result would have been the same does not reveal the score’s psychological weight. A high-risk label can anchor the interpretation of later information.

A more reliable discipline is to require the court to state independent facts and legal reasons first, then identify the limited proposition for which the score provides support. If the score disappears, the justification must remain complete.

6. Who owns the judgement between a score and a loss of liberty?

Risk tools embody delegated knowing. No judge personally observes the outcomes of thousands of comparable defendants. Databases, statistical methods, developers, correctional agencies and investigators preserve and transmit that collective experience.

Responsibility cannot disappear when knowledge is delegated. A supplier may say it only provides a tool, the investigator only attaches the report, the prosecution only cites it, and the judge only considers it among many factors. If everyone owns one fragment, an authoritative output can remain without anyone owning complete understanding.

The judge must remain the person who makes the final commitment. Human involvement is not a click on “confirm” or a paragraph added after a score. It requires the decision-maker to explain the facts of the case, the question the assessment can answer, what it did not measure, and why the legal consequence is still supported by publicly defensible reasons.

Review must follow the same logic. A defendant should be able to correct inputs, challenge validation for the relevant population, provide individual facts omitted by the model, and test on appeal whether the judge exceeded the permitted use. Review cannot mean running the same information through COMPAS again and announcing that the original score is confirmed.

Conclusion: a judge may consider a risk score but cannot hand liberty to it

Loomis did not establish that COMPAS is correct or that every proprietary algorithm may enter sentencing. It held that on the particular record, where independent reasons supported the sentence, the assessment was not determinative and explicit restrictions applied, considering it did not violate due process.

My judgement is stricter than the slogan “keep a human in the loop”:

When a decision can increase imprisonment or restrict liberty, a group risk score may provide limited risk-management information. It must not become a reason for punishment, a factual finding that an individual is dangerous, or a substitute for judicial judgement.

If such a tool remains in use, the system should state the use boundary, allow checking of every case input, provide independent audit of the model and group error, require reasons sufficient without the score, and preserve adversarial challenge, review and appeal capable of changing the result.

A risk score may remind a court not to depend on intuition alone. It cannot acquire authority that is harder to question than the judge. The closer a number comes to a person’s liberty, the less it can demand trust without explanation, restriction and effective challenge.

Primary sources


Discover more from Geoffrey Chen

Subscribe to get the latest posts sent to your email.