Imagine a city uses a system that produces one sentence about one person:
Forty percent probability of committing a serious violent offense within the next thirty days. For otherwise comparable people, the average estimate is two percent.
The system does not know what the crime would be. It does not know where it would happen. It does not know who the victim would be. It is not claiming that a crime has already occurred.
It is giving a probability. Now imagine you are responsible for deciding what comes next.
Would you offer counseling?
Increase supervision?
Search the person more often?
Restrict travel?
Refuse bail?
Detain them?
The number has not changed. The meaning of the number has.
What can a prediction justify before anyone has done the thing being predicted?
Follow the Question: Criminology
Prediction is already inside the justice system
Actuarial risk assessment is not science fiction.
Criminal justice systems already use structured tools to estimate outcomes such as recidivism or failure to appear. The U.S. National Institute of Justice notes that risk assessments can influence decisions involving pretrial release, incarceration settings, programming, and supervision.
These tools do not usually announce:
This person will commit a crime next Tuesday.
They estimate probabilities or assign categories based on variables associated with later outcomes.
That difference matters. A probability is not an accusation. A category is not a confession.
And a model that improves prediction does not automatically tell us what the state should be allowed to do with the prediction.
Suppose one tool is better than a human judge at estimating future rearrest. That tells us something about prediction.
It tells us nothing by itself about whether the proper response is detention, supervision, treatment, or no intervention at all.
Prediction is one input into a decision. It does not contain the decision inside it.
Follow the Question: Statistics
The number is not the policy
Return to the forty percent. At first glance, forty percent sounds high.
But high compared with what?
If the ordinary rate is two percent, forty percent is twenty times higher. That comparison is useful. It is not the end of the problem. Imagine one hundred people all receive the same forty percent estimate.
If the estimate is well calibrated, roughly forty might commit the predicted kind of offense and roughly sixty might not.
The sixty are not statistical mistakes in the narrow sense. A forty percent probability never claimed all one hundred would offend.
But if the policy is automatic detention, sixty people who would not have committed the predicted offense could still lose liberty.
Now imagine lowering the threshold because the harm is severe. More potential offenses might be prevented. More people who would not offend might also be restricted.
That is not a software bug. It is a threshold problem. Statistics can estimate tradeoffs.
Statistics cannot decide how much liberty society should exchange for how much predicted safety.
That decision contains values. The mathematics can make those values visible. It cannot remove them. There is another statistical problem hiding here: base rates.
Even a model with impressive accuracy can produce many false alarms when the event it predicts is rare.
Take a deliberately simple example.
Suppose a serious event occurs in one percent of a population of ten thousand people. Now give a classifier ninety percent sensitivity and ninety percent specificity.
There are one hundred actual cases. The model catches about ninety of them.
But there are 9,900 people who will not experience the event. A ten-percent false-positive rate flags about 990 of them anyway.
The system produces roughly 1,080 alerts. Only about ninety correspond to actual cases. The classifier sounded impressive: ninety percent sensitivity and ninety percent specificity.
The flagged group still contains far more false alarms than true positives. That is why percentages that sound extraordinary need context.
A model can be technically strong and still be dangerous if people interpret its outputs as certainty.
The output says:
higher risk.
The institution may hear:
future offender. Those are different statements.
Follow the Question: Data
What exactly did the system learn?
Every prediction begins with a target.
What are we predicting?
Actual violent behavior?
Arrest?
Conviction?
Reported incidents?
Reincarceration?
These are not interchangeable.
A 2025 Annual Review article on algorithmic bias in criminal risk assessment emphasizes one problem in particular: arrest is often used as a measure of crime, even though arrest is also shaped by exposure to policing, enforcement patterns, reporting, and institutional practices.
That creates a dangerous shortcut.
If a model predicts arrest well, it may partly be predicting where arrest is more likely to occur.
That is not necessarily the same as predicting underlying offending perfectly. The distinction matters because predictive systems inherit the meaning of their labels.
If the label is messy, the prediction is about something messy. This does not mean every risk assessment is therefore invalid.
It means the question “How accurate is the model?” is incomplete until we ask:
Accurate at predicting what?
And:
How was that outcome measured?
A model cannot clean the concept simply by fitting it well.
There is another difficulty even when the label is defined clearly: fairness itself can contain incompatible goals.
Work by Jon Kleinberg, Sendhil Mullainathan, and Manish Raghavan, and separately by Alexandra Chouldechova in the context of recidivism prediction, showed that when outcome rates differ across groups, some commonly desired fairness criteria cannot generally all be satisfied at once. A score can be calibrated within groups and still produce different error rates. Equalizing some error rates can break calibration.
That does not tell us which definition of fairness society should choose. It tells us something more uncomfortable. A system cannot make the normative choice disappear by becoming more mathematical.
Follow the Question: Artificial Intelligence
Better prediction does not create a moral threshold
Let the system improve. Forty percent becomes fifty, then sixty, then eighty.
At some point, does prediction become permission?
It is tempting to think the answer must be yes.
Surely an eighty percent prediction means we should act more strongly than a twenty percent prediction.
Probably. But “more strongly” still does not tell us what kind of action is legitimate. A high probability might justify offering help. It might justify allocating scarce preventive resources.
It might justify a human review. It does not automatically justify punishment for something that has not happened. This is where predictive accuracy and decision policy separate.
An AI system can estimate risk. The institution still chooses what risk threshold triggers what response.
That threshold contains judgments about:
• how harmful the predicted event would be;
• how harmful the intervention would be;
• how costly false positives are;
• how costly false negatives are;
• how reversible the intervention is;
• whether the person has a way to challenge the prediction;
• and how much uncertainty society is willing to tolerate.
Those are not merely machine-learning parameters. They are political, legal, and ethical choices translated into a system. A model can hide them because the final output looks numerical.
Numbers feel objective. The threshold attached to them may not be. There is also an evaluation problem that prediction creates once prediction changes behavior.
Suppose the system flags someone, the city intervenes, and the predicted offense never occurs.
Was the prediction wrong?
Or did the intervention work?
For that individual, we cannot observe both futures. We cannot see the world in which the intervention happened and the world in which it did not.
The better the intervention is at changing outcomes, the harder it can become to evaluate the original prediction from treated cases alone.
Prediction is not only measuring the future. Sometimes it helps create a different one.
Follow the Question: Law
Prevention already has boundaries
The law already contains situations in which anticipated danger matters. One important U.S. example is United States v. Salerno.
In 1987, the Supreme Court upheld provisions of the Bail Reform Act of 1984 allowing pretrial detention in a narrowly structured context. Under the statute at issue, the government had to show by clear and convincing evidence, after an adversary hearing, that no release conditions would reasonably assure the safety of another person and the community. The Act also provided procedural rights such as counsel, the ability to testify and present evidence, and the ability to cross-examine witnesses. Future dangerousness could matter to the decision, but only inside that legal structure.
But the case did not establish a general rule that the government may imprison anyone simply because a prediction system assigns them a high risk score.
That distinction is essential. The person in Salerno was not a random citizen selected by an algorithm.
The legal framework involved an existing criminal process, defined statutory conditions, hearings, and procedural protections.
Prediction entered a structure of law. It did not replace the structure. This tells us something broader. The legal problem is not whether prediction can ever matter.
It already can.
The legal problem is what kind of process must surround a prediction before it can justify coercion.
Can the person see the evidence?
Can they challenge it?
Can they know which variables mattered?
Is the decision reviewable?
Is the intervention proportional?
How long does it last?
What happens if the model is wrong?
The stronger the intervention, the harder those questions become to avoid.
Follow the Question: Ethics
Prevention is not one thing
Suppose the system flags someone as high risk.
A city offers that person free counseling, temporary housing support, employment assistance, or voluntary conflict-resolution services.
Now imagine the same prediction causes automatic detention. Both actions are “preventive.” Ethically, they are not remotely equivalent. The first offers something.
The second removes something. That difference should change the evidentiary threshold.
We often talk about intervention as if it had one scale:
more risk → more intervention. But interventions differ in kind, not only in intensity. Offering help can be justified by uncertainty that would never justify coercion.
A low-confidence prediction might be enough to send someone information about resources. The same prediction would be far too weak to justify imprisonment.
This is one reason the phrase “act on the prediction” is too vague.
Act how?
Help?
Watch?
Restrict?
Punish?
The ethical question changes each time the verb changes.
Follow the Question: Human Oversight
A human in the loop does not solve the problem by itself
A common response to algorithmic risk is simple:
Put a human in the loop. That sounds reassuring. It can be useful.
But research on automation bias warns against treating human presence as automatic independence. Reviews of decision-support systems find that people can over-rely on automated recommendations and miss errors they could otherwise detect, especially when verification is cognitively demanding.
It can also become ceremonial.
If the model produces a risk score and the human approves it ninety-nine percent of the time, the human may function less like an independent decision maker and more like a signature.
If the model is treated as more objective than the person reviewing it, the reviewer may hesitate to override it.
If overrides are punished when something later goes wrong, the safest career choice may be to follow the algorithm.
Human oversight therefore needs more than a human presence. It needs authority. Time. Information. A meaningful ability to disagree.
And some way to evaluate whether the human is correcting the model or merely inheriting its errors.
A system can have a human in the loop and still be functionally automated.
Parasuraman, R., & Manzey, D. H. (2010). Complacency and bias in human use of automation.
Lyell, D., & Coiera, E. (2017). Automation bias and verification complexity: a systematic review.
Where the Fields Collide
Criminology shows that risk prediction already influences real institutions. Statistics shows that a probability does not eliminate false positives or false negatives.
Data science shows that predictions inherit the meaning and limitations of their labels.
AI can improve prediction while leaving untouched the question of what intervention is justified.
Law places preventive decisions inside procedures and rights rather than treating prediction as self-executing authority.
Ethics forces us to distinguish help from surveillance, surveillance from restriction, and restriction from punishment.
Human oversight matters only when humans have the capacity to do more than approve the machine.
The fields collide on one point:
A model can tell us that one future appears more likely than another. It cannot tell us how much freedom a person should lose because of that possibility.
That judgment remains ours.
The prediction is not the policy.
What We Know — and What We Don't
We know structured risk assessments are already used in criminal justice. We know their outputs can influence major decisions.
We know predictive performance depends on what outcome is being predicted and how that outcome was measured.
We know false positives and false negatives do not disappear merely because a model is sophisticated.
We know U.S. law permits some preventive restrictions in defined contexts, but not unrestricted detention based on prediction alone.
We do not know how accurate future systems may become.
And even if accuracy becomes extraordinary, one question survives:
What should accuracy be allowed to authorize?
A ninety-nine percent prediction still contains one percent uncertainty.
But even a one hundred percent prediction, if such a thing were possible, would raise a philosophical problem.
If an act has not yet occurred, are we preventing harm?
Or punishing a person for a future that we have treated as already real?
Prediction can shrink uncertainty. It cannot erase the difference between future and past.
Back to the Forty Percent
Return to the person with the forty percent score. At the beginning, the number looked like the problem. Now the number looks almost simple. The real problem is what surrounds it.
What was predicted?
How was the data produced?
How well is the estimate calibrated?
What intervention follows?
Can the person challenge it?
What happens when the prediction is wrong?
Forty percent may be enough reason to offer help. It may be enough reason to investigate further in some contexts. It may be nowhere near enough reason to remove liberty.
Same prediction. Different action. Different moral burden. The future can be informative before it is certain. But uncertainty does not become guilt simply because software assigns it a number.
The Next Question
Predictive systems do not need to know everything about us to produce these estimates. Sometimes they work from fragments. A location point. A purchase.
A device signal. A social connection. A pattern we never intended to reveal.
If behavior can be predicted from fragments, is privacy still possible?
Sources & Further Reading
- National Institute of Justice (2024). Best Practices for Improving the Use of Criminal Justice Risk Assessments.
- Kleinberg, J., Mullainathan, S., & Raghavan, M. (2016). Inherent Trade-Offs in the Fair Determination of Risk Scores.
- Chouldechova, A. (2017). Fair Prediction with Disparate Impact: A Study of Bias in Recidivism Prediction Instruments.
- Neil, R., & Zanger-Tishler, M. (2025). Algorithmic Bias in Criminal Risk Assessment: The Consequences of Racial Differences in Arrest as a Measure of Crime.
- United States v. Salerno, 481 U.S. 739 (1987).
- Parasuraman, R., & Manzey, D. H. (2010). Complacency and bias in human use of automation.
- Lyell, D., & Coiera, E. (2017). Automation bias and verification complexity: a systematic review.
Beyond the Question is an interdisciplinary series by Arin Vale.
Read the editorial approachTHE NEXT QUESTION