Imagine you are filling out a form. It asks for your address. You leave it blank. It asks where you work. Blank again.
Religion?
Blank.
Health conditions?
Blank.
Political beliefs?
Blank. You close the form feeling reasonably private. You did not give away the sensitive information. Now imagine the company behind the form has something else.
A few location points. The times you usually leave home. A list of purchases. The songs you play late at night. The devices you use. Which links you click.
Who you communicate with most often. The places your phone regularly pauses. No single item looks like the fact you refused to provide.
But what if the fragments are enough?
If someone can infer what you never told them, was the information ever really private?
Follow the Question: Data Science
Ordinary points become unusual in combination
One location coordinate says very little. A second may say more. A third can begin to form a pattern.
In 2013, Yves-Alexandre de Montjoye and colleagues analyzed a dataset containing mobility information from 1.5 million people. In that dataset, four randomly selected spatiotemporal points were enough to uniquely characterize 95 percent of individual mobility traces.
The careful wording matters.
That does not mean four location points will identify 95 percent of every population in every database.
It means that in the dataset the researchers studied, human mobility patterns were strikingly unique.
The result is important because it challenges a simple idea of anonymity. Remove the name. Remove the email address. Remove the phone number. The pattern may still be distinctive.
Data does not need to contain your identity in one field if the combination of fields behaves like a fingerprint.
The same principle appears beyond location. A set of harmless-looking observations can become identifying when joined together. Time. Sequence. Frequency.
Regularity. Association. The fragments acquire meaning from their relationship. This is why privacy is not merely about individual data points. It is also about what becomes possible when data points are combined.
de Montjoye, Y.-A., et al. (2013). Unique in the Crowd: The privacy bounds of human mobility.
Follow the Question: Cybersecurity
Removing a name is not the same as removing identifiability
Anonymization sounds like a binary condition. Either the data identifies someone or it does not. Reality is less tidy.
A dataset may remove direct identifiers and still remain vulnerable to re-identification when combined with outside information.
Consider a table containing only:
• age range;
• broad location;
• purchase dates;
• and travel patterns.
No names. No email addresses. No account numbers.
Now add another dataset containing public event attendance, social posts, or workplace schedules.
The first table may become more revealing because the second one exists. Privacy therefore depends partly on the surrounding information environment.
A dataset that was difficult to identify ten years ago may become easier to identify later if more auxiliary data becomes available.
That changes the security question.
It is not enough to ask:
Did we remove the obvious identifiers?
We also have to ask:
What could someone infer if they joined this with other information?
The answer can change over time.
And unlike a password, a mobility pattern or a social network is not something a person can simply rotate after a breach.
Follow the Question: Machine Learning
Attribute inference turns absence into a prediction problem
A classic example arrived before today's generative-AI boom.
In 2013, Michal Kosinski, David Stillwell, and Thore Graepel analyzed Facebook Likes from more than 58,000 volunteers and showed that relatively ordinary digital traces could predict a range of personal attributes, including political and religious views, personality traits, age, gender, and other sensitive characteristics.
The point was not that one Like exposes one secret. It was that many weak signals can become a strong inference system.
NIST uses the term attribute inference attack for attempts to infer sensitive attributes about a record using partial knowledge.
The language is useful because it reveals the shift. The attacker does not necessarily steal the hidden field. The attacker predicts it.
Suppose a model is trained on people for whom both ordinary behavior and a sensitive attribute are known.
It may learn relationships between the two. Then, for a new person whose sensitive attribute is hidden, the model can produce an estimate. The estimate may be wrong.
That matters. Inference is not omniscience. A model can produce false positives, false negatives, uncertainty, and systematic errors. But privacy risk does not require perfect inference.
A business may still act on a probability. An advertiser may still change what it shows you.
An insurer, employer, platform, or political campaign may treat a prediction as useful enough to alter behavior even when the hidden fact is not known with certainty.
That creates a strange situation.
You can truthfully say:
“I never told them that.”
And they can truthfully say:
“We did not need you to.”
Follow the Question: Business
Prediction can be more valuable than possession
Companies do not always need to know who you are in a philosophical sense. They need to know what you are likely to do.
Will you click?
Will you buy?
Will you cancel?
Will you repay?
Will you respond to a discount?
Will you watch another video?
Will you leave after ten seconds or stay for an hour?
From a commercial perspective, prediction can sometimes be more useful than biography.
A company may not care about the full story of your life if a small set of signals predicts the next decision well enough.
That changes the value of seemingly minor data. A single click may be trivial. A pattern of clicks can estimate interest. A sequence of interests can estimate intent.
Intent can change price, ranking, advertising, timing, and offers. This does not mean every company is secretly inferring every sensitive fact it can.
It means the economic incentive has changed. Data is valuable not only for what it directly says. It is valuable for what it makes predictable.
And once prediction becomes the product, the boundary of “personal information” becomes harder to draw around only the fields a person explicitly entered.
Follow the Question: Psychology
We leak patterns without meaning to communicate
Human behavior is repetitive. We wake at similar times. Travel familiar routes. Buy recurring items. Call the same people. Read the same kinds of stories.
Visit places connected to work, family, health, worship, recreation, and habit. Most of these actions are not experienced as disclosures.
You do not walk into a pharmacy thinking:
I am publishing a behavioral feature. You are buying something.
You do not drive to a friend's house thinking:
I am updating a social graph. You are visiting a friend. The data layer translates lived behavior into variables after the fact. That creates an asymmetry between intention and observability.
A person may understand the obvious disclosure — the email address typed into a box — while having almost no intuitive sense of what hundreds of ordinary traces reveal in combination.
Privacy therefore becomes partly a problem of legibility. We are readable in ways we did not evolve to perceive.
Follow the Question: Law
Regulation is beginning to notice inference
European privacy law gives us one useful language for this problem.
The General Data Protection Regulation defines profiling in terms that include automated processing used to evaluate or predict aspects such as work performance, economic situation, health, personal preferences, interests, reliability, behavior, location, or movements.
The important word is predict. Privacy law is not limited to literal fields handed over directly.
The European Data Protection Board's final version of Guidelines 3/2025, adopted in September 2026, also recognizes that special-category information can be derived or inferred through profiling. Its example describes inferred religious beliefs from geolocation — such as visits to places of worship — or shopping habits such as purchases of particular food products.
That matters because it rejects the idea that inferred data is somehow harmless simply because nobody asked the person the question directly.
But law faces a difficult problem.
What exactly should be regulated?
The raw data?
The inference?
The decision based on the inference?
The accuracy of the inference?
The purpose for which it is used?
Suppose a model infers a sensitive trait incorrectly. Privacy harm can still occur if the false inference changes how the person is treated. So the legal problem is not only about truth.
It is about power attached to prediction.
European Union. General Data Protection Regulation, Article 4(4) — profiling.
A deeper problem: privacy can fail without a leak
Traditional privacy stories have a clear villain. A database is breached. A document is stolen. A secret is exposed. Inference is different. Nothing has to leak.
The system may work exactly as designed. A platform collects ordinary behavioral data under its normal rules. A model identifies patterns. A sensitive attribute is estimated.
No hacker enters the system. No file is stolen. No employee exports a secret spreadsheet. And yet something the person tried not to disclose has become operationally visible.
That is a different kind of privacy failure. The system did not lose the information. It created an estimate of it.
The asymmetry of inference
There is another problem.
The person making the inference may know much more about the process than the person being inferred about.
A company can test models across millions of records. An individual sees only their own experience. The company may know which variables are predictive.
The individual may not even know which variables were observed. The company can measure uncertainty. The individual may only see the consequence.
A different advertisement. A different ranking. A different price. A rejected application. A flagged account. The prediction can shape the world around a person without announcing itself as a prediction.
That makes contestability important.
If a system acts on an inferred trait, should the person be told?
Should they be allowed to correct it?
What if correcting it requires revealing the very private fact they refused to provide in the first place?
The attempt to protect privacy can become the condition for proving that privacy was violated.
What can actually reduce the risk?
Privacy needs limits before prediction begins
Diagnosis is not enough.
If inference is the problem, one response is to reduce the amount of material available for inference in the first place.
The GDPR's principles of purpose limitation and data minimisation point in that direction: collect data for specified purposes and limit collection to what is necessary for those purposes.
That does not make inference impossible. It narrows the raw material.
Technical approaches can help too. NIST's 2025 guidance on differential privacy describes a mathematical framework for quantifying privacy loss when information from a dataset is released or analyzed. Differential privacy is not a magic shield, and poor implementations can make strong-sounding promises without delivering strong protection. But it illustrates a different philosophy: do not merely promise that names are removed. Design the analysis so that the contribution of any one person is harder to expose.
Other protections are institutional rather than mathematical. Limit secondary uses. Delete data that no longer serves its purpose. Make consequential inferences contestable.
Tell people when inferred attributes materially affect decisions. Separate what can be predicted from what is permitted to be used.
None of these measures restores a world in which privacy means simply keeping a secret in a drawer.
They recognize the new problem:
The risk begins before the sensitive fact is ever explicitly collected.
European Union. GDPR Article 5 — purpose limitation and data minimisation.
NIST SP 800-226 (2025). Guidelines for Evaluating Differential Privacy Guarantees.
Where the Fields Collide
Data science shows that combinations of ordinary traces can become highly distinctive.
Cybersecurity shows that removing explicit identifiers does not always remove re-identification risk.
Machine learning turns hidden attributes into prediction targets. Business gives prediction economic value even when perfect knowledge is unnecessary.
Psychology explains why ordinary behavior produces patterns people do not experience as deliberate disclosure.
Law is beginning to recognize that profiling and inference can implicate privacy even when sensitive data was not directly typed into a form.
The fields meet at one change in the meaning of privacy:
That is a harder problem because inference cannot be solved simply by saying less. Sometimes the pattern survives.
Privacy is no longer only about controlling what others receive from you. It is also about what they can derive from what they already have.
What We Know — and What We Don't
We know human mobility can be highly distinctive in some datasets. We know attribute inference is a recognized privacy and security risk.
We know modern profiling systems can use behavioral data to predict preferences, interests, behavior, location, and other attributes.
We know European privacy frameworks explicitly address prediction and profiling.
We do not know that every sensitive attribute can be inferred accurately from a handful of data points.
We do not know that every company collecting behavioral data is making every possible inference.
And we should not turn striking research findings into universal claims.
Four spatiotemporal points uniquely characterizing 95 percent of traces in one dataset does not mean four points identify 95 percent of all people everywhere.
Inference is powerful. It is also probabilistic, context-dependent, and fallible. That is exactly why the problem is difficult. Privacy harm can arise from an inference that is accurate.
It can also arise from an inference that is wrong.
Back to the Form
Return to the form. Address: blank. Workplace: blank. Religion: blank. Health: blank. Political beliefs: blank. You did not disclose the answers.
That fact still matters. But it no longer completes the privacy story. The next question is not only what you gave away. It is what the surrounding system could learn from everything else.
Privacy used to feel like a door. Open or closed. Prediction makes it look more like a shadow. You can hide the object and still reveal its shape.
The Next Question
Prediction teaches us that value can be created by context around information, not only by the information itself.
Scarcity reveals a parallel problem with objects. Sometimes nothing about the object changes. Only the future around it does. Availability narrows.
Delay becomes costly. And suddenly the same thing feels more valuable.
Why does scarcity change what we think something is worth?
Sources & Further Reading
- de Montjoye, Y.-A., et al. (2013). Unique in the Crowd: The privacy bounds of human mobility.
- NIST. Attribute inference attacks — Glossary.
- Kosinski, M., Stillwell, D., & Graepel, T. (2013). Private traits and attributes are predictable from digital records of human behavior.
- European Union. General Data Protection Regulation, Article 4(4) — profiling.
- European Data Protection Board. Guidelines 3/2025 on the interplay between the DSA and the GDPR — final version (17 September 2026).
- NIST SP 800-226 (2025). Guidelines for Evaluating Differential Privacy Guarantees.
Beyond the Question is an interdisciplinary series by Arin Vale.
Read the editorial approachTHE NEXT QUESTION