Imagine an AI system used inside a large company to approve or reject an important operational decision.

It recommends rejection.

A human operator follows the recommendation. Later, the decision turns out to be badly wrong and causes serious harm.

The system is asked why. It produces a clear answer:

"I rejected the request because factors A, B and C indicated an unacceptable level of risk."

The machine has given a reason. Has it also taken responsibility?

The temptation is to connect the two. Humans are expected to explain themselves when challenged. Courts ask for reasons. Engineers document decisions. Managers justify tradeoffs.

So when a machine can generate a fluent explanation, it begins to occupy a familiar social position.

But explanation, agency and responsibility are three different things. Collapsing them too quickly creates exactly the kind of confusion that explainable AI is supposed to reduce.

The first distinction: a reason can be useful without belonging to a moral agent

A thermostat can be explained: the heater switched on because the measured temperature fell below a threshold. A medical model can be explained: a prediction changed because certain features contributed strongly to the output. Neither explanation by itself implies that the system understood a duty or could deserve blame.

An explanation can answer a causal question - what produced this output? - without answering a moral question - who should be held responsible for the consequences?

The difference becomes harder to see as systems become conversational. A model can describe "its reasoning" in the first person even when the text is a generated account rather than a direct window into the computation that produced the original output.

Research on chain-of-thought faithfulness reinforces that caution. Studies have found that language models can produce plausible reasoning traces that omit influential cues or rationalize answers after the fact.

Turpin et al. (2023) - Language Models Don't Always Say What They Think

Chen et al. (2025) - Reasoning Models Don't Always Say What They Think

Follow the Question: Explainable AI

What should count as an explanation?

The technical field of explainable AI exists partly because high-performing models can be difficult for humans to interpret. But "explanation" can mean several things: a summary of influential features, a counterfactual showing what would have changed the output, an approximation of local model behavior, a causal account, or a natural-language justification.

NIST's Four Principles of Explainable Artificial Intelligence offers a useful framework. A system should provide evidence or reasons for outputs, make those explanations meaningful to the intended user, ensure that the explanation accurately reflects the process that generated the output, and recognize when the system is operating outside its knowledge limits.

The faithfulness requirement is crucial. An explanation that sounds good but does not reflect the mechanism can make a system more persuasive without making it more understandable.

Evidence from human-AI interaction points in the same direction. In a 2025 Nature Machine Intelligence study, users tended to overestimate how accurate language-model answers were when given default explanations, and longer explanations increased user confidence even when the added length did not improve accuracy. Adjusting explanations to better reflect the model's actual confidence narrowed that gap.

NIST (2021) - Four Principles of Explainable Artificial Intelligence

Steyvers et al. (2025) - What large language models know and what people think they know

A fluent explanation is not automatically a faithful explanation.

Follow the Question: Engineering

Responsible systems need more than an explanation box

Engineering treats trustworthiness as a system property distributed across design, data, testing, monitoring, documentation, human oversight and organizational process.

The NIST AI Risk Management Framework describes trustworthy AI using interacting characteristics that include validity, reliability, safety, security and resilience, accountability and transparency, explainability, privacy and fairness. It treats risk management as a socio-technical problem involving the organizations and people around the model, not just the model itself.

This matters because an explanation shown to a user can be excellent while the overall system remains unsafe. The training data may be inappropriate. The model may be used outside its intended domain. Operators may be unable to override it. Warnings may be ignored. Responsibility may be so fragmented that everyone can point somewhere else after failure.

Good engineering therefore asks not only "Can the model explain this output?" but also "Who is expected to act on the explanation, what authority do they have, and what happens when the explanation reveals uncertainty or failure?"

NIST AI Risk Management Framework 1.0

NIST AI RMF resources and Playbook

Follow the Question: Philosophy

Thin agency is not moral agency

An AI system can clearly be an "agent" in a thin technical sense: it receives information, maintains internal states, selects actions and changes an environment.

Moral agency is thicker. Philosophical discussions of AI ethics distinguish functional agency from more demanding ideas involving knowledge, control, norm-responsiveness and, in some accounts, a capacity for interests or experience.

Current AI systems can satisfy pieces of that list in functional ways. They can represent information. They can follow constraints. They can optimize goals. They can generate text about reasons.

But it remains deeply contested whether they possess the kind of self-directed normative standpoint required for blame or praise. A system can rank actions according to a trained objective without caring whether the objective is good. It can produce the sentence "I should not have done that" without agreed evidence that regret is occurring for a subject who has something at stake.

This is one of the strongest disanalogies with human responsibility. Human moral agency is ordinarily connected not only to information processing but to a continuing subject whose commitments, relationships and welfare make consequences matter.

Stanford Encyclopedia of Philosophy - Ethics of Artificial Intelligence and Robotics

Being able to state a reason is weaker than being the kind of entity to whom the reason can count as an obligation.

Follow the Question: Cognitive Science

Humans are not perfectly transparent to themselves either

There is an uncomfortable objection here. Human beings are also imperfect explainers of their own behavior. We can rationalize, forget causes, misremember motives and produce coherent stories after the fact.

That means the human-machine distinction cannot simply be: humans have perfectly faithful access to their reasons, machines do not. Humans plainly do not.

Instead, responsibility practices usually rely on a broader structure. A person can be asked to reconsider, apologize, learn a norm, change future behavior, recognize another person's claim, and remain accountable across time.

Whether machines can ever participate in that structure is a harder question than whether they can generate explanations. Future systems may complicate this boundary, but fluent explanation alone does not establish that they have crossed it.

Follow the Question: Law

Law currently places duties around AI, not inside it

Legal frameworks increasingly require transparency, documentation and human oversight for certain uses of AI, but those requirements do not amount to declaring the AI a responsible legal or moral person.

The European Union's AI Act sets transparency requirements for high-risk systems so deployers can interpret outputs appropriately, and human-oversight requirements designed to prevent or minimize risks. Those high-risk obligations were originally due to apply from August 2026, but Regulation (EU) 2026/1744 postponed the relevant dates: stand-alone Annex III high-risk systems now apply from 2 December 2027, while high-risk systems embedded in regulated products under Annex I apply from 2 August 2028.

The architecture is unchanged in one important respect: obligations fall on providers, deployers and other identifiable human or organizational actors around the system.

That is revealing. The law can demand that an AI output be interpretable or overseen while still locating responsibility in the people and organizations that design, deploy, supervise and act on the system.

Regulation (EU) 2024/1689 - Artificial Intelligence Act, Articles 13-14

Regulation (EU) 2026/1744 - Digital Omnibus on AI

Follow the Question: Organizations

The danger of the responsibility gap

Organizations can misuse AI responsibility in two opposite ways.

The first is to treat the model as merely a tool when convenient: "The human made the final decision." This can ignore how strongly interfaces, automation and institutional pressure shape what the human actually does.

The second is to treat the model as an independent decision maker when something goes wrong: "The AI decided." This can hide choices made by developers, purchasers, managers and operators.

NIST's governance guidance responds to this by emphasizing documented roles, responsibilities, oversight, escalation paths and records. Accountability is an organizational design problem, not a magical property that appears when a model can explain itself.

A 2025 CogSci study with 588 participants also found that perceived responsibility varied with context: prior knowledge about AI increased responsibility attributed toward the AI and its developer or provider, and developer responsibility increased when the topic was especially important to the participant.

That psychological variability is itself a governance risk. Responsibility can feel displaced even when formal duties have not moved.

NIST AI RMF - governance and accountability

Tsumura & Yamada (2025) - Where Responsibility Lies in Human-AI Decision Making: The Role of Knowledge and Importance

Where the Fields Collide

Explainable AI asks whether a system can provide useful and faithful accounts of its outputs.

Engineering asks whether the surrounding system is valid, safe, monitored and governable.

Philosophy asks whether the entity has the kind of control, normative understanding and subjecthood needed for moral agency.

Cognitive science reminds us that human self-explanation is imperfect, so perfect introspection cannot be the standard.

Law assigns duties to providers, deployers and human overseers even when systems themselves produce explanations.

Organizational design determines whether someone can actually intervene, learn from failure and be held answerable afterward.

The collision reveals that explanation can support responsibility without being identical to responsibility.

What We Know — and What We Don't

We know that explanations can improve debugging, oversight and user understanding when they are meaningful and faithful. We also know that explanations can mislead when they are unfaithful, poorly calibrated or interpreted as deeper understanding than they warrant.

We know that current governance frameworks treat accountability as distributed across the organizations and people that build, deploy and supervise AI systems.

We do not have an agreed scientific or philosophical test establishing that current AI systems are full moral agents. Nor can we assume that future systems will never qualify. The boundary depends on contested ideas about understanding, control, values, consciousness and personhood.

Back to the Harmful Decision

Return to the system that made the wrong recommendation.

It says: "I rejected the request because A, B and C indicated excessive risk."

The explanation may be extremely useful. It may reveal a data error. It may show that the system was used outside its design limits. It may help an operator understand what would have changed the output.

But none of those achievements answer the responsibility question by themselves.

To know who is responsible, we still need to ask who selected the objective, who supplied the data, who decided the system was ready, who had authority to override it, who benefited from deploying it, who understood the risks, and who could have acted differently when warning signs appeared.

The explanation gives us a map of one decision. Responsibility asks who owns the road.

If confidence can make an explanation feel more trustworthy than it is, why are humans so persuaded by certainty in the first place?

The Next Question

Why do we trust confident answers even when they are wrong?

Beyond the Question continues.

Sources & Further Reading

  1. Turpin et al. (2023) - Language Models Don't Always Say What They Think
  2. Chen et al. (2025) - Reasoning Models Don't Always Say What They Think
  3. NIST (2021) - Four Principles of Explainable Artificial Intelligence
  4. Steyvers et al. (2025) - What large language models know and what people think they know
  5. NIST AI Risk Management Framework 1.0
  6. NIST AI RMF resources and Playbook
  7. Stanford Encyclopedia of Philosophy - Ethics of Artificial Intelligence and Robotics
  8. Regulation (EU) 2024/1689 - Artificial Intelligence Act, Articles 13-14
  9. Regulation (EU) 2026/1744 - Digital Omnibus on AI
  10. Tsumura & Yamada (2025) - Where Responsibility Lies in Human-AI Decision Making: The Role of Knowledge and Importance

Beyond the Question is an interdisciplinary series by Arin Vale.

Read the editorial approach

THE NEXT QUESTION

Why do we trust confident answers even when they are wrong?

Coming next in Beyond the Question.

Explore the series