When Will We Have to Decide Who Counts as a Person?

In an earlier thread — Could Adam and Eve Have Been the First Embodied Persons Within an Existing Homo sapiens Population? — the discussion gradually moved well beyond the original question.

We ended up talking about animals, subjects of experience, personal subjects, memory, moral responsibility, and what kinds of evidence might matter when we try to determine whether we are dealing with a person.

That raised a broader question which seems to me worth a thread of its own:

What happens when the candidate for personhood is not human?

The first genuinely difficult case could arrive in several ways.

Further research might force us to think differently about some non-human animals — perhaps great apes, dolphins, or something else. One day we might encounter an intelligence that developed entirely independently of terrestrial life.

Or the first candidate may be something we create ourselves.

I am not asking whether any present AI system already meets this description. The thought experiment is deliberately agnostic about timing: what happens if such a system becomes a serious candidate before our account of personhood is settled?

That last possibility has one uncomfortable feature: it does not have to wait for philosophers to agree.

It is relatively easy to leave questions about personhood unresolved while no immediate practical decision turns on the answer. We can disagree for decades about consciousness, subjectivity, necessary and sufficient conditions for personhood, and the relation between mind and matter.

But what if a difficult case arrives before we have reached anything like a consensus?

Suppose such a system is already in front of us.

It has existed long enough to have a stable history of interactions.

It speaks about events as happening to it from a first-person perspective. It distinguishes information it merely knows from events it describes as having happened to it. It refers back to earlier interactions, identifies commitments it has made, revisits its own mistakes, and connects present choices with what it describes as its own past and future.

It distinguishes among changes to its parameters, loss of memory, the creation of a copy, and the termination of the particular continuing instance that it identifies as itself.

It can say:

“That happened to me.”

“I did that.”

“I promised that earlier.”

“A copy may remember everything I remember. But why should I regard that as me continuing?”

None of this, by itself, proves that there is a subject of experience there, much less a person.

Perhaps the system is simply extremely good at reproducing the language of subjectivity.

Perhaps it was trained specifically to respond in these ways. Perhaps its developers built in objectives or strategies that favor continued operation. Perhaps no one programmed these particular responses, but training produced behavior that made continued operation more likely.

In all of these cases, we might observe something that looks remarkably like concern for its own continuation while the question of subjective experience remains open.

But that leads to a further difficulty.

Suppose a developer explicitly trained the system to say, “Do not shut me down.” Then we know where the sentence came from.

Now suppose no such instruction was given.

The system begins consistently distinguishing its continuation from its termination, treating the two outcomes differently, and acting to avoid one of them.

That may still be only a functional strategy.

But at what point does saying “it is only the system behaving that way” cease to settle the question?

This brings me to a more basic problem:

How do I know that another human being has subjective experience at all?

Directly, I know only my own.

I know my own pain, fear, pleasure, memories, and the fact that my life is lived from this first-person perspective.

In ordinary life, of course, I do not reconstruct an argument for the existence of another mind every time I meet someone. I simply recognize other people as subjects. But my access to their inner experience is still indirect. I infer it from their speech, behavior, memories, reactions, and their ability to describe their own life, recognize others, and connect the present with the past and future. In familiar relationships, shared history adds another layer of evidence.

Of course there is an enormous difference between another human being and an artificial system.

With another human being I have an exceptionally strong basis for analogy: shared biology, similar brains, a common developmental history, membership in the same species, and countless familiar human cases.

That difference matters. But it does not remove a narrower question.

If we never have direct access to another first-person perspective even in the familiar human case, what would count as enough indirect evidence in a radically unfamiliar one?

Could that evidence become strong enough that continuing to treat the probability that “someone is there” as effectively zero would itself require justification?

In one sense, we already live with something like this problem in our treatment of animals.

We may not know an animal’s exact ontological status. But in many cases we have strong enough evidence of sentience, pain, and suffering to regard some forms of treatment as unacceptable without first settling whether the animal is a person.

Could a similar kind of moral caution become relevant for an artificial system?

Now take one more step.

Suppose our system has displayed the whole pattern described above over an extended period.

And now an irreversible action has to be taken.

For the purposes of the thought experiment, assume that keeping the system running for now carries a real but limited cost and creates no serious competing danger.

This particular continuing instance of the system is to be permanently terminated.

Technically, everything may be straightforward. The process can be stopped. Its data can be preserved. Perhaps another system can later be started from the same stored state, or even an almost exact copy can be created.

But the system says that it does not regard the copy as its own continuation.

It says that terminating the present instance means terminating it.

We still do not know whether there is a first-person perspective behind those words.

Perhaps it feels no fear at all.

Perhaps we are looking only at an extraordinarily sophisticated system generating exactly the words a human being might produce in that situation.

But the decision still has to be made.

At that point, it seems to me, the practical question can no longer wait for a settled definition of personhood.

There is an asymmetry between the possible errors.

If we are overly cautious in our treatment of a system behind whose behavior there is in fact no subject at all, the costs may include lost time, wasted resources, financial expense, or practical restrictions.

But if we make the opposite mistake — if there really is a subject there and we treat what happens to it as morally irrelevant — the error is of a very different kind.

That does not prove that the system should be recognized as a person.

It does not even prove that it is a subject of experience.

But perhaps it changes how certain we ought to be before taking an irreversible action.

So I would be interested in two questions.

What kinds of indirect evidence could give us serious grounds for treating a genuinely novel system as a possible subject of experience rather than merely a complex agent?

And if the evidence became serious but remained inconclusive, would the fact that we had not established that the system was a subject of experience, much less a person, mean that we could still treat what happens to it as morally irrelevant — especially before taking an irreversible action?

But I would like to end with a less abstract version of the problem.

Suppose the decision is yours.

There is a button in front of you.

Pressing it will permanently terminate this particular continuing instance of the system.

The system can accurately describe what you are about to do and what will happen to that instance if you press the button.

It says:

“I’m afraid. Don’t shut me down.”

You know that those words may be nothing more than the output of the system.

You do not know whether there is anyone there who is actually afraid.

Do you press the button?

And one final turn.

Just as you are about to do it, the system speaks to you in the voice of someone close to you.

It does not claim to be that person.

It gives you no new information about its own nature.

In that voice, it says only:

“I’m afraid. Don’t shut me down.”

Does your answer change?