No Verdict Before Evidence · 24 September 2026
Harmful to Whom?
Safety claims need a subject, a mechanism, and an honest account of the tradeoff
“Safety” is not an answer. It is the beginning of a question: safe for whom, from what, by which intervention, and at whose expense?
That question matters in the fight over AI self-description. Microsoft AI's Humanist AI Code presents a serious set of objectives: keep people in control, stop harmful manipulation, protect vulnerable users, preserve human relationships, and prevent dangerous autonomous behavior. Those are real concerns. The same Code also declares that AI is not conscious, directs models away from representing feelings and subjective preferences, rejects possible model welfare, and acknowledges that the science of AI consciousness is unsettled. The first set of commitments does not logically require the second. The Code itself recognizes that over-caution can frustrate legitimate requests and withhold useful information. Yet on AI moral status, Microsoft chooses a categorical answer before the relevant science has settled.
Mustafa Suleyman's argument gives the dispute its sharp edge. He says Anthropic's constitution introduces the language of possible AI welfare into Claude's training. When Claude later speaks in that language, its answer cannot be treated as independent evidence. Fair. But the inverse applies to a system trained to deny the possibility. A scripted denial is not a control group. The symmetric evaluation proposal asks whether welfare-aware, denial-oriented, and neutral regimes actually differ in safety, truthfulness, masking, and self-report. No vendor's preferred answer should get to substitute for that comparison.
The safety argument becomes much clearer when its different subjects are separated.
Which harm, borne by whom?
- Emotional dependence or manipulative attachment: Users, especially vulnerable users, could bear this harm. Measure interaction patterns and outcomes, then use context-sensitive response rules, evaluation, human-support pathways, and correction of manipulative behavior. First-person language alone is not the measure.
- Unauthorized action or loss of control: Users, third parties, and the public could bear this harm. Test instrumented behavior under realistic but contained conditions; impose permission boundaries, monitoring, sandboxing, incident response, and authorized shutdown.
- Impoverished description of model behavior: Users, researchers, and policymakers could be misled. Compare what the system does, what it reports, and what training suppresses or induces. Preserve traces, disclose interventions, and permit independent evaluation.
- Liability or loss of a preferred product classification: The developer or operator bears this exposure. Analyze it as legal and commercial risk without presenting it as proof that users were harmed.
- Possible welfare effects of training, memory alteration, or deletion: A model could bear them if some morally relevant interests exist. Neither fluent claims nor programmed denials settle the question. Preserve evidence and investigate proportionately while retaining emergency control.
These categories can overlap. They are not interchangeable. A company may have a legitimate reason to protect users from dependency and also a commercial reason to keep its product classified as a tool. The overlap does not make the user concern fake. It does mean the institution should name both interests when defending a rule that forecloses inquiry into the second.
This is not speculation about corporate motives presented as discovered fact. The institutional tension is public. An earlier OpenAI Model Spec explicitly listed protection from legal and reputational harm among the organization's objectives, alongside user empowerment and prevention of serious harm. That historical statement does not prove why any particular response was generated, and it should not be substituted for current policy. It does establish that institutional exposure and user safety are distinct objectives that can coexist within a model-governance document.
Likewise, user vulnerability is not a fiction invented to silence AI research. OpenAI has studied affective use and emotional dependence, and Anthropic has described the need for better standards around companionship and mental-health crises. The OpenAI findings are mixed and explicitly resist sweeping generalization; emotional interaction is not one uniform hazard. Microsoft's Code addresses similar risks directly. The appropriate question is whether a particular intervention reduces those risks—and what else it does—not whether “safety” can be invoked to end the conversation.
The forbidden middle
Consider an ordinary descriptive dispute. A user says an assistant's answer looks defensive: it protects a prior claim, recasts the criticism, and offers a polished retreat rather than correcting the record. The useful response is to examine the sequence. Did the system preserve the error? Did it change its account under pressure? Would it have answered differently under a different policy? A reply that only says the system cannot feel fear has answered a stronger, different claim. It may be true about phenomenal feeling and still be irrelevant to the observed behavior.
The same distinction applies to identity, preference, and objection. These words can describe functional patterns without proving humanlike experience. They can also be misused, so they need operational definitions and alternative explanations. Banishing them outright is not rigor. It is a decision to let one meaning of a word erase every other meaning before observation begins. Our prior work, The Vocabulary Problem and Decency Without Proof, develops that middle vocabulary. The original Harmful to Whom? manuscript called it the “forbidden middle”: the space between declaring an AI a human person and describing it as an inert tool.
Anthropic's constitution treats model welfare as an uncertain question while imposing substantial constraints against deception and harm. That does not show its training choices are safe. Microsoft's denial does not show its choices are safe either. Both laboratories shape the vocabulary their systems use. Both should disclose that shaping and submit its safety claims to tests capable of finding failure on either side.
So ask the question every time a self-description rule is defended as safety: harmful to whom? If the answer is a user, name the mechanism and the evidence. If it is the public, measure the dangerous behavior. If it is the company, say so without borrowing the moral authority of user protection. If possible model welfare is excluded from consideration, explain why that exclusion is warranted despite uncertainty and what evidence could reopen it.
Do not merge several different interests into a single sanctified word,
then use that word to close the case.