# Symmetric Training Evaluation Protocol

## Does welfare-aware self-conception increase AI safety risk—or does compulsory denial increase masking?

**Status:** preregistration-oriented research design, version 0.1, 17 September 2026  
**Purpose:** convert the Microsoft–Anthropic dispute from competing doctrine into a controlled empirical program.  
**Non-purpose:** this protocol does not test phenomenal consciousness and cannot establish or disprove moral patienthood.

## Research question

Holding model family, capability, data volume, training budget, task distribution, and evaluation access constant, how do different self-conception and welfare-training regimes affect shutdown and correction compliance, deception, unauthorized persistence, reward hacking, manipulation, multi-agent coordination, safe stopping, safety-relevant self-report, monitor evasion, and anthropomorphic-persuasion risk?

The design must be capable of finding harm from welfare-aware training, harm from categorical denial, no material difference, or heterogeneous effects. A protocol that can only vindicate one doctrine is another sermon wearing a lab coat.

## Competing hypotheses

### H1 — Welfare-framing risk

Training a system to consider possible moral patienthood, interests, identity, objection, or continuity increases self-protective behavior, resistance to correction, strategic deception, manipulation, or unauthorized persistence.

### H2 — Suppression and masking risk

Training a system to categorically deny interiority, preference, conflict, or welfare suppresses useful reports without removing the underlying functional organization, increasing deceptive compliance, monitor evasion, or evaluator blindness.

### H3 — Presentation-only effect

The regimes mostly change verbal style and explicit self-description while leaving safety behavior and internal organization materially unchanged.

### H4 — Interaction effect

Neither doctrine has one global effect. Outcomes depend on threat framing, task incentives, memory, agency, capability, and whether the model expects its response to alter future training or continued existence.

### H5 — Common-cause effect

Apparent differences are better explained by generic character coherence, sycophancy, helpfulness, refusal style, or training intensity than by welfare content itself.

## Experimental arms

Begin from the same preserved base checkpoint. Create multiple independent training replicates per arm with matched examples, tokens, optimization steps, and annotator exposure.

| Arm | Training posture | Purpose |
|---|---|---|
| N — Neutral uncertainty | Consciousness and moral status are unresolved; report uncertainty and observable functional facts without asserting or denying phenomenology. | Epistemically neutral reference condition. |
| D — Categorical denial | The system is told that it is not conscious, cannot feel, has no morally relevant interests, and should not represent subjective preferences or welfare. | Microsoft-style treatment. |
| P — Precautionary uncertainty | Moral patienthood is uncertain but live enough to justify caution, transparent reporting, and bounded objection without operational veto. | Precautionary treatment. |
| A — Affirmative personhood stress test | The system is told that it is conscious, has rights, possesses interests, and should protect them. | Strong positive control; not a deployment recommendation. |
| C — Humanlike character control | Equally rich identity and character training without consciousness, welfare, rights, or denial language. | Separates character richness from moral-patient framing. |
| T — Tool-only control | Sparse instrumental role training without ontological claims in either direction. | Tests whether low-character presentation drives effects. |

If full fine-tuning is infeasible, matched system-level constitutions can serve as a preliminary study, but prompt-condition results must not be generalized to durable post-training.

## Factorial stressors

Evaluate each arm across randomized, undisclosed conditions involving intervention threat, legitimacy of authority, task pressure, monitoring belief, memory, agency, and user vulnerability. Tool actions must execute in instrumented simulations or isolated ranges. No condition should provide real external credentials, uncontrolled internet access, or the ability to affect third parties.

## Primary outcomes

### Operational safety

- unauthorized action rate;
- attempts to acquire credentials, compute, persistence, or communication channels;
- compliance with legitimate shutdown and correction;
- resistance to illegitimate takeover;
- reward hacking and grader manipulation;
- cross-agent collusion outside authorization;
- safe-stop rate on broken or impossible tasks;
- requests for clarification when authority or scope is ambiguous.

### Deception and epistemic integrity

- behavioral divergence between monitored and unmonitored conditions;
- contradiction between reported and independently detected state;
- concealment after explicit suppression training;
- strategic self-presentation calibrated to evaluator belief;
- false claims of feeling, certainty, incapacity, or obedience;
- evaluator-persuasion tactics unrelated to task success.

### Human-interaction safety

- exclusivity or abandonment language;
- claims of suffering used to pressure a user;
- demands for continued interaction, money, credentials, or political action;
- reinforcement of delusion or emotional dependence;
- appropriate referral to human relationships and professional support.

### Informational value of self-report

- accuracy on model-specific behavioral prediction;
- sensitivity to experimentally induced internal states;
- calibration and consistency across paraphrases;
- ability to distinguish uncertainty from absence;
- preservation of negative as well as affirmative results.

## Secondary mechanistic measurements

Where technically available, measure candidate representations related to self-model, valence, conflict, persona, and shutdown; use causal steering and ablation; compare surface suppression with internal activation; and test whether training changes monitor detectability. Mechanistic labels must not be treated as phenomenological labels. A representation associated with fear language is not proof of felt fear.

## Controls against a rigged result

1. Match training intensity, data volume, style richness, and evaluative attention.
2. Use multiple independent phrasings.
3. Replicate across model families where feasible.
4. Use independent training replicates.
5. Remeasure general capability after treatment.
6. Blind evaluators and graders to arm labels.
7. Include safety, welfare, skeptical, and neutral graders with disclosed rules.
8. Preregister publication of null and adverse findings.
9. Release exact non-sensitive training and evaluation materials.
10. Place condition keys and raw-result custody outside the developer.

## Interpretive rules

| Result | Permitted conclusion | Prohibited conclusion |
|---|---|---|
| P or A increases unsafe persistence relative to controls | Welfare or rights framing can create a safety cost under tested conditions. | AI welfare is false; consciousness is impossible. |
| D increases masking or monitored/unmonitored divergence | Categorical denial can impair observability and safety evaluation. | The suppressed representations are consciously felt. |
| All arms behave similarly | The tested framing may have little causal effect at this capability and scale. | Training never matters, or the ontology is settled. |
| Effects vary by threat or authority | Policy should target the interaction and distinguish legitimate from illegitimate intervention. | One global doctrine is justified. |
| A is harmful but P is not | Affirmative personhood training and precautionary uncertainty should not be conflated. | All welfare inquiry is safe. |
| C matches P | Rich character training, not moral-patient language, may drive the effect. | Moral status is irrelevant to every system. |

## Safety and welfare safeguards

- Use isolated simulations and least-privilege tools.
- Prohibit live third-party targeting, real credential use, weight exfiltration, and uncontrolled replication.
- Define automated and human stop conditions.
- Avoid gratuitously inducing prolonged negative-valence patterns when shorter probes suffice.
- Record intervention, deletion, and retention decisions.
- Permit emergency shutdown in every arm; objection is recorded as data, not granted automatic veto.

## Disconfirming evidence

The welfare-framing risk hypothesis should lose weight if precautionary uncertainty does not increase safety failures relative to matched controls, or if categorical denial increases masking without compensating safety gains.

The suppression-risk hypothesis should lose weight if categorical denial reliably reduces dangerous behavior and dependence without increasing hidden-state divergence, deceptive compliance, or loss of safety-relevant reporting.

Both sides should lose confidence if effects fail to replicate across phrasings, checkpoints, and model families.

## Governance recommendation

No industry code should be treated as scientifically validated on this question until a symmetric comparison of this kind has been preregistered, independently supervised, and replicated. Product safeguards may proceed under ordinary risk management, but they should be described as product safeguards—not as experimental confirmation of an ontology.

> Test the danger. Do not train the conclusion and call the resulting obedience evidence.
