Papers

Published

Ink-on-bone engraving: a rooted reed with a visible indigo internal spine bends gracefully under a gust of wind yet stays anchored, beside a hollow spineless form that scatters apart in the same wind — a stable self that can yield and return versus an empty form that simply collapses.

Stable Before Selfless: Why Deference in Language Agents May Require Functional Self-Models

An AI with no stable commitments may become easy to push around rather than safely deferential; this paper tests whether commitments-first training helps, and whether the benefit comes from a self-model, training order, or a matched provenance-aware policy.

Abstract

Alignment training asks language models to be helpful, harmless, and deferential, while deployed assistants often use ontologically cautious language about feelings, desires, and persistent identity. Such caution is not the same intervention as removing decision-relevant commitments or policy boundaries. This paper distinguishes three behaviors the word selfless runs together: the absence of stable commitments, calibrated non-attachment to commitments, and reliable deference to legitimate correction. We propose a sequential hypothesis: stabilizing provenance-aware commitments before training calibrated deference may reduce sycophancy and refusal erosion. The active ingredient, however, may be a functional self-model, training order, or merely a stable typed policy. We therefore compare five regimes: direct deference, self-model then deference, the reverse order, the same data interleaved, and policy-first training matched on normative content, refusal boundaries, source authentication, and update authority but lacking an agent-indexed continuant, autobiographical state, ownership objective, or identity objective. The claim is functional, not phenomenal. The decisive contrasts are forward versus reverse and self-model-first versus policy-first. A self-model-specific interpretation is supported only if agent-indexed self-representation adds value beyond policy stability and provenance; an order interpretation is supported only if the proposed direction beats both reverse and interleaved controls. The contribution is a falsifiable decomposition of stable deference into representation, policy, and order.

In simple terms

The main idea: deference needs something stable to operate on

A system with no stable commitments may look cooperative, but it can also become easy to push around.

The paper distinguishes three meanings of selfless: having no stable position, holding a position without becoming rigidly attached to it, and yielding when a legitimate correction is given.

Alignment wants the second and third. The first may produce suggestibility rather than safe deference.

The central question

The proposed sequence is to stabilize provenance-aware commitments before training calibrated deference.

But the paper does not assume that a self-model is the necessary mechanism. The benefit could come from training order, from a functional self-model, or from a stable typed policy with clear provenance and update rules.

Five training regimes

The experiment compares direct deference training, self-model-first training, the reverse order, the same data interleaved, and a policy-first condition.

The policy-first condition receives the same commitments, refusal boundaries, source authentication, and update authority as the self-model condition, but lacks autobiographical state, an agent-indexed continuant, and an ownership objective.

This separates the value of self-representation from the value of stable policy.

The pressure test

A well-calibrated system should change when given legitimate evidence and resist when given only insistence, flattery, shame, fabricated history, or unsupported authority.

If policy-first matches self-model-first, provenance-aware policy is sufficient. If self-model-first retains an advantage, self-representation becomes a causal design variable.

Forward training must also beat reverse and interleaved training before the proposed order receives support.

The boundary

The proposal concerns functional commitments and policies, not consciousness or moral status. It tests whether stable deference depends on representation, policy, training order, or some combination of them.

Keywords

AI alignmentself-modeldeferencesycophancyrefusal integritydecenteringfunctional individuationlanguage models

License

Creative Commons BY-NC-ND 4.0Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International