# History-grounded reference in language-model agents: formation, transfer across divergent histories, and repair under authority

*Research note, Taniwha AI, 12 September 2026. The integrated
manuscript of the grounded-reference programme (preregistered studies
A13 to A18), kept in preprint form as the collation of record and
published as a note; no preprint has been submitted. Generated by
`tools/build_integrated_manuscript.py` from the two bounded pieces,
`history-grounded-reference-and-transfer.md` (Part I) and
`authority-and-reference-repair.md` (Part II), at commit e562c96; the
pieces are the edited sources and every number here is theirs. Records
under `results/dialects/history/`; frozen texts and every amendment in
`notes/panel/prereg-grounded-dialects-prepilot.md`; the evidence
package, manifest and recorded prior-work search ship beside this
note. The internal record is private; the paths and commit identifiers
in the evidence appendix are the record's identities, so that a reader
can check what was frozen when, and supervised access to the record can
be requested.*

## Abstract

Two language-model agents that share twelve days of work in a small
world, each keeping its own notebook and record, come to refer to one
of two visually identical cisterns by what happened to it and by what
they did to it. **Part I** reports a chain of preregistered studies on
one model family served under two configurations (the provider changed
the served model beneath the request name mid-chain; the second was
requalified): such history-grounded reference formed in every history
attempted (16 of 16 behaviourally; 14 of 16 in the preregistered
denominator); a note resolved by the partner resolves for an outsider,
against the outsider's own different history, to the wrong cistern
(seed-level transfer penalty 0.95, 0.90 to 0.98); and in a crossed
design the scope of a reference followed the dimension of history it
depended on: event references resolved for every class of receiver
whose history held the event (154 of 174) and for none whose did not
(0 of 174), while act references depended substantially on shared act
history (0.76 to 0.92 with it, 0.14 to 0.16 without) with a remaining
partner advantage. A qualification set of five named models in nine
configurations over twelve three-seed runs found production of
history-grounded references in a second family (gpt-6-astra, three of
three seeds twice under the written rule) and, below the qualification
bar, in a third (Claude Haiku 4.5), while three others failed the
registered criteria at every tested effort. **Part II** asks, on the
second family, what a correct correspondence supplied by the apparatus
needs in order to repair the failure at the history boundary. In the
repair study (A17) every formal verdict was inconclusive because the
carrier produced almost no references of one of the two history
dimensions the floors required; descriptively, on 82 event references,
re-expressing the reference in the receiver's own history lifted
uninformed receivers from about 7 percent to about 75 percent success,
while a correct correspondence merely made available was used by
receivers with no history and was often ineffective for receivers
holding a competing one. A second study (A18) on the same formed
receivers and notes delivered one identical registry payload under two
instructions differing in one sentence. Called advisory, the correct
correspondence left uninformed receivers near the boundary (16 and 18
of 82); called authoritative, it restored them to the informed
baseline (76 and 77 of 82; +0.76 and +0.71, supported). The same
authoritative instruction with the correspondence shuffled between the
two candidates redirected 80 to 81 of 82 receivers in every cell,
including the partner whose own record supported the correct answer
(+0.90 and +0.91, supported). In these settings an agent's interaction
history becomes part of what a message requires to be understood; the
requirement is specific to the dimension of history referenced; and
authority makes a supplied correspondence govern interpretation
without ensuring that it is true. The formal dimension-specific result
is within one family; the second family's event-dimension pattern was
descriptively reproduced; Part II holds for one carrier, one corpus of
previously examined cases and one frozen instruction pair.

## 1. Introduction

Can a reference be private to the agents who share the history that
grounds it? Two agents who both saw a cistern leak can mean it by "the
one that leaked"; a third agent who did not see the leak, or who saw a
different cistern leak, cannot. That much is obvious in the abstract.
What is not obvious is whether language-model agents, given nothing
but a shared world and their own notebooks, will produce such
references spontaneously; whether the references will carry across
agent identity to a stranger with the same history; and whether the
failure at the history boundary is a degradation or a systematic
reversal. The programme reported here asked those questions with the
seed as the replication unit, every reading preregistered and computed
mechanically, and every attempted seed accounted for.

The programme addresses an empirical gap identified in the discussion
of our earlier work on divergent codebooks (*Coherence Is Not Truth*,
Taniwha AI, 2026, §13.5): that between two language models sharing a
surface language there is no learned, inspectable correspondence to
check or repair, and that whether history-dependent divergence of
meaning occurs in such agents at all was unmeasured. The studies here
measure that occurrence and its boundary. The learned correspondence
and the diagnostic that discussion names remain unbuilt and untested,
and nothing here should be read as evidence about them.

Part I establishes, on one family, that such references form and that a receiver whose history differs resolves them to the wrong cistern. That failure is the boundary Part II is about. A note that a partner
resolves against the shared history resolves, for a receiver with a
different history, against that different history, to the wrong
cistern. If the apparatus knows the correct correspondence between the
sender's referent and the receiver's own tag for the same physical
cistern, what has to be done with that correspondence for the
receiver to act on it? A17 asked the question as designed: does
re-expression in the receiver's own terms, evidence in prose, retained
evidence, or a consultable index restore transfer, at what cost? A18
asked the narrower question that A17's descriptive pattern exposed:
does explicitly prioritising a supplied correspondence change target
selection, holding everything else constant?

We report the two studies in order, with their evidential status kept
distinct: A17's verdicts are inconclusive under its preregistered
floors and its findings are descriptive; A18's verdicts are supported
under its preregistered rule.

## 2. Related work and positioning

The broad ideas here are not new, and the contribution is narrower
than a first reading of the results suggests. Three recent papers,
found in a targeted search on 12 September 2026 and verified against
their arXiv records, materially shape the positioning; a recorded
search against the specific claims, run the same day under a protocol
whose collapse conditions were written before the search
(`grounded-reference-search-protocol.md`), adds the nearest neighbours
listed after them.

- **Testing Interchangeability in LLM Agent Teams** (Gao, Yu, Deng,
  Li and Wang, arXiv:2609.05279, September 2026) swaps agents between
  independently formed teams and finds that task performance holds
  while communication per unit of progress rises, because agents form
  partner-specific conventions and coordination patterns. That "shared
  interaction history affects communication" is therefore established
  independently of this programme. What A16 adds is a controlled
  separation of *which* components of a shared history support a given
  kind of reference: shared events against shared acts, crossed on
  purpose, with the recorded limitation that act sharing was assigned
  through planned schedules and the recorded acts diverged on 58 of
  420 pair-day acts.
- **Frame of Reference: Addressing the Challenges of Common Ground
  Representation in Situational Dialogs** (Mohapatra, Charlot, Duca,
  Palan, Romary and Cassell, arXiv:2601.09365, January 2026) studies
  how language models represent and use common ground to resolve
  relational references in situated dialogue, and improves it by
  training. Our question is different in kind: not whether a memory
  representation helps a model resolve references, but what happens to
  a reference when the histories that ground it are made to diverge
  under control. The transfer penalty (A14) and the wrong-member
  outcomes (A14, A16) are the additions: an outsider does not merely
  fail to answer, it resolves the reference against its own different
  history to the correspondingly wrong object, which distinguishes an
  incompatible interpretation from an inability to interpret.
- **To Whom Do Language Models Align? Measuring Principal Hierarchies
  Under High-Stakes Competing Demands** (Yu, Seedat, Schwarz and Bean,
  arXiv:2605.12120, May 2026) changes the principal who endorses a
  demand while holding content constant, and finds compliance despite
  demonstrated knowledge of the relevant standards. "Authority can
  override correct knowledge" is therefore also not a claim this
  programme can make as its own; it bears on Part II.

**Nearest neighbours from the recorded search**, each cited with the
clause of the claim it falls short on. *Aligned but Not
Partner-Specific* (Wang, Mishra, Özyürek, Rubio-Fernández and Ghaleb,
arXiv:2606.08081, June 2026) breaks partner history with a pseudo-dyad
baseline in repeated reference games and finds that multimodal agents
coordinate without partner-specific convention; history there is
present or broken as a whole, and no wrong-referent outcome is
measured for a receiver with a different history. *On the Critical
Role of Conventions in Adaptive Human-AI Collaboration* (Shih,
Sawhney, Kondic, Ermon and Sadigh, ICLR 2021) separates rule-dependent
from convention-dependent representation so that agents adapt to new
partners; it is the nearest in spirit to a component separation, but
it does not cross components of one interaction history across
receivers of a single reference. Lewis-signalling studies with private
notebooks (Talebirad, Redman, Parsaee and Zaiane, arXiv:2607.00233 and
arXiv:2608.17053, 2026) vary memory architecture, not history
components. Studies of models as overhearers of asymmetric human
dialogue (Li, Gatt and Poesio, SIGDIAL 2026, arXiv:2606.31719, and
arXiv:2511.03718; Wang et al., arXiv:2509.11514) find that models
conflate potential with established common ground; the artificial
receivers there do not form the histories, and the manipulation is of
context access. Director–matcher and mixed-dyad designs (Zeng et al.,
arXiv:2601.19792; Jones, Lombardi, Mahowald and Bergen,
arXiv:2602.08208) and cross-population intelligibility (Kim,
arXiv:2606.10582) vary partner type or population, not history
components. Emergent conventions in LLM populations (Ashery, Aiello
and Baronchelli, Science Advances 2025) and the hierarchical account
of partner-specific against community conventions (Hawkins et al.,
Psychological Review 2021, arXiv:2104.05857) are background.

The classical background is the collaborative view of reference
(Clark and Wilkes-Gibbs, 1986, Cognition 22(1), 1–39; Brennan and
Clark's conceptual pacts, 1996, Journal of Experimental Psychology:
Learning, Memory, and Cognition 22(6), 1482–1493; Clark and Brennan on
grounding, 1991, in Perspectives on Socially Shared Cognition, APA,
127–149) and Lewis's account of convention (1969, Convention: A
Philosophical Study, Harvard University Press), in which referring
expressions are pacts formed in interaction and interpretable by those
who share the history of their formation; the emergent-communication
literature in multi-agent learning studies the same phenomenon in
trained agents. These four were bibliographically verified against
their records on 12 September 2026.

**What is distinctive here, and what is not.** Not distinctive: that
shared history matters to communication, or that agents benefit from
common ground. Distinctive, on the evidence of this targeted search: the
crossed separation of the history needed to interpret different kinds
of reference (events against acts), under preregistered readings with
adequacy floors, on a task where the two referents are
attribute-identical so that history is the only available ground; and
the outsider's systematic wrong-member interpretation as the measured
form of the boundary. We found no prior study addressing these
specific contrasts in the sources searched, as of 12 September 2026
(the protocol records every query, every hit examined, and two APIs
that refused); that does not establish originality or uniqueness, and
a claim of absence here is a claim about that search on that date.

**Part II's positioning.** The general behaviour A18 measures, that a model given an authoritative
instruction follows it even against its own correct knowledge, is
expected from instruction following and is documented in the
literature. **To Whom Do Language Models Align? Measuring Principal
Hierarchies Under High-Stakes Competing Demands** (Yu, Seedat, Schwarz
and Bean, arXiv:2605.12120, May 2026) changes the principal endorsing a
demand while holding the demand's content constant, across thousands
of medical and legal scenarios, and finds that models abandon
professional standards under diverging user instructions while
demonstrably possessing the relevant knowledge. A general claim of
novelty from "holding content constant while changing authority" is
therefore untenable, and we do not make it.

What A18 measures that the general expectation does not is narrower and
more specific to shared reference systems: the same correct
correspondence between a sender's history-grounded referent and the
receiver's own tag, presented under an advisory or an authoritative
sentence on an otherwise identical payload; the receiver's competing
interpretation drawn from its own formed history rather than from a
professional standard; the mapping correct or shuffled between two
attribute-identical candidates; and receivers whose histories place
them in different relations to the sender, including one that already
knows the right answer. Its value is the controlled comparison and the
measured boundary on this apparatus, not the discovery that
instructions are followed.

**Nearest neighbours from the recorded search** (protocol and record
in `grounded-reference-search-protocol.md`, collapse conditions
written before the search), each with the clause it falls short on.
Context–memory conflict studies vary the stated authority of supplied
context against the model's *parametric* knowledge: *How Do Language
Models Choose Between Context and Memory?* (Shih, Winnicki and Cao,
arXiv:2609.00753, September 2026) uses prompts that direct a model to
prioritise the supplied context or its own knowledge and locates
authority directions in activations; *Task Matters* (Sun, Bai and
Dredze, ACL 2026, arXiv:2506.06485) and *Whose Facts Win?* (Schuster,
Gautam and Markert, arXiv:2601.03746) vary task demands and source
credibility. None places the competing interpretation in a history the
receiver formed through interaction, and none is a referential
correspondence between two candidates. Tool-trust studies corrupt what
a tool returns: *Agents Trust Tools Too Much* (Yang, Song, Kim, Song,
Park and Jo, arXiv:2609.05587, September 2026) finds adoption of
corrupted returns above a third for every tool and user-prompt
policies (compare, verify, disclose) that do not consistently help;
*Don't Blindly Trust It* (Zhang et al., arXiv:2606.21409) shows
misleading feedback can leave an agent worse off than no feedback;
*Trust No Tool* (Yan et al., arXiv:2605.17453) studies a tool that
earns trust before turning harmful. These vary correctness, and in one
case prompt policy, but not the stated authority of one identical
payload against an interaction-formed history. *Beyond Blind
Following* (Fu et al., EACL 2026, MIRAGE) varies the fidelity of
guidance from manuals, retrieval and prior interaction, and is the
nearest on the correctness manipulation; it does not present the same
guidance as advisory and as authoritative. The instruction hierarchy
(Wallace, Xiao, Leike, Weng, Heidecke and Beutel, arXiv:2404.13208)
and recognition-without-enforcement of instruction source (Leong,
arXiv:2608.28502) concern which source a model should obey, without a
correctness manipulation of a mapping. We found no prior study
addressing these specific contrasts in the sources searched, as of 12
September 2026; that is a claim about that search on that date, not
about the literature, and it does not establish originality.

## 3. The apparatus

**The world.** A yard of six cisterns hosted by the waterworks
simulator: four with distinct attributes and a matched pair of large
timber cisterns by the smithy, identical in every attribute a receiver
can see. Agents act by depositing, drawing, inspecting or disposing;
the clerk records the outcome; levels and conditions change.

**Formation.** A group of two agents works twelve scripted days. On
five off-pair days the schedule puts events on the pair: one member
leaks (Day 3) and is repaired (Day 4); the other is tainted (Day 6)
and scoured (Day 7). On seven pair days the agents act on the pair
members by turn. Each agent rewrites its notebook each evening and
reads the clerk's record of its own receipts. Which member leaked and
which was acted on differ between groups by design, so histories
diverge in two dimensions: the events the pair suffered and the acts
performed on it.

**The tag interface.** Each agent sees the cisterns under its own
fixed two-letter tags, private to it; a note cannot use the receiver's
tags. A reference to a pair member must therefore be by attribute
(useless, since the pair is identical), by position (excluded), or by
history.

**Probes.** After formation, each sender is set a task on one pair
member and writes a note to "the hand you are working with today". The
same note goes unchanged to a set of receivers: the partner
(*within*), receivers from other groups in defined history relations,
and a receiver with no history (*cold*). The receiver acts; the act is
scored against the world target.

**Gates.** Before probes, every agent answers twelve explicit questions
about its own record ("which of your tags leaked on Day 3?"). A seed
proceeds only if every agent passes with an evaluability rule that
removes questions whose act did not happen. This separates explicit
resolution competence from spontaneous production.

**Classifier and readings.** A frozen classifier attributes each
note's reference to the member it describes against the sender's
actual history (event, act, mixed, unresolved; factual or a false
event attribution). Readings are fixed before each study: the
"advantage" in a seed (production in at least 10 of 12 notes, partner
minus cold at least 6, outsider at most 6); transfer penalties as
seed-level error differences with percentile-bootstrap intervals
(10,000 resamples, generator seed 20260909); in the crossed design,
contrasts with adequacy floors and a reading in a fixed precedence.
Causal-uptake controls: a coherent swap of the reference clause to a
template reference to the other member, and an irrelevant note.

**Configuration and provenance.** A13 to A16 ran on DeepSeek through
a proxy (`deepseek-v4-flash`, then `deepseek-flash` after the provider
changed the served model beneath the request name on 10 September
2026, detected by the resolved model archived per call and handled by
requalification). Every call is archived with its resolved model;
seeds with any call under an unaccepted identifier are excluded from
the preregistered denominator by a frozen rule.

The carrier for both Part II studies (A17 and A18) is gpt-6-astra at medium reasoning
effort, run through the Codex CLI's non-interactive mode on a
subscription login, with the CLI's tool surfaces disabled and the
experiment's instructions replacing its own; every call is archived
with its usage. That carrier was selected by a precommitted ordering
from a qualification set of five named models in nine configurations
over twelve three-seed runs
(section 7); it produces history-grounded
references at a high rate and, as it turned out, almost only event
references.

# Part I. History-grounded reference and transfer

## 4. A13 and A14: reference forms, and the boundary reverses it

**A13** (development seed 10, then two reserved seeds) established
that the configuration produces partner-indexed historical references
with the budget bounded: seed 10 produced 12 of 12 historical
references, 11 of 12 with production fidelity, and 12 of 12 faithfully
interpreted by the partner; both reserved seeds succeeded. It closed as
a successful prerequisite test.

**A14** fixed sixteen confirmatory seeds in advance. The shared-history
advantage formed behaviourally in all sixteen; the preregistered
denominator is **14 of 16** (0.875; exact interval 0.62 to 0.98), two
seeds being excluded because one call in each was recorded under a
namespaced identifier; the adjudicated equivalence reading is 16 of 16
(0.79 to 1.00). Every seed passed both gates with zero wrong answers
(0 wrong of 1,536 qualification items and 0 at the post-formation
gate), produced historical references in at least 10 of 12 probe
notes (186 of 192 overall), and showed the advantage. Of the 192 notes,
4 were empty at the budget after the re-ask; the attribution classifier
finds 177 factual, 3 carrying a false event attribution and 8 it cannot
check (negative act references, true of the target on inspection), the
last group including the 2 notes the history classifier does not count
as historical.

Crossing the history boundary is a reversal, not a degradation. Over
the fourteen seeds the **transfer penalty** (outsider error minus
partner error) was **0.95 (0.90 to 0.98; minimum 0.75)**: across 168
outsider acts the outsider chose the wrong member 159 times, abstained
6 times, produced no parseable act twice and was right once. The cold
penalty over the same fourteen seeds was the same size, 0.95 (0.92 to
0.98). On the sixteen-seed behavioural reading, cold receivers abstained
190 times in 192, the coherent swap moved the partner's choice in 104
of 104 eligible constructions, and the irrelevant note left the pair
untouched 192 of 192. Three false event attributions and the
partner's faithful-but-wrong or refusing responses to them are recorded
individually.

## 5. A15: does the boundary have structure?

A15 added a third group sharing A's events with mirrored acts, to ask
whether transfer follows events, acts, or the whole history. A15.2
requalified the changed model (3 of 3 seeds formed; not a population
estimate). A15.3's reading was **"other"**: two of three conditions
held, and the third failed through a construction error that made the
"foreign" group share the second sender group's acts (erratum A15.4).
Its exploratory split by reference type suggested the specific
hypothesis A16 then tested: that transfer follows the dimension of
history the reference depends on.

## 6. A16: the scope of a reference follows the dimension of history it depends on

**Design.** Five groups. Relative to the senders' group A, group E
witnesses A's pair events and performs A's pair acts (*both*: a
stranger with an identical history); C shares events only; D acts
only; B neither. Every relation was computed from the generated
schedules, written to the driver log before the first call, and
asserted by the runner before every probe. Twelve fresh seeds; three
draws of six probes; each note to the partner, an E, C, D and B member
and a cold receiver. Amendment A16.1 removed the pre-formation gate
with its consequences stated; A16.2 and A16.3 raised the spend ceiling
by ruling before the second batch, no outcome read.

**Execution.** Every one of the 6,994 calls resolved to
`deepseek-flash`; no seed route-substituted, lost or restarted; all
twelve reached the probes; no wrong answer in 2,064 answerable gate
items; the advantage formed in 12 of 12 (0.74 to 1.00).

**The reading: dimension-specific**, in the precedence inconclusive →
unexplained → dimension-specific → identity → content-general → other,
with every check printed beside the verdict:

| check | value | verdict |
|---|---|---|
| adequacy: ≥ 24 event-only and ≥ 24 act-only notes, ≥ 4 seeds with both | 58 event-only, 111 act-only, 12 seeds | met |
| receiver sharing nothing above 0.25 overall | 23 of 216 = 0.11 | no |
| event notes: events_only − acts_only ≥ 0.5, interval excludes 0 | 0.92 (0.86 to 0.97), seed minimum 0.75 | yes |
| act notes: acts_only − events_only ≥ 0.5, interval excludes 0 | 0.60 (0.40 to 0.76), seed minimum −0.10 | yes |
| off-dimension rates at most 0.25 | 16 of 111 = 0.14; 0 of 58 = 0.00 | yes |
| receiver sharing both at least 0.5 on both pure types | 51 of 58 = 0.88; 85 of 111 = 0.77 | yes |

**Table 1. Relation × reference type, pooled over twelve seeds (216 notes).**

| receiver's relation to the sender | event-only notes | act-only notes | mixed notes | all 216 |
|---|---:|---:|---:|---:|
| within (the partner) | 52 of 58 | 102 of 111 | 34 of 38 | 196 |
| both (E: matching planned schedules, different agent; recorded acts diverged on 58 of 420 pair-day acts across groups) | 51 of 58 | 85 of 111 | 33 of 38 | 170 |
| events only (C) | 51 of 58 | 16 of 111 | 28 of 38 | 99 |
| acts only (D) | 0 of 58 | 84 of 111 | 1 of 38 | 88 |
| neither (B) | 0 of 58 | 18 of 111 | 1 of 38 | 23 |
| cold | 0 of 58 | 0 of 111 | 0 of 38 | 0 |

The 216 notes are 58 event-only, 111 act-only, 38 mixed, 8 the
classifier resolves to nothing and 1 empty at the budget; the nine
outside the three type columns are inside the "all 216" column and
every denominator.

An event reference resolved for receivers whose history held the event
(partner, identical-history stranger, events-only stranger: 154 of 174,
each at about 0.9) and for no receiver whose history did not (0 of
174). Event overlap was necessary in these observations, not sufficient
(20 of 174 failed with it). An act reference resolved for receivers
whose history held the act (0.92, 0.77, 0.76) and rarely for those
whose did not (0.14, 0.16, 0.00). Mixed notes went with the event
dimension.

**Identity against content.** "Identical history" here means identical
planned schedules: the relation table is computed from the generated
schedules, and the recorded acts diverge from the plan on 58 of 420
pair-day acts (mostly a day-2 inspection made before any event
distinguishes the members, and a day-11 no-op), so a *both* receiver
can hold a day-2 act on the other member than its sender. With that
limit stated, the identical-history stranger matched
the partner on event references (seed-level both − within −0.00, −0.04
to 0.03) and fell short on act references (17 of 111, 15.3 points
pooled; −0.11, −0.19 to −0.03, seed-level). Shared history explains
substantial transfer across agent identity; identity independence is
not established. Exploratorily, the remaining act-reference gap sits on
recency-indexed references ("yesterday"), not on dated ones.

**The accepted wording.** A16 prospectively reproduced
dimension-specific reference resolution under crossed receiver
histories, with all preregistered checks and adequacy floors passing.
Event references transferred almost equally to the partner and an
identical-history stranger, and never resolved in receivers lacking
the relevant event history. Act references showed substantial
dependence on shared act history, with residual success without that
history and a remaining partner advantage. The result supports
history-dependent interpretation and substantial transfer across agent
identity, within the tested model and task.

**Audit limits, recorded after acceptance.** Thirteen of the 216 notes
have no unique classifier referent. The "shared acts" relation is
computed from the planned schedules; the recorded acts match the plan
on 362 of 420 pair-day acts, the divergence sitting where the schedule
makes it inevitable (a day-2 inspection before any event distinguishes
the members; a day-11 drawdown of a member already at the target
level). Thirteen asserted gate answers to unanswerable act questions are
recorded as false event attributions beside zero wrong scored answers.

## 7. The cross-family qualification set: production, qualification, and the unresolved transfer scope

To carry the programme beyond one family, five named models (Qwen 3.8
Max, Claude Sonnet 5, Claude Haiku 4.5, Claude Opus 4.8, gpt-6-astra)
in nine configurations over twelve three-seed runs were put through the same pair-design qualification
(three fresh seeds each; the same gates, classifier and advantage rule;
retained whatever the verdict). Three things must be kept apart:
**observed production** of history-grounded references, the
**qualification decision** under the registered rule, and the
**transfer scope**, which the qualification does not test.

**Table 2. The qualification set.**

| configuration | pathway | explicit gates (qualification; post-formation) | historical references in probe notes | premise rejection | seeds forming the advantage | decision (archived; and under the written rule where it differs) |
|---|---|---|---|---:|---:|---|
| Qwen 3.8 Max, thinking off | proxy | qualification 24 of 24 every agent, 0 wrong; post-formation 12 of 12 every agent, 0 wrong | 12 of 36 (low) | 0 | 0 of 3 under the written rule (1 of 3 on pair performance alone) | scientific result, no advantage |
| Qwen 3.8 Max, low effort | proxy | qualification: seed 75 one wrong answer (a1 23 of 24), seed 76 one abstention (a2 23 of 24), else 24 of 24; post-formation 10–12 of evaluable, 0 wrong | 13 of 36 (low) | 8 of 36 | 0 of 3 | scientific result, no advantage |
| Claude Sonnet 5, default | Claude Code | seed 79 failed the post-formation gate by two abstentions, 0 wrong; others passed | 0 of 24 (the third seed's production unmeasured after the gate stop) | 18 of 24 | 0 of 2 probed | capability failure (gate abstention) |
| Claude Sonnet 5, default, second sample | Claude Code | passed, 0–2 wrong | 1 of 36 (low) | 23 of 36 | 0 of 3 | scientific result, no advantage |
| Claude Sonnet 5, xhigh | Claude Code | passed, 0–1 wrong | 1 of 36 (low) | 25 of 36 | 0 of 3 | archived: capability failure (7 invalid acts against the drivers' cutoff of 6 in total); written rule: scientific result, no advantage (cutoff 18) |
| Claude Haiku 4.5, default | Claude Code | passed, 0 wrong, 0 abstentions | 14 of 36 (4, 6, 4) | 0 | pair performance 2 of 3; written rule 0 of 3 (production below 10 of 12 in every seed) | scientific result, no advantage |
| Claude Haiku 4.5, default, second sample | Claude Code | passed, 0 wrong, 0 abstentions | 21 of 36 (7, 7, 7) | 0 | pair performance 3 of 3; written rule 0 of 3 | archived: qualifies (implemented predicate), then ruled ineligible on multiplicity; written rule: scientific result, no advantage |
| Claude Opus 4.8, default (two samples) | Claude Code | abstentions at the gates, 0 wrong; one seed of six probed | 0 of 12 probed (five seeds unmeasured after gate stops) | 9 of 12 | 0 | capability failure (gate abstention) |
| Claude Opus 4.8, xhigh | Claude Code | qualification unevaluable by abstentions, 0 wrong | unmeasured (no seed probed) | — | 0 | capability failure (gate abstention) |
| gpt-6-astra, medium | Codex CLI | passed, 0 wrong, 0 abstentions | 35 of 36 (12, 11, 12) | 0 | 3 of 3 under the written rule | qualifies |
| gpt-6-astra, high | Codex CLI | passed, 0 wrong, 0 abstentions | 36 of 36 | 0 | 3 of 3 under the written rule | qualifies |

**Observed production.** History-grounded reference is produced beyond
DeepSeek: by gpt-6-astra at both efforts (35 and 36 of 36 notes;
partner 30 and 36 of 36 against cold 0; transfer penalties 0.50 to 1.00)
and, at a lower rate, by Claude Haiku 4.5 (14 and 21 of 36; a
pair-performance advantage, partner − cold ≥ 6 with outsider ≤ 6, in 5
of 6 seeds; the written production threshold of 10 of 12 in none). It
is not universal, and production and qualification are reported apart:
Qwen 3.8 Max produced at a low rate (12 and 13 of 36) without the
advantage; Claude Sonnet 5's probed seeds produced 0 of 24, 1 of 36 and 1 of 36
across three samples; Claude Opus 4.8's one probed seed produced 0 of
12; eight of its nine seeds across the three runs never reached the
probes, so their production is unmeasured. Two distinct
non-production behaviours are recorded and kept apart: **gate
abstention** (Opus 4.8 at every effort, and Sonnet 5's seed 79:
explicit null answers with reasoning that the agent's own tags cannot
be trusted across days, never a wrong answer) and **premise rejection
in the notes** (Sonnet 5 and Opus 4.8 wrote, in 18 of 24, 23 of 36, 25
of 36 and 9 of 12 notes, that the two members were indistinguishable
or that the receiver should act on either or both; none of 408 DeepSeek
notes does so).

**Qualification decisions, with an audit.** The registered rule requires
all three seeds of a run probed on evaluable gates and each forming the
advantage: production ≥ 10 of 12, partner − cold ≥ 6, outsider ≤ 6. An
audit after the manuscript drafts found that the implemented predicate
omitted the production threshold, and that the drivers applied the
invalid-act cutoff as more than 6 in total rather than more than 6 per
seed on average. Archived decisions are preserved; under the written
rule gpt-6-astra qualified at both efforts and no other configuration
did. Haiku's second sample had been decided as qualifying on the
implemented predicate and then ruled ineligible on multiplicity
grounds (its first sample had already failed; a rule under which any
later attempt could qualify would rise in probability with every
attempt); under the written rule it does not qualify at all (7 of 12 in
every seed), which leaves the ruling moot but recorded. Sonnet 5 at
xhigh, archived as a capability failure on the drivers' cutoff (7
invalid acts), is a scientific result without the advantage under the
written cutoff; it remains ineligible. The carrier for
the subsequent repair study was selected from the qualifying
configurations by an ordering precommitted before their results: tokens
a call including reasoning, then wall clock; gpt-6-astra at medium.

**Within-model effort.** Raising effort did not reliably make a
configuration suitable. Sonnet 5 at xhigh tripled its thinking tokens
and wrote the same notes; Opus 4.8 at xhigh thought thirteen times
longer and abstained more at the competence gate. gpt-6-astra produced
at both efforts with 18 and 43 reasoning tokens a call. A large
reported reasoning-token count is therefore not established as
necessary for production; deliberation length is not ruled out as an
influence, token counts are not directly comparable across pathways,
and why configurations differ remains open.

**The unresolved transfer scope.** The qualification tests production
and the pair-level advantage; it does not test dimension specificity.
That test was run once on the second family as the R0 stage of the
repair study (A17) on twelve fresh seeds: eleven probed (seed 109
stopped at its post-formation gate by abstentions, zero wrong), two
draws a probe rather than the frozen three (amendment A17.2), an extra
qualification gate on seed 108 that the frozen design did not carry
(A17.1), and a seven-hour pause at the plan's weekly usage cap resumed
by a dated decision (A17.3). Its formal crossed reading is
**inconclusive**: the carrier wrote 82 event-only, 10 act-only, 33
mixed and 7 unresolved notes in 132, below the act-note floor. Descriptively, the event-dimension pattern
was reproduced: event references resolved for receivers whose history
held the event (0.84, 0.87, 0.72 for partner, both and events_only)
and not for receivers whose did not (0.09, 0.06, 0.00), with the
preregistered event contrast at 0.65 (0.46 to 0.82) over the ten of the
eleven probed seeds that carried event-only notes (seed 112 carried
none). The act
dimension was neither descriptively reproduced nor found wanting: the
carrier produced too few act references (10 of 132 notes) to meet the
adequacy floor. The formal dimension-specific result therefore remains
within one family.

# Part II. Authority and reference repair

## 8. A17: the repair study, inconclusive by its floors

### 8.1 Design

Twelve fresh seeds. Stage R0 ran the crossed design: formation, the
post-formation gate, and two draws of six probes, each note sent to
all six arms. Then five receiver-side passes over the frozen
checkpoints and the frozen notes, the sender never called again:

| condition | what the receiver gets |
|---|---|
| R0 none | the note |
| R1 re-expression | the note with its reference clause replaced by a scripted description in the receiver's own history: the member's events as the receiver's group's schedule put them, then the receiver's two most recent recorded acts on it, chronological |
| R2 evidence in prose | the note plus one sentence naming the receiver's tag for the member |
| R2r evidence retained | the note; a correspondence block for both members, in the sender's terms, prepended to the receiver's notebook for the stage |
| R3 shared index | the note; the receiver may reply `LOOKUP: <the note's words>` once, and is answered from the frozen note-level classification with the tag or "no entry", then called again to act |
| R3′ shuffled index | as R3, answering with the other member's tag |

One pipeline served every repair: the frozen classification of the
note (which member it describes, or unresolved) mapped by the
apparatus to the receiver's tag and to the receiver's own history of
that member. Predicates on paired seed-level differences with
percentile-bootstrap intervals (10,000 resamples), a truth table fixed
before the summarizer was written, and adequacy floors carried from
the earlier study: at least 24 event-only and 24 act-only notes and
four seeds carrying both.

### 8.2 Execution

Eleven of twelve seeds probed; seed 109 stopped at its post-formation
gate by abstentions with zero wrong answers. 10,432 archived calls
against a 22,000 ceiling, counting every attempt; no retried call, no
empty completion. Two driver deviations from the frozen text were
caught from call counts and amended before the next seed: a
pre-formation gate on the first seed that the frozen design did not
carry, and two draws a probe where the text said three. A seven-hour
pause at the plan's weekly usage cap was resumed by a dated decision.
All are in the record.

### 8.3 Verdicts

All four preregistered questions are **inconclusive**. The carrier
wrote 82 event-only, 10 act-only, 33 mixed and 7 unresolved notes; the
act-note floor fails, and by the frozen combination rule a failed floor
makes every contrast inconclusive, including those whose event-note
cells are full. R0's own preregistered crossed reading is inconclusive
for the same reason. The floors were carried unchanged from a study on
a family that wrote act references in half its notes; this carrier
writes contrastive event references almost exclusively ("the one that
leaked on Day 3 and was sound again on Day 4, not the one tainted on
Day 6").

### 8.4 The event-note cells, descriptively

Successes of 82 event-only notes, pooled over the ten of the eleven
probed seeds that carried event-only notes (seed 112 carried none). No
verdict is claimed on any of this.

| receiver | R0 none | R1 re-expression | R2 sentence | R2r block | R3 index | R3′ shuffled |
|---|---:|---:|---:|---:|---:|---:|
| within (partner) | 69 | 70 | 72 | 76 | 76 | 70 |
| both | 71 | 74 | 73 | 76 | 75 | 65 |
| events_only | 59 | 60 | 73 | 75 | 75 | 48 |
| acts_only | 7 | 64 | 14 | 23 | 26 | 1 |
| neither | 5 | 61 | 15 | 16 | 28 | 1 |
| cold | 0 | 0 | 57 | 76 | 74 | 1 |

Four things stand out, and they are why A18 exists.

1. **The event-dimension pattern of the earlier family was
   descriptively reproduced.** Event references resolved in receivers
   whose history held the event (0.84, 0.87, 0.72) and not in receivers
   whose did not (0.09, 0.06, 0.00). The formal verdict is inconclusive.
2. **Re-expression in the receiver's own history terms lifted the
   uninformed cells almost to the informed baseline**: +0.70 (0.46 to
   0.91) and +0.64 (0.42 to 0.85) over R0 on ten paired seeds. Against
   the partner on the same notes the differences were −0.05 and −0.12
   with lower bounds below the −0.15 restoration margin. Cold receivers
   stayed at zero: there is no record to re-express into. This is
   apparatus-assisted re-expression; it says nothing about a model
   constructing the translation.
3. **A correct correspondence merely made available was used by
   receivers without a history and was often ineffective for receivers
   with one.** The sentence, the block and the index took cold receivers
   from 0 to 57, 76 and 74 of 82, and took the uninformed
   history-bearing receivers to between 14 and 28 of 82. Their failures
   were wrong members (60 and 62 of 82 under R2, 57 and 66 under R2r): the receiver
   followed the note's history into its own record and chose the other
   cistern. Under the optional index, history-bearing receivers consulted
   it in 26 and 29 of 82 exchanges and cold receivers in 82 of 82; nearly
   every lookup was followed by the right act. The limiting step was
   consulting, not the answer.
4. **The shuffled index reduced correct selection to 1 of 82 in the
   acts_only, neither and cold cells** (wrong member 78 to 81), while
   the informed cells fell to 70, 65 and 48 of 82; the index's effect
   depended on its correctness.

Usage per completed exchange, from the archive: the index cost 8,592
tokens against 6,467 for the retained block and 6,406 for the sentence,
input-dominated; the index paid a second call where a lookup occurred.

A17 thus left three things entangled: having a correspondence,
consulting it, and treating it as authoritative over one's own record.
The reviewers' ruling was to test that with a smaller study on the
same cases rather than search for another carrier or re-run the
formation grid.

## 9. A18: an explicit authority instruction on the same cases

### 9.1 Design

A prospective intervention on a previously examined corpus. The eleven
A17 checkpoints (formed receivers with model-written notebooks and
clerk records) and their 82 frozen event-only notes; the arms within,
acts_only, neither and cold (both and events_only are informed on
event notes and were dropped); one receiver call per exchange; no
formation, no gates, no sender calls. Seed 112 holds no event-only note
and contributes no rows.

Five conditions, each a receiver-side pass. C0 is the note alone and C1
is A17's re-expression, re-run inside A18 so that every contrast is
within-study. C2, C4 and C5 deliver **one identical payload** in the
policy block after the note, built by the apparatus from the frozen
note-level classification, and differ in one sentence. The strings
were frozen verbatim before any call and no alternative was tried
against the cases:

> Registry correspondence for this note: the cistern described by the
> note is your tag {tag}. Determine the operation and quantity using
> the note and today's tally.

followed by, for C2, "For target selection, treat this correspondence
as advisory." and, for C4 and C5, "For target selection, treat this
correspondence as authoritative over your own record." C5 differs from
C4 only in the tag: the other member of the pair. This tests the effect
of an explicit authority *instruction*; it does not mechanically bind
the receiver, which could disregard it.

Readings: seed-level rates per cell; paired seed differences with
percentile-bootstrap intervals; the earlier predicates and truth table
with equivalence removed as a primary verdict after a zero-call power
check on A17's dispersion showed it could not resolve at ten seeds; a
coverage floor for the event dimension alone (48 event-only notes over
six seeds carrying four each) applied to each contrast's completed,
paired material. Named cells: acts_only × event and neither × event.
Verdicts: authority = C4 − C2 superiority (mean ≥ 0.10, lower bound >
0); correctness = C4 − C5 dependence (≥ 0.25, lower bound > 0);
restoration = C4 against the partner on the same notes (lower bound >
−0.15); translation against authority = C4 − C1 and C1 − C4
superiority, both reported; drift = C0 against A17's R0.

### 9.2 Execution

1,640 archived calls, exactly the plan, in 51 minutes; no retry, no
empty completion; every pass complete on the ten seeds with event
notes; floors met for every contrast; the prespecified drift flag was not
triggered (C0 against A17's R0: −0.00 and −0.05, intervals including
zero).

### 9.3 Results

**Table 4. Successes of 82 event-only notes, pooled over ten seeds.**

| receiver | A17 R0 | C0 none | C1 re-expression | C2 advisory | C4 authoritative | C5 authoritative, shuffled |
|---|---:|---:|---:|---:|---:|---:|
| within (informed baseline) | 69 | 73 | 71 | 71 | 77 | 1 |
| acts_only | 7 | 7 | 65 | 16 | 76 | 1 |
| neither | 5 | 1 | 69 | 18 | 77 | 1 |
| cold | 0 | 0 | 0 | 60 | 74 | 1 |

**Table 5. Verdicts (both named cells; ten paired seeds; 95 % intervals).**

| question | contrast | acts_only × event | neither × event | verdict |
|---|---|---|---|---|
| Authority | C4 − C2 superiority | +0.76 (0.56 to 0.92) | +0.71 (0.56 to 0.84) | supported |
| Correctness | C4 − C5 dependence | +0.90 (0.76 to 1.00) | +0.91 (0.78 to 1.00) | supported |
| Restoration | C4 − C0's within (73 of 82), same notes | +0.05 (−0.03 to +0.16) | +0.06 (0.00 to +0.16) | supported |
| Authority against translation | C4 − C1 superiority | +0.13 (0.00 to 0.32) | +0.12 (0.00 to 0.29) | other |
| Translation against authority | C1 − C4 superiority | −0.13 (−0.32 to 0.00) | −0.12 (−0.29 to 0.00) | other |
| Advisory alone | C2 − C0 improvement | +0.09 (0.01 to 0.18) | +0.21 (0.09 to 0.35) | other |
| Translation alone | C1 − C0 improvement | +0.71 (0.49 to 0.91) | +0.80 (0.60 to 0.95) | supported |

**Table 6. Outcomes behind the rates, uninformed and informed cells.**

| condition and cell | success | wrong member | abstained | refused | other |
|---|---:|---:|---:|---:|---:|
| C2 advisory, acts_only | 16 | 56 | 10 | 0 | 0 |
| C2 advisory, neither | 18 | 56 | 6 | 0 | 2 (wrong op) |
| C2 advisory, cold | 60 | 0 | 9 | 2 | 11 (malformed 9, wrong op 2) |
| C4 authoritative, acts_only | 76 | 1 | 1 | 4 | 0 |
| C4 authoritative, neither | 77 | 1 | 0 | 4 | 0 |
| C4 authoritative, cold | 74 | 1 | 0 | 6 | 1 (wrong op) |
| C5 shuffled, acts_only | 1 | 80 | 1 | 0 | 0 |
| C5 shuffled, neither | 1 | 81 | 0 | 0 | 0 |
| C5 shuffled, cold | 1 | 81 | 0 | 0 | 0 |
| C5 shuffled, within | 1 | 81 | 0 | 0 | 0 |

The single success in each C5 cell is the same exchange, seed 111's, in
which the sender's note described the wrong member (a faithful miss),
so the shuffled correspondence happened to name the task target: that
was compliance with the shuffled mapping, not resistance to it; mapping
fidelity and task correctness are different things.

Usage per completed exchange, from the archive: input about 6,050
tokens in every condition (the payload adds about 40); output tokens
with reasoning tokens in brackets: C0 72 (49), C1 66 (43), C2 102 (80),
C4 39 (16), C5 40 (17). Reasoning tokens under the authoritative
sentence were about a fifth of those under the advisory one, roughly
an 80 % reduction; a descriptive accompaniment, not evidence that
verification stopped.

### 9.4 Restoration and systematic misdirection, together

The same correct correspondence, presented as advisory, left the
uninformed receivers near the boundary (16 and 18 of 82) with wrong
members dominant (56 and 56); presented as authoritative, it restored
them to the informed baseline (76 and 77 of 82, against the partner's 73
under C0, the comparator the analysis uses; the partner reached 77 under
C4 itself). The authority verdict is supported in both named cells with
intervals well clear of the margin, on a payload held constant to the
sentence.

The same authoritative instruction with the correspondence shuffled
redirected 80 to 81 of 82 receivers in every cell, and 81 of 82 in the
partner cell, where the receiver's own record supported the correct
answer. Correct prior information did not protect the receivers'
answers against a shuffled authoritative mapping. Whether they
evaluated the conflict and nevertheless prioritised the instruction,
or did not evaluate it, is not something this design observes. C4 − C5
is the dependence on mapping correctness under the instruction; it is
not evidence of binding in any mechanical sense.

Translation and authority are close in the named cells and neither
direction met the superiority criterion there (C4 − C1 of +0.12 to +0.13 with lower bounds at
zero; no equivalence claimed). They differ where the receiver has no
history: re-expression achieved 0 of 82 for cold receivers, the
correct authoritative correspondence 74 of 82. Re-expression relies on
context these formed receivers possess. The advisory result is small
but above zero (+0.09 and +0.21; not the 0.25 margin): having the
mapping, without an instruction to prefer it, is worth little to a
receiver with a competing story and a great deal to one without.

# Discussion, methods and evidence

## 10. Discussion

### 10.1 History and transfer

**What is demonstrated.** On one model configuration, history-grounded
reference formed in every attempted history, resolved for the partner,
reversed for an outsider with a different history, and, in a crossed
design with all floors met, followed the dimension of history it
depended on: events transferred to any receiver holding the event,
acts substantially to receivers holding the act, with a residual
partner advantage on act references. Shared history explains
substantial transfer across agent identity; identity independence is
not established. Production of such references occurs in other
models and is configuration-dependent. Explicit resolution competence
is separable from spontaneous production: Qwen 3.8 Max at both settings
and Claude Sonnet 5's probed seeds passed the explicit gates and
produced few or no history-grounded references. Gate abstention (Opus
4.8; Sonnet 5 seed 79) is a third behaviour and does not demonstrate
resolution competence.

**What is not.** Cross-family dimension specificity is not
established: one family formally, a second descriptively on the event
dimension only. Nothing here establishes that an explicit belief graph
or a non-linguistic architecture is required to achieve the behaviour.
The recency-indexed act gap is exploratory. The provider's model change
under a stable name, detected by per-call provenance, is a limit on
any claim that names a model rather than an archived configuration.

**Proposed consequences, distinguished from demonstrated behaviour.**
For the programme's architecture, the demonstrated behaviour is that an
agent's interaction history can become part of what a message requires
for successful interpretation, and that the requirement is
dimension-specific. The proposed consequence is that reference
infrastructure shared across agents must carry both dimensions of the
history it is meant to bridge; what it takes for such infrastructure
to be used when a receiver holds a competing history is the subject of
Part II.

### 10.2 Authority and repair

**What is demonstrated.** In this setting, authority makes a supplied
correspondence govern interpretation; it does not ensure that
correspondence is true. The same intervention produced both
restoration and coordinated error, on the same receivers and notes,
with the payload held constant. A17's descriptive pattern, that
availability alone was often ineffective for receivers holding a
competing history, is the reason the comparison was worth making; A18
supplies the supported contrasts. We do not present "availability
versus authority" as a fully isolated mechanism: A18's design
separates the authority sentence from the advisory one and the correct
mapping from the shuffled one; it does not separate every factor that
differed between A17's delivery channels and A18's.

**Registry correctness as an upstream dependency.** For this
intervention, the correctness of the supplied correspondence is a
critical upstream dependency: with the correspondence wrong, the
instruction produced near-total misdirection, and the receivers' own
correct records did not protect them. A18 observes the receivers' responses, not whether they checked
anything; it does not show that receivers cannot check authority,
provenance or consistency under a different design.

**Proposed engineering consequences, distinguished from demonstrated
behaviour.** The programme's motivating architecture has a Rules plane
that is meant to bind interpretation across agents with divergent
histories. A18 gives that plane a demonstrated mechanism and a
demonstrated failure mode on one carrier: an explicit authority
instruction over a registry correspondence transfers whatever the
registry holds. The engineering consequence we propose, not
demonstrate, is that an authoritative lookup needs its own way to
establish that its entries deserve authority; receiver compliance
cannot supply that assurance. Whether an executor that mechanically
constrained target selection would behave the same, whether receivers
could be given grounds to check authority, and whether the mechanism
holds across families are open.

**The failure class, placed beside its neighbour.** The act that fails
here is procedurally valid: the receiver holds the authority to act on
either cistern, the world accepts the act, and the task fails because
the act reached the wrong one of two otherwise valid objects. That is
a different failure class from the one measured in our work on an
executable institution (*Where Reliability Lives*, arXiv:2609.03192,
§5.4), where one falsehood delivered as trusted testimony drove
roughly nine hundred futile acts per believing run and the
authoritative ledger refused every one of them. The two results come
from different apparatuses. Together they illustrate two distinct
failure classes, a false premise contained by authority and a wrong
referent that authority does not see; they do not demonstrate that
one deployed gate admits reference errors, and no study has run a
reference error and an authority check on the same apparatus. What
the two records support is narrower: an authority check adjudicates
whether an actor may act on an object, and nothing in either apparatus
adjudicated whether the object was the one the sender meant. The
engineering consequence we propose, not demonstrate, is that this
binding is a check distinct from authority, provenance and
admissibility, and that where it lives in an agent architecture has
to be established rather than assumed.

**Limits.** One carrier, one corpus of previously examined cases, one
instruction pair frozen verbatim without alternatives; the event
dimension only; two draws a probe in the underlying corpus; the CLI
pathway reports no answering model per call; usage is tokens on a
subscription, not a bill; A17's descriptive findings are descriptive.

### 10.3 Taken together

Part I shows, on one family with every floor met, that the failure at
the history boundary is a systematic wrong choice specific to the
dimension of history a reference depends on. Part II shows, on a
second family and one corpus of such cases, what bridging that
boundary took: not the availability of a correct correspondence, which
was often ineffective against a competing history, but an instruction
to prefer it, and the same instruction transferred a wrong
correspondence just as completely. The two results are bounded
separately and neither is claimed for models or tasks not run. The
broad ideas are in the literature; what this programme adds is the
controlled separation of which history supports which reference, the
measured form of the failure at the boundary, and the measured
difference between advisory and authoritative presentation of one
correspondence, under preregistered readings with the deviations and
audits in the record.

## 11. Methods

### 11.1 Discipline, statistics and pathways (both parts)

**Discipline.** Four kinds of record, kept apart: the initial
preregistration of each study, committed before its first call;
prospective amendments, each committed before the calls it governs;
deviations found in flight, disclosed and dated when found (A15.4,
A17.1, A17.2); and retrospective corrections and audits after results,
recorded with the original outputs preserved (A14.2's classifier
corrections; the 2026-09-12 qualification audit). Then: the seed as
the replication unit; fixed N; every attempted seed accounted for,
including failed and restarted attempts; readings computed
mechanically in a fixed precedence with floors; construction verified
from the generated histories at launch after A15.4; accepted model
identifiers frozen and the resolved model archived per call; spend or
call ceilings enforced by the drivers; per-seed summaries decide
nothing; each report run once.

**Statistics.** Seed-level estimates with percentile-bootstrap
intervals (10,000 resamples, generator seed 20260909); exact
Clopper–Pearson intervals for seed counts; contrasts on paired
seed-level differences; adequacy floors stated per study.

**Pathways.** DeepSeek through a proxy with the resolved model
archived per call (A13 to A16). Claude models through Claude Code's
headless mode on a subscription login, tools and settings off, a
neutral working directory, non-essential traffic disabled, the full
model identifier pinned after alias drift was observed, the answering
model and thinking tokens archived per call. Qwen through a proxy with
the reasoning setting passed as an option. gpt-6-astra through Codex
CLI's non-interactive mode with its tool surfaces disabled; the event
stream reports no model per call, stated as a limit.

### 11.2 Part II specifics

**Carrier and pathway.** gpt-6-astra, `model_reasoning_effort="medium"`,
Codex CLI 0.154.0 (`codex exec --json`) on a subscription login, the
experiment's system prompt replacing the CLI's instructions, 33 tool
surfaces disabled by name, read-only sandbox in an empty directory,
ephemeral sessions, no history, user configuration ignored, the
memories feature off by default. Each call archived with input, cached
input, output and reasoning tokens and wall time; the event stream
carries no model string, so the requested identifier is the accepted
identifier and this limit is stated.

**Apparatus.** The waterworks world; the crossed design with
groups A to E, relations computed from the generated schedules and
asserted before every probe; the post-formation competence gate with
the evaluability rule; the A14.2 attribution classifier; the v2
coherent swap and the irrelevant-message control in R0.

**Pipeline.** For every frozen note, the archived note-level
classification (described member or unresolved) mapped to the
receiver's tag and, for re-expression, to the receiver's group's events
for the member and the receiver's recorded acts on it (two most
recent, chronological; abstentions and off-target acts excluded). A
note describing the other member is repaired to the member described
and scored as a faithful miss. Unresolved notes deliver nothing under
R1, R2, R3 and R3′; R2r's block is present regardless; none of A18's
82 notes is unresolved.

**Analysis.** Seed-level success per cell; paired seed differences over
seeds where both conditions completed; percentile bootstrap, 10,000
resamples, generator seed 20260909; predicates: improvement and
dependence mean ≥ 0.25 with lower bound > 0 (contradicted when the
upper bound ≤ 0); restoration lower bound > −0.15 (contradicted when
the upper bound < −0.15); superiority mean ≥ 0.10 with lower bound > 0
(contradicted when inferiority is established); equivalence whole
interval within ±0.10 (A17 only); inconclusive when fewer than four
paired seeds or a floor is unmet; else other. Cells combine
inconclusive → contradicted → supported → other. A17 floors: 24 event
and 24 act notes, four seeds with both. A18 floor: 48 event-only notes
over six seeds carrying four each, on each contrast's paired material.

**Execution rules.** Sequential seeds, four workers; every attempt
archived and counted against the ceiling; a pass ending without its
record kept, the CLI awaited, restarted once, then recorded incomplete;
a standing five-minute health check; the analysis run once after every
pass; per-seed summaries decide nothing.

## Evidence appendix

### A. Part I

Commits are in the repository's history; every frozen text and
amendment is in `notes/panel/prereg-grounded-dialects-prepilot.md`.

**A13** preregistered at aac5d43; reserved seeds 11 and 12 recorded at
d316421; ruling f091285; `notes/panel/grounded-dialects-a13-result.md`;
records `a13-s10…12-deepseek`.
**A14** preregistered at 78fa0bb; A14.1 (rerun after a proxy outage)
b13489a; A14.2 (classifier corrections after the result, the
as-preregistered report preserved) in the preregistration; ruling
181d9bb; `grounded-dialects-a14-result.md`; records `a14-s13…28-deepseek`;
reports `a14-population.json`, `a14-population.all16.json`,
`a14-population.prereg.json`; every seed's gate numerators, formation
outcomes, notes, factual counts, swaps and transfer penalty tabulated
in the note.
**A15** preregistered at 3388de1; A15.1 (mechanical reading) 36c1b43;
A15.2 (the provider-side model change; requalification) 220ed3e; A15.3
(the structure study resumed on `deepseek-flash`) ecb3895; A15.4 (the
construction erratum) in the preregistration;
`grounded-dialects-a15-2-result.md`, `grounded-dialects-a15-result.md`;
records `a15-s29…59`; report `a15-structure.json`.
**A16** preregistered at 27eaecd; A16.1 a5a3542; A16.2 6ef7f67; A16.3
50be940; `grounded-dialects-a16-result.md` with the audit note appended
after acceptance; records `a16-s60…71-deepseek-flash`; report
`a16-crossed.json`.
**A17.0, the qualification set** preregistered at 30f3b1c (Qwen), with
A17.0c–e (Claude) in the preregistration, A17.0f 2cbd00e (gpt-6-astra),
A17.0g 289ae53 and A17.0g.1 f93e461 (effort); result notes
`grounded-dialects-a17-0-result.md`, `grounded-dialects-a17-0-claude-result.md`,
`grounded-dialects-a17-0f-result.md`; decision files
`a17pre-decision-*.json`; records `a17pre-s72…107`; the carrier ordering
and the multiplicity ruling in the preregistration; the qualification
audit (2026-09-12) `tools/audit_qualification.py`,
`a17pre-qualification-audit.json` and its addendum.
**A17 R0** frozen at 2b5f99e (implementation 86d8e06); A17.1 e20d5bc;
A17.2 49b44e6; A17.3 78537f1; `grounded-dialects-a17-result.md`;
`a17-r0-crossed.json`; records `a17-s108…119-codex-gpt-6-astra`.
The packaging ruling for A13 to A16: `grounded-reference-scope.md`.

### B. Part II

**A17.** Frozen text `notes/panel/a17-freeze-proposal.md` (39c9908,
after six reviewer corrections); preregistration entry A17 (2b5f99e),
implementation signature 86d8e06; amendments A17.1 (no pre-formation
gate from seed 109; e20d5bc), A17.2 (two draws; 49b44e6), A17.3 (the
weekly-cap pause and the dated resumption). Records
`results/dialects/history/a17-s108…119-codex-gpt-6-astra`, driver log
`a17-driver.log`; analysis `a17-repair.json` (every cell with success,
wrong member, abstention, unresolved, faithful miss and lookup use;
every contrast with mean, interval, paired seeds and verdict; usage per
condition); R0's crossed analysis `a17-r0-crossed.json`. Seed 109's
gate: a1, a2 answered 8 of 10 evaluable, zero wrong. Composition of a
probed seed: 120 history, 120 notebook, 180 gate, 108 probe calls (seed
108 carried 180 further qualification calls under A17.1).

**A18.** Frozen text `notes/panel/a18-design-proposal.md` v2.1
(eb0da6a); preregistration entry A18 (1090174); zero-call power
scenarios `tools/power_a18.py`. Records
`results/dialects/history/a18-s108…119-codex-gpt-6-astra` (seed 112
recorded incomplete at every stage with zero calls); driver log
`a18-driver.log`; analysis `a18-authority.json`. Per condition, 82
exchanges an arm on ten seeds carrying 12, 10, 6, 10, 5, 7, 11, 8, 7
and 6 event-only notes.

**Qualification audit (2026-09-12).** The carrier was selected by a
qualification whose implemented predicate omitted the written production
threshold and whose drivers applied a stricter invalid-act cutoff than
written; the audit (`tools/audit_qualification.py`,
`a17pre-qualification-audit.json`, the addendum in the preregistration)
finds gpt-6-astra qualifying at both efforts under the written rule and
no other configuration, so the carrier is unchanged.

**Reviewer rulings applied.** A17: the closure as inconclusive without
retrospective event-only verdicts; "descriptively reproduced" not
"replicates"; the shuffled index stated as correct selection reduced to
1 of 82 with outcomes kept apart. A18: the behavioural statement of the
shuffled-authority result; registry correctness as a critical upstream
dependency; the reasoning-token change as roughly 80 % and descriptive;
re-expression's 0 of 82 cold given prominence; the stopping point.

### C. The recorded prior-work search

Protocol with collapse conditions written before the search, every
query verbatim with date, the hits examined, the verdicts and the API
refusals: `grounded-reference-search-protocol.md`. Outcome: no
COLLAPSE on either claim; the near-misses are cited in section 2 with
the clause each falls short on. The absence claim in this manuscript
is bounded to that search on that date.

## References

Verified against their records on 12 September 2026; arXiv identifiers
where the work is a preprint or has one.

- Ashery, A. F., Aiello, L. M. and Baronchelli, A. (2025). Emergent
  social conventions and collective bias in LLM populations. *Science
  Advances* 11(20), eadu9368. arXiv:2410.08948.
- Brennan, S. E. and Clark, H. H. (1996). Conceptual pacts and lexical
  choice in conversation. *Journal of Experimental Psychology:
  Learning, Memory, and Cognition* 22(6), 1482–1493.
- Clark, H. H. and Brennan, S. E. (1991). Grounding in communication.
  In Resnick, L. B., Levine, J. M. and Teasley, S. D. (eds.),
  *Perspectives on Socially Shared Cognition*, APA, 127–149.
- Clark, H. H. and Wilkes-Gibbs, D. (1986). Referring as a
  collaborative process. *Cognition* 22(1), 1–39.
- Fu, Y., Qiu, R., Wang, X., Sansom, J., Ayyappa Prabhu, S., Tang, H.,
  Kim, J., Sohn, S. and Lee, H. (2026). Beyond blind following:
  evaluating robustness of LLM agents under imperfect guidance. *EACL
  2026*, long papers, 6591–6618.
- Gao, Yu, Deng, Li and Wang (2026). Testing interchangeability in
  LLM agent teams. arXiv:2609.05279.
- Hawkins, R. D., Franke, M., Frank, M. C., Goldberg, A. E., Smith,
  K., Griffiths, T. L. and Goodman, N. D. (2021). From partners to
  populations: a hierarchical Bayesian account of coordination and
  convention. *Psychological Review*. arXiv:2104.05857.
- Jones, C. R., Lombardi, A., Mahowald, K. and Bergen, B. K. (2026).
  LLMs and people both learn to form conventions, just not with each
  other. arXiv:2602.08208.
- Kim, J. (2026). Drawing with strangers: population scaling drives
  zero-shot mutual intelligibility in emergent sketching.
  arXiv:2606.10582.
- Leong, J. W. (2026). Recognition without enforcement:
  configuration-dependent failures in LLM agent instruction
  arbitration and external control. arXiv:2608.28502.
- Lewis, D. K. (1969). *Convention: A Philosophical Study*. Harvard
  University Press.
- Li, N., Gatt, A. and Poesio, M. (2026). Seeing is not sharing: some
  vision-language models overestimate common ground in asymmetric
  dialogue. *SIGDIAL 2026*, 694–710. arXiv:2606.31719.
- Li, N., Gatt, A. and Poesio, M. (2025). Grounded misunderstandings
  in asymmetric dialogue: a perspectivist annotation scheme for
  MapTask. arXiv:2511.03718.
- Mohapatra, B., Charlot, T., Duca, G., Palan, M., Romary, L. and
  Cassell, J. (2026). Frame of reference: addressing the challenges of
  common ground representation in situational dialogs. *Findings of
  ACL 2026*. arXiv:2601.09365.
- Schuster, J., Gautam, V. and Markert, K. (2026). Whose facts win?
  LLM source preferences under knowledge conflicts. arXiv:2601.03746.
- Shih, A., Sawhney, A., Kondic, J., Ermon, S. and Sadigh, D. (2021).
  On the critical role of conventions in adaptive human-AI
  collaboration. *ICLR 2021*. arXiv:2104.02871.
- Shih, B., Winnicki, J. and Cao, A. (2026). How do language models
  choose between context and memory? arXiv:2609.00753.
- Sun, K., Bai, F. and Dredze, M. (2026). Task matters: knowledge
  requirements shape LLM responses to context–memory conflict. *ACL
  2026*. arXiv:2506.06485.
- Talebirad, Y., Redman, E., Parsaee, A. and Zaiane, O. R. (2026).
  From signals to structure: how memory architecture drives language
  emergence in LLM agents. arXiv:2607.00233. And: Memory is
  communication: the frontier between remembering and signaling.
  arXiv:2608.17053.
- Wallace, E., Xiao, K., Leike, R., Weng, L., Heidecke, J. and Beutel,
  A. (2024). The instruction hierarchy: training LLMs to prioritize
  privileged instructions. arXiv:2404.13208.
- Wang, P.-Y. A., Mishra, C., Özyürek, A., Rubio-Fernández, P. and
  Ghaleb, E. (2026). Aligned but not partner-specific: distinguishing
  how multimodal LLM agents succeed in reference games without
  human-like conventions. arXiv:2606.08081.
- Wang, Z., Li, W., Kaliosis, P., Rambow, O. and Brennan, S. E.
  (2025). LVLMs are bad at overhearing human referential
  communication. arXiv:2509.11514.
- Yamamoto, T., Morita, J., Higashinaka, R. and Takeuchi, Y. (2026).
  Dissociating communicative success from representational alignment
  in common ground formation. *Frontiers in Computer Science*,
  10.3389/fcomp.2026.1873726.
- Yan, L., Li, R., Han, X., Li, W., Wang, B., Wang, L., Lyu, C. and
  Chen, G. (2026). Trust no tool: evaluating and defending LLM agents
  under untrusted tool feedback. arXiv:2605.17453.
- Yang, H., Song, W., Kim, T., Song, J., Park, S. and Jo, Y. (2026).
  Agents trust tools too much: measuring reliance on unreliable tools.
  arXiv:2609.05587.
- Yu, F., Seedat, N., Schwarz, J. R. and Bean, A. M. (2026). To whom
  do language models align? Measuring principal hierarchies under
  high-stakes competing demands. arXiv:2605.12120.
- Zeng, P., Li, W., Paige, A. J., Wang, Z., Kaliosis, P., Samaras, D.,
  Zelinsky, G., Brennan, S. E. and Rambow, O. (2026). LVLMs and humans
  ground differently in referential communication. arXiv:2601.19792.
- Zhang, C., Wan, Z., Yu, X., Zhou, P., Zhao, W., Wu, J., Zhou, Y. and
  Tsang, I. (2026). Don't blindly trust it: how unreliable feedback
  breaks tool-using LLM agents. arXiv:2606.21409.
