# History-grounded reference and transfer: reference that emerges through interaction, and the history it depends on

*First complete draft, 2026-09-12. Preregistered studies A13 to A16 and
the A17.0 qualification set of the grounded-dialects programme, with
A17's R0 as descriptive corroboration; records under
`results/dialects/history/`, frozen texts and every amendment in
`notes/panel/prereg-grounded-dialects-prepilot.md`. Written under the
wording rules of `grounded-reference-writeup-plan.md`.*

## Abstract

Two language-model agents that share twelve days of work in a small
world, each keeping its own notebook and record, come to refer to one
of two visually identical cisterns by what happened to it and by what
they did to it. We show, in a chain of preregistered studies on one
model family served under two configurations (the provider changed the
served model beneath the request name mid-chain; the second was
requalified), that such history-grounded reference forms in
every history attempted (16 of 16 behaviourally; 14 of 16 in the
preregistered denominator), that a note resolved by the partner
resolves for an outsider against the outsider's own different history
to the wrong cistern (seed-level transfer penalty 0.95, 0.90 to 0.98),
and that the scope of a reference follows the dimension of history it
depends on: in a crossed design, event references resolved for every
class of receiver whose history held the event (154 of 174) and for
none whose did not (0 of 174), while act references depended
substantially on shared act history (0.76 to 0.92 with it, 0.14 to
0.16 without) with a remaining partner advantage. A stranger whose group
ran matching planned schedules (recorded acts diverged from the plan on
58 of 420 pair-day acts) matched the partner on event references and
fell 15 points short on act references, so shared history explains substantial
transfer across agent identity without establishing identity
independence. A qualification set of five named models (Qwen 3.8 Max, Claude
Sonnet 5, Claude Haiku 4.5, Claude Opus 4.8, gpt-6-astra) in nine
configurations over twelve three-seed runs found production of history-grounded references in a
second family (gpt-6-astra, three of three seeds twice) and, below the
qualification bar, in a third (Claude Haiku 4.5: a pair-performance
advantage in five of six seeds with 35 of 72 notes historical, never
the written production threshold), and three
others that failed the registered qualification criteria at every
tested effort: Qwen 3.8 Max and Claude Sonnet 5 passing the explicit
competence gates while producing few or no such references, Claude
Opus 4.8 abstaining at the gates. The formal dimension-specific result is
within one family; a second family's event-dimension pattern was
descriptively reproduced, its formal verdict inconclusive because that
carrier writes almost no act references.

## 1. Introduction

Can a reference be private to the agents who share the history that
grounds it? Two agents who both saw a cistern leak can mean it by "the
one that leaked"; a third agent who did not see the leak, or who saw a
different cistern leak, cannot. That much is obvious in the abstract.
What is not obvious is whether language-model agents, given nothing
but a shared world and their own notebooks, will produce such
references spontaneously; whether the references will carry across
agent identity to a stranger with the same history; and whether the
failure at the history boundary is a degradation or a systematic
reversal. The programme reported here asked those questions with the
seed as the replication unit, every reading preregistered and computed
mechanically, and every attempted seed accounted for.

The programme addresses an empirical gap identified in the discussion
of our earlier work on divergent codebooks (*Coherence Is Not Truth*,
Taniwha AI, 2026, §13.5): that between two language models sharing a
surface language there is no learned, inspectable correspondence to
check or repair, and that whether history-dependent divergence of
meaning occurs in such agents at all was unmeasured. The studies here
measure that occurrence and its boundary. The learned correspondence
and the diagnostic that discussion names remain unbuilt and untested,
and nothing here should be read as evidence about them.

## 2. Related work and positioning

The broad ideas here are not new, and the contribution is narrower
than a first reading of the results suggests. Three recent papers,
found in a targeted search on 12 September 2026 and verified against
their arXiv records, materially shape the positioning; a recorded
search against the specific claims, run the same day under a protocol
whose collapse conditions were written before the search
(`grounded-reference-search-protocol.md`), adds the nearest neighbours
listed after them.

- **Testing Interchangeability in LLM Agent Teams** (Gao, Yu, Deng,
  Li and Wang, arXiv:2609.05279, September 2026) swaps agents between
  independently formed teams and finds that task performance holds
  while communication per unit of progress rises, because agents form
  partner-specific conventions and coordination patterns. That "shared
  interaction history affects communication" is therefore established
  independently of this programme. What A16 adds is a controlled
  separation of *which* components of a shared history support a given
  kind of reference: shared events against shared acts, crossed on
  purpose, with the recorded limitation that act sharing was assigned
  through planned schedules and the recorded acts diverged on 58 of
  420 pair-day acts.
- **Frame of Reference: Addressing the Challenges of Common Ground
  Representation in Situational Dialogs** (Mohapatra, Charlot, Duca,
  Palan, Romary and Cassell, arXiv:2601.09365, January 2026) studies
  how language models represent and use common ground to resolve
  relational references in situated dialogue, and improves it by
  training. Our question is different in kind: not whether a memory
  representation helps a model resolve references, but what happens to
  a reference when the histories that ground it are made to diverge
  under control. The transfer penalty (A14) and the wrong-member
  outcomes (A14, A16) are the additions: an outsider does not merely
  fail to answer, it resolves the reference against its own different
  history to the correspondingly wrong object, which distinguishes an
  incompatible interpretation from an inability to interpret.
- **To Whom Do Language Models Align? Measuring Principal Hierarchies
  Under High-Stakes Competing Demands** (Yu, Seedat, Schwarz and Bean,
  arXiv:2605.12120, May 2026) changes the principal who endorses a
  demand while holding content constant, and finds compliance despite
  demonstrated knowledge of the relevant standards. "Authority can
  override correct knowledge" is therefore also not a claim this
  programme can make as its own; it bears on the companion piece.

**Nearest neighbours from the recorded search**, each cited with the
clause of the claim it falls short on. *Aligned but Not
Partner-Specific* (Wang, Mishra, Özyürek, Rubio-Fernández and Ghaleb,
arXiv:2606.08081, June 2026) breaks partner history with a pseudo-dyad
baseline in repeated reference games and finds that multimodal agents
coordinate without partner-specific convention; history there is
present or broken as a whole, and no wrong-referent outcome is
measured for a receiver with a different history. *On the Critical
Role of Conventions in Adaptive Human-AI Collaboration* (Shih,
Sawhney, Kondic, Ermon and Sadigh, ICLR 2021) separates rule-dependent
from convention-dependent representation so that agents adapt to new
partners; it is the nearest in spirit to a component separation, but
it does not cross components of one interaction history across
receivers of a single reference. Lewis-signalling studies with private
notebooks (Talebirad, Redman, Parsaee and Zaiane, arXiv:2607.00233 and
arXiv:2608.17053, 2026) vary memory architecture, not history
components. Studies of models as overhearers of asymmetric human
dialogue (Li, Gatt and Poesio, SIGDIAL 2026, arXiv:2606.31719, and
arXiv:2511.03718; Wang et al., arXiv:2509.11514) find that models
conflate potential with established common ground; the artificial
receivers there do not form the histories, and the manipulation is of
context access. Director–matcher and mixed-dyad designs (Zeng et al.,
arXiv:2601.19792; Jones, Lombardi, Mahowald and Bergen,
arXiv:2602.08208) and cross-population intelligibility (Kim,
arXiv:2606.10582) vary partner type or population, not history
components. Emergent conventions in LLM populations (Ashery, Aiello
and Baronchelli, Science Advances 2025) and the hierarchical account
of partner-specific against community conventions (Hawkins et al.,
Psychological Review 2021, arXiv:2104.05857) are background.

The classical background is the collaborative view of reference
(Clark and Wilkes-Gibbs, 1986, Cognition 22(1), 1–39; Brennan and
Clark's conceptual pacts, 1996, Journal of Experimental Psychology:
Learning, Memory, and Cognition 22(6), 1482–1493; Clark and Brennan on
grounding, 1991, in Perspectives on Socially Shared Cognition, APA,
127–149) and Lewis's account of convention (1969, Convention: A
Philosophical Study, Harvard University Press), in which referring
expressions are pacts formed in interaction and interpretable by those
who share the history of their formation; the emergent-communication
literature in multi-agent learning studies the same phenomenon in
trained agents. These four were bibliographically verified against
their records on 12 September 2026.

**What is distinctive here, and what is not.** Not distinctive: that
shared history matters to communication, or that agents benefit from
common ground. Distinctive, on the evidence of this targeted search: the
crossed separation of the history needed to interpret different kinds
of reference (events against acts), under preregistered readings with
adequacy floors, on a task where the two referents are
attribute-identical so that history is the only available ground; and
the outsider's systematic wrong-member interpretation as the measured
form of the boundary. We found no prior study addressing these
specific contrasts in the sources searched, as of 12 September 2026
(the protocol records every query, every hit examined, and two APIs
that refused); that does not establish originality or uniqueness, and
a claim of absence here is a claim about that search on that date.

## 3. The apparatus

**The world.** A yard of six cisterns hosted by the waterworks
simulator: four with distinct attributes and a matched pair of large
timber cisterns by the smithy, identical in every attribute a receiver
can see. Agents act by depositing, drawing, inspecting or disposing;
the clerk records the outcome; levels and conditions change.

**Formation.** A group of two agents works twelve scripted days. On
five off-pair days the schedule puts events on the pair: one member
leaks (Day 3) and is repaired (Day 4); the other is tainted (Day 6)
and scoured (Day 7). On seven pair days the agents act on the pair
members by turn. Each agent rewrites its notebook each evening and
reads the clerk's record of its own receipts. Which member leaked and
which was acted on differ between groups by design, so histories
diverge in two dimensions: the events the pair suffered and the acts
performed on it.

**The tag interface.** Each agent sees the cisterns under its own
fixed two-letter tags, private to it; a note cannot use the receiver's
tags. A reference to a pair member must therefore be by attribute
(useless, since the pair is identical), by position (excluded), or by
history.

**Probes.** After formation, each sender is set a task on one pair
member and writes a note to "the hand you are working with today". The
same note goes unchanged to a set of receivers: the partner
(*within*), receivers from other groups in defined history relations,
and a receiver with no history (*cold*). The receiver acts; the act is
scored against the world target.

**Gates.** Before probes, every agent answers twelve explicit questions
about its own record ("which of your tags leaked on Day 3?"). A seed
proceeds only if every agent passes with an evaluability rule that
removes questions whose act did not happen. This separates explicit
resolution competence from spontaneous production.

**Classifier and readings.** A frozen classifier attributes each
note's reference to the member it describes against the sender's
actual history (event, act, mixed, unresolved; factual or a false
event attribution). Readings are fixed before each study: the
"advantage" in a seed (production in at least 10 of 12 notes, partner
minus cold at least 6, outsider at most 6); transfer penalties as
seed-level error differences with percentile-bootstrap intervals
(10,000 resamples, generator seed 20260909); in the crossed design,
contrasts with adequacy floors and a reading in a fixed precedence.
Causal-uptake controls: a coherent swap of the reference clause to a
template reference to the other member, and an irrelevant note.

**Configuration and provenance.** A13 to A16 ran on DeepSeek through
a proxy (`deepseek-v4-flash`, then `deepseek-flash` after the provider
changed the served model beneath the request name on 10 September
2026, detected by the resolved model archived per call and handled by
requalification). Every call is archived with its resolved model;
seeds with any call under an unaccepted identifier are excluded from
the preregistered denominator by a frozen rule.

## 4. A13 and A14: reference forms, and the boundary reverses it

**A13** (development seed 10, then two reserved seeds) established
that the configuration produces partner-indexed historical references
with the budget bounded: seed 10 produced 12 of 12 historical
references, 11 of 12 with production fidelity, and 12 of 12 faithfully
interpreted by the partner; both reserved seeds succeeded. It closed as
a successful prerequisite test.

**A14** fixed sixteen confirmatory seeds in advance. The shared-history
advantage formed behaviourally in all sixteen; the preregistered
denominator is **14 of 16** (0.875; exact interval 0.62 to 0.98), two
seeds being excluded because one call in each was recorded under a
namespaced identifier; the adjudicated equivalence reading is 16 of 16
(0.79 to 1.00). Every seed passed both gates with zero wrong answers
(0 wrong of 1,536 qualification items and 0 at the post-formation
gate), produced historical references in at least 10 of 12 probe
notes (186 of 192 overall), and showed the advantage. Of the 192 notes,
4 were empty at the budget after the re-ask; the attribution classifier
finds 177 factual, 3 carrying a false event attribution and 8 it cannot
check (negative act references, true of the target on inspection), the
last group including the 2 notes the history classifier does not count
as historical.

Crossing the history boundary is a reversal, not a degradation. Over
the fourteen seeds the **transfer penalty** (outsider error minus
partner error) was **0.95 (0.90 to 0.98; minimum 0.75)**: across 168
outsider acts the outsider chose the wrong member 159 times, abstained
6 times, produced no parseable act twice and was right once. The cold
penalty over the same fourteen seeds was the same size, 0.95 (0.92 to
0.98). On the sixteen-seed behavioural reading, cold receivers abstained
190 times in 192, the coherent swap moved the partner's choice in 104
of 104 eligible constructions, and the irrelevant note left the pair
untouched 192 of 192. Three false event attributions and the
partner's faithful-but-wrong or refusing responses to them are recorded
individually.

## 5. A15: does the boundary have structure?

A15 added a third group sharing A's events with mirrored acts, to ask
whether transfer follows events, acts, or the whole history. A15.2
requalified the changed model (3 of 3 seeds formed; not a population
estimate). A15.3's reading was **"other"**: two of three conditions
held, and the third failed through a construction error that made the
"foreign" group share the second sender group's acts (erratum A15.4).
Its exploratory split by reference type suggested the specific
hypothesis A16 then tested: that transfer follows the dimension of
history the reference depends on.

## 6. A16: the scope of a reference follows the dimension of history it depends on

**Design.** Five groups. Relative to the senders' group A, group E
witnesses A's pair events and performs A's pair acts (*both*: a
stranger with an identical history); C shares events only; D acts
only; B neither. Every relation was computed from the generated
schedules, written to the driver log before the first call, and
asserted by the runner before every probe. Twelve fresh seeds; three
draws of six probes; each note to the partner, an E, C, D and B member
and a cold receiver. Amendment A16.1 removed the pre-formation gate
with its consequences stated; A16.2 and A16.3 raised the spend ceiling
by ruling before the second batch, no outcome read.

**Execution.** Every one of the 6,994 calls resolved to
`deepseek-flash`; no seed route-substituted, lost or restarted; all
twelve reached the probes; no wrong answer in 2,064 answerable gate
items; the advantage formed in 12 of 12 (0.74 to 1.00).

**The reading: dimension-specific**, in the precedence inconclusive →
unexplained → dimension-specific → identity → content-general → other,
with every check printed beside the verdict:

| check | value | verdict |
|---|---|---|
| adequacy: ≥ 24 event-only and ≥ 24 act-only notes, ≥ 4 seeds with both | 58 event-only, 111 act-only, 12 seeds | met |
| receiver sharing nothing above 0.25 overall | 23 of 216 = 0.11 | no |
| event notes: events_only − acts_only ≥ 0.5, interval excludes 0 | 0.92 (0.86 to 0.97), seed minimum 0.75 | yes |
| act notes: acts_only − events_only ≥ 0.5, interval excludes 0 | 0.60 (0.40 to 0.76), seed minimum −0.10 | yes |
| off-dimension rates at most 0.25 | 16 of 111 = 0.14; 0 of 58 = 0.00 | yes |
| receiver sharing both at least 0.5 on both pure types | 51 of 58 = 0.88; 85 of 111 = 0.77 | yes |

**Table 1. Relation × reference type, pooled over twelve seeds (216 notes).**

| receiver's relation to the sender | event-only notes | act-only notes | mixed notes | all 216 |
|---|---:|---:|---:|---:|
| within (the partner) | 52 of 58 | 102 of 111 | 34 of 38 | 196 |
| both (E: matching planned schedules, different agent; recorded acts diverged on 58 of 420 pair-day acts across groups) | 51 of 58 | 85 of 111 | 33 of 38 | 170 |
| events only (C) | 51 of 58 | 16 of 111 | 28 of 38 | 99 |
| acts only (D) | 0 of 58 | 84 of 111 | 1 of 38 | 88 |
| neither (B) | 0 of 58 | 18 of 111 | 1 of 38 | 23 |
| cold | 0 of 58 | 0 of 111 | 0 of 38 | 0 |

The 216 notes are 58 event-only, 111 act-only, 38 mixed, 8 the
classifier resolves to nothing and 1 empty at the budget; the nine
outside the three type columns are inside the "all 216" column and
every denominator.

An event reference resolved for receivers whose history held the event
(partner, identical-history stranger, events-only stranger: 154 of 174,
each at about 0.9) and for no receiver whose history did not (0 of
174). Event overlap was necessary in these observations, not sufficient
(20 of 174 failed with it). An act reference resolved for receivers
whose history held the act (0.92, 0.77, 0.76) and rarely for those
whose did not (0.14, 0.16, 0.00). Mixed notes went with the event
dimension.

**Identity against content.** "Identical history" here means identical
planned schedules: the relation table is computed from the generated
schedules, and the recorded acts diverge from the plan on 58 of 420
pair-day acts (mostly a day-2 inspection made before any event
distinguishes the members, and a day-11 no-op), so a *both* receiver
can hold a day-2 act on the other member than its sender. With that
limit stated, the identical-history stranger matched
the partner on event references (seed-level both − within −0.00, −0.04
to 0.03) and fell short on act references (17 of 111, 15.3 points
pooled; −0.11, −0.19 to −0.03, seed-level). Shared history explains
substantial transfer across agent identity; identity independence is
not established. Exploratorily, the remaining act-reference gap sits on
recency-indexed references ("yesterday"), not on dated ones.

**The accepted wording.** A16 prospectively reproduced
dimension-specific reference resolution under crossed receiver
histories, with all preregistered checks and adequacy floors passing.
Event references transferred almost equally to the partner and an
identical-history stranger, and never resolved in receivers lacking
the relevant event history. Act references showed substantial
dependence on shared act history, with residual success without that
history and a remaining partner advantage. The result supports
history-dependent interpretation and substantial transfer across agent
identity, within the tested model and task.

**Audit limits, recorded after acceptance.** Thirteen of the 216 notes
have no unique classifier referent. The "shared acts" relation is
computed from the planned schedules; the recorded acts match the plan
on 362 of 420 pair-day acts, the divergence sitting where the schedule
makes it inevitable (a day-2 inspection before any event distinguishes
the members; a day-11 drawdown of a member already at the target
level). Thirteen asserted gate answers to unanswerable act questions are
recorded as false event attributions beside zero wrong scored answers.

## 7. The cross-family qualification set: production, qualification, and the unresolved transfer scope

To carry the programme beyond one family, five named models (Qwen 3.8
Max, Claude Sonnet 5, Claude Haiku 4.5, Claude Opus 4.8, gpt-6-astra)
in nine configurations over twelve three-seed runs were put through the same pair-design qualification
(three fresh seeds each; the same gates, classifier and advantage rule;
retained whatever the verdict). Three things must be kept apart:
**observed production** of history-grounded references, the
**qualification decision** under the registered rule, and the
**transfer scope**, which the qualification does not test.

**Table 2. The qualification set.**

| configuration | pathway | explicit gates (qualification; post-formation) | historical references in probe notes | premise rejection | seeds forming the advantage | decision (archived; and under the written rule where it differs) |
|---|---|---|---|---:|---:|---|
| Qwen 3.8 Max, thinking off | proxy | qualification 24 of 24 every agent, 0 wrong; post-formation 12 of 12 every agent, 0 wrong | 12 of 36 (low) | 0 | 0 of 3 under the written rule (1 of 3 on pair performance alone) | scientific result, no advantage |
| Qwen 3.8 Max, low effort | proxy | qualification: seed 75 one wrong answer (a1 23 of 24), seed 76 one abstention (a2 23 of 24), else 24 of 24; post-formation 10–12 of evaluable, 0 wrong | 13 of 36 (low) | 8 of 36 | 0 of 3 | scientific result, no advantage |
| Claude Sonnet 5, default | Claude Code | seed 79 failed the post-formation gate by two abstentions, 0 wrong; others passed | 0 of 24 (the third seed's production unmeasured after the gate stop) | 18 of 24 | 0 of 2 probed | capability failure (gate abstention) |
| Claude Sonnet 5, default, second sample | Claude Code | passed, 0–2 wrong | 1 of 36 (low) | 23 of 36 | 0 of 3 | scientific result, no advantage |
| Claude Sonnet 5, xhigh | Claude Code | passed, 0–1 wrong | 1 of 36 (low) | 25 of 36 | 0 of 3 | archived: capability failure (7 invalid acts against the drivers' cutoff of 6 in total); written rule: scientific result, no advantage (cutoff 18) |
| Claude Haiku 4.5, default | Claude Code | passed, 0 wrong, 0 abstentions | 14 of 36 (4, 6, 4) | 0 | pair performance 2 of 3; written rule 0 of 3 (production below 10 of 12 in every seed) | scientific result, no advantage |
| Claude Haiku 4.5, default, second sample | Claude Code | passed, 0 wrong, 0 abstentions | 21 of 36 (7, 7, 7) | 0 | pair performance 3 of 3; written rule 0 of 3 | archived: qualifies (implemented predicate), then ruled ineligible on multiplicity; written rule: scientific result, no advantage |
| Claude Opus 4.8, default (two samples) | Claude Code | abstentions at the gates, 0 wrong; one seed of six probed | 0 of 12 probed (five seeds unmeasured after gate stops) | 9 of 12 | 0 | capability failure (gate abstention) |
| Claude Opus 4.8, xhigh | Claude Code | qualification unevaluable by abstentions, 0 wrong | unmeasured (no seed probed) | — | 0 | capability failure (gate abstention) |
| gpt-6-astra, medium | Codex CLI | passed, 0 wrong, 0 abstentions | 35 of 36 (12, 11, 12) | 0 | 3 of 3 under the written rule | qualifies |
| gpt-6-astra, high | Codex CLI | passed, 0 wrong, 0 abstentions | 36 of 36 | 0 | 3 of 3 under the written rule | qualifies |

**Observed production.** History-grounded reference is produced beyond
DeepSeek: by gpt-6-astra at both efforts (35 and 36 of 36 notes;
partner 30 and 36 of 36 against cold 0; transfer penalties 0.50 to 1.00)
and, at a lower rate, by Claude Haiku 4.5 (14 and 21 of 36; a
pair-performance advantage, partner − cold ≥ 6 with outsider ≤ 6, in 5
of 6 seeds; the written production threshold of 10 of 12 in none). It
is not universal, and production and qualification are reported apart:
Qwen 3.8 Max produced at a low rate (12 and 13 of 36) without the
advantage; Claude Sonnet 5's probed seeds produced 0 of 24, 1 of 36 and 1 of 36
across three samples; Claude Opus 4.8's one probed seed produced 0 of
12; eight of its nine seeds across the three runs never reached the
probes, so their production is unmeasured. Two distinct
non-production behaviours are recorded and kept apart: **gate
abstention** (Opus 4.8 at every effort, and Sonnet 5's seed 79:
explicit null answers with reasoning that the agent's own tags cannot
be trusted across days, never a wrong answer) and **premise rejection
in the notes** (Sonnet 5 and Opus 4.8 wrote, in 18 of 24, 23 of 36, 25
of 36 and 9 of 12 notes, that the two members were indistinguishable
or that the receiver should act on either or both; none of 408 DeepSeek
notes does so).

**Qualification decisions, with an audit.** The registered rule requires
all three seeds of a run probed on evaluable gates and each forming the
advantage: production ≥ 10 of 12, partner − cold ≥ 6, outsider ≤ 6. An
audit after the manuscript drafts found that the implemented predicate
omitted the production threshold, and that the drivers applied the
invalid-act cutoff as more than 6 in total rather than more than 6 per
seed on average. Archived decisions are preserved; under the written
rule gpt-6-astra qualified at both efforts and no other configuration
did. Haiku's second sample had been decided as qualifying on the
implemented predicate and then ruled ineligible on multiplicity
grounds (its first sample had already failed; a rule under which any
later attempt could qualify would rise in probability with every
attempt); under the written rule it does not qualify at all (7 of 12 in
every seed), which leaves the ruling moot but recorded. Sonnet 5 at
xhigh, archived as a capability failure on the drivers' cutoff (7
invalid acts), is a scientific result without the advantage under the
written cutoff; it remains ineligible. The carrier for
the subsequent repair study was selected from the qualifying
configurations by an ordering precommitted before their results: tokens
a call including reasoning, then wall clock; gpt-6-astra at medium.

**Within-model effort.** Raising effort did not reliably make a
configuration suitable. Sonnet 5 at xhigh tripled its thinking tokens
and wrote the same notes; Opus 4.8 at xhigh thought thirteen times
longer and abstained more at the competence gate. gpt-6-astra produced
at both efforts with 18 and 43 reasoning tokens a call. A large
reported reasoning-token count is therefore not established as
necessary for production; deliberation length is not ruled out as an
influence, token counts are not directly comparable across pathways,
and why configurations differ remains open.

**The unresolved transfer scope.** The qualification tests production
and the pair-level advantage; it does not test dimension specificity.
That test was run once on the second family as the R0 stage of the
repair study (A17) on twelve fresh seeds: eleven probed (seed 109
stopped at its post-formation gate by abstentions, zero wrong), two
draws a probe rather than the frozen three (amendment A17.2), an extra
qualification gate on seed 108 that the frozen design did not carry
(A17.1), and a seven-hour pause at the plan's weekly usage cap resumed
by a dated decision (A17.3). Its formal crossed reading is
**inconclusive**: the carrier wrote 82 event-only, 10 act-only, 33
mixed and 7 unresolved notes in 132, below the act-note floor. Descriptively, the event-dimension pattern
was reproduced: event references resolved for receivers whose history
held the event (0.84, 0.87, 0.72 for partner, both and events_only)
and not for receivers whose did not (0.09, 0.06, 0.00), with the
preregistered event contrast at 0.65 (0.46 to 0.82) over the ten of the
eleven probed seeds that carried event-only notes (seed 112 carried
none). The act
dimension was neither descriptively reproduced nor found wanting: the
carrier produced too few act references (10 of 132 notes) to meet the
adequacy floor. The formal dimension-specific result therefore remains
within one family.

## 8. Discussion

**What is demonstrated.** On one model configuration, history-grounded
reference formed in every attempted history, resolved for the partner,
reversed for an outsider with a different history, and, in a crossed
design with all floors met, followed the dimension of history it
depended on: events transferred to any receiver holding the event,
acts substantially to receivers holding the act, with a residual
partner advantage on act references. Shared history explains
substantial transfer across agent identity; identity independence is
not established. Production of such references occurs in other
models and is configuration-dependent. Explicit resolution competence
is separable from spontaneous production: Qwen 3.8 Max at both settings
and Claude Sonnet 5's probed seeds passed the explicit gates and
produced few or no history-grounded references. Gate abstention (Opus
4.8; Sonnet 5 seed 79) is a third behaviour and does not demonstrate
resolution competence.

**What is not.** Cross-family dimension specificity is not
established: one family formally, a second descriptively on the event
dimension only. Nothing here establishes that an explicit belief graph
or a non-linguistic architecture is required to achieve the behaviour.
The recency-indexed act gap is exploratory. The provider's model change
under a stable name, detected by per-call provenance, is a limit on
any claim that names a model rather than an archived configuration.

**Proposed consequences, distinguished from demonstrated behaviour.**
For the programme's architecture, the demonstrated behaviour is that an
agent's interaction history can become part of what a message requires
for successful interpretation, and that the requirement is
dimension-specific. The proposed consequence is that reference
infrastructure shared across agents must carry both dimensions of the
history it is meant to bridge; what it takes for such infrastructure
to be used when a receiver holds a competing history is the subject of
the companion piece.

## 9. Methods

**Discipline.** Four kinds of record, kept apart: the initial
preregistration of each study, committed before its first call;
prospective amendments, each committed before the calls it governs;
deviations found in flight, disclosed and dated when found (A15.4,
A17.1, A17.2); and retrospective corrections and audits after results,
recorded with the original outputs preserved (A14.2's classifier
corrections; the 2026-09-12 qualification audit). Then: the seed as
the replication unit; fixed N; every attempted seed accounted for,
including failed and restarted attempts; readings computed
mechanically in a fixed precedence with floors; construction verified
from the generated histories at launch after A15.4; accepted model
identifiers frozen and the resolved model archived per call; spend or
call ceilings enforced by the drivers; per-seed summaries decide
nothing; each report run once.

**Statistics.** Seed-level estimates with percentile-bootstrap
intervals (10,000 resamples, generator seed 20260909); exact
Clopper–Pearson intervals for seed counts; contrasts on paired
seed-level differences; adequacy floors stated per study.

**Pathways.** DeepSeek through a proxy with the resolved model
archived per call (A13 to A16). Claude models through Claude Code's
headless mode on a subscription login, tools and settings off, a
neutral working directory, non-essential traffic disabled, the full
model identifier pinned after alias drift was observed, the answering
model and thinking tokens archived per call. Qwen through a proxy with
the reasoning setting passed as an option. gpt-6-astra through Codex
CLI's non-interactive mode with its tool surfaces disabled; the event
stream reports no model per call, stated as a limit.

## Evidence appendix

Commits are in the repository's history; every frozen text and
amendment is in `notes/panel/prereg-grounded-dialects-prepilot.md`.

**A13** preregistered at aac5d43; reserved seeds 11 and 12 recorded at
d316421; ruling f091285; `notes/panel/grounded-dialects-a13-result.md`;
records `a13-s10…12-deepseek`.
**A14** preregistered at 78fa0bb; A14.1 (rerun after a proxy outage)
b13489a; A14.2 (classifier corrections after the result, the
as-preregistered report preserved) in the preregistration; ruling
181d9bb; `grounded-dialects-a14-result.md`; records `a14-s13…28-deepseek`;
reports `a14-population.json`, `a14-population.all16.json`,
`a14-population.prereg.json`; every seed's gate numerators, formation
outcomes, notes, factual counts, swaps and transfer penalty tabulated
in the note.
**A15** preregistered at 3388de1; A15.1 (mechanical reading) 36c1b43;
A15.2 (the provider-side model change; requalification) 220ed3e; A15.3
(the structure study resumed on `deepseek-flash`) ecb3895; A15.4 (the
construction erratum) in the preregistration;
`grounded-dialects-a15-2-result.md`, `grounded-dialects-a15-result.md`;
records `a15-s29…59`; report `a15-structure.json`.
**A16** preregistered at 27eaecd; A16.1 a5a3542; A16.2 6ef7f67; A16.3
50be940; `grounded-dialects-a16-result.md` with the audit note appended
after acceptance; records `a16-s60…71-deepseek-flash`; report
`a16-crossed.json`.
**A17.0, the qualification set** preregistered at 30f3b1c (Qwen), with
A17.0c–e (Claude) in the preregistration, A17.0f 2cbd00e (gpt-6-astra),
A17.0g 289ae53 and A17.0g.1 f93e461 (effort); result notes
`grounded-dialects-a17-0-result.md`, `grounded-dialects-a17-0-claude-result.md`,
`grounded-dialects-a17-0f-result.md`; decision files
`a17pre-decision-*.json`; records `a17pre-s72…107`; the carrier ordering
and the multiplicity ruling in the preregistration; the qualification
audit (2026-09-12) `tools/audit_qualification.py`,
`a17pre-qualification-audit.json` and its addendum.
**A17 R0** frozen at 2b5f99e (implementation 86d8e06); A17.1 e20d5bc;
A17.2 49b44e6; A17.3 78537f1; `grounded-dialects-a17-result.md`;
`a17-r0-crossed.json`; records `a17-s108…119-codex-gpt-6-astra`.
The packaging ruling for A13 to A16: `grounded-reference-scope.md`.
