Research · Note
This note ships with
The note (Markdown)Overview — the plain-language readPart I as a bounded report: history-grounded reference and transfer (Markdown)Part II as a bounded report: authority and reference repair (Markdown)Recorded prior-work search — protocol and recordEvidence package — how to check and regenerate every figureEvidence manifest (SHA-256 identities)History-grounded reference in language-model agents: formation, transfer across divergent histories, and repair under authority
Research note, Taniwha AI, 12 September 2026. The integrated
manuscript of the grounded-reference programme (preregistered studies
A13 to A18), kept in preprint form as the collation of record and
published as a note; no preprint has been submitted. Generated by
tools/build_integrated_manuscript.py from the two bounded pieces,
history-grounded-reference-and-transfer.md (Part I) and
authority-and-reference-repair.md (Part II), at commit e562c96; the
pieces are the edited sources and every number here is theirs. Records
under results/dialects/history/; frozen texts and every amendment in
notes/panel/prereg-grounded-dialects-prepilot.md; the evidence
package, manifest and recorded prior-work search ship beside this
note. The internal record is private; the paths and commit identifiers
in the evidence appendix are the record’s identities, so that a reader
can check what was frozen when, and supervised access to the record can
be requested.
Abstract
Two language-model agents that share twelve days of work in a small world, each keeping its own notebook and record, come to refer to one of two visually identical cisterns by what happened to it and by what they did to it. Part I reports a chain of preregistered studies on one model family served under two configurations (the provider changed the served model beneath the request name mid-chain; the second was requalified): such history-grounded reference formed in every history attempted (16 of 16 behaviourally; 14 of 16 in the preregistered denominator); a note resolved by the partner resolves for an outsider, against the outsider’s own different history, to the wrong cistern (seed-level transfer penalty 0.95, 0.90 to 0.98); and in a crossed design the scope of a reference followed the dimension of history it depended on: event references resolved for every class of receiver whose history held the event (154 of 174) and for none whose did not (0 of 174), while act references depended substantially on shared act history (0.76 to 0.92 with it, 0.14 to 0.16 without) with a remaining partner advantage. A qualification set of five named models in nine configurations over twelve three-seed runs found production of history-grounded references in a second family (gpt-6-astra, three of three seeds twice under the written rule) and, below the qualification bar, in a third (Claude Haiku 4.5), while three others failed the registered criteria at every tested effort. Part II asks, on the second family, what a correct correspondence supplied by the apparatus needs in order to repair the failure at the history boundary. In the repair study (A17) every formal verdict was inconclusive because the carrier produced almost no references of one of the two history dimensions the floors required; descriptively, on 82 event references, re-expressing the reference in the receiver’s own history lifted uninformed receivers from about 7 percent to about 75 percent success, while a correct correspondence merely made available was used by receivers with no history and was often ineffective for receivers holding a competing one. A second study (A18) on the same formed receivers and notes delivered one identical registry payload under two instructions differing in one sentence. Called advisory, the correct correspondence left uninformed receivers near the boundary (16 and 18 of 82); called authoritative, it restored them to the informed baseline (76 and 77 of 82; +0.76 and +0.71, supported). The same authoritative instruction with the correspondence shuffled between the two candidates redirected 80 to 81 of 82 receivers in every cell, including the partner whose own record supported the correct answer (+0.90 and +0.91, supported). In these settings an agent’s interaction history becomes part of what a message requires to be understood; the requirement is specific to the dimension of history referenced; and authority makes a supplied correspondence govern interpretation without ensuring that it is true. The formal dimension-specific result is within one family; the second family’s event-dimension pattern was descriptively reproduced; Part II holds for one carrier, one corpus of previously examined cases and one frozen instruction pair.
1. Introduction
Can a reference be private to the agents who share the history that grounds it? Two agents who both saw a cistern leak can mean it by “the one that leaked”; a third agent who did not see the leak, or who saw a different cistern leak, cannot. That much is obvious in the abstract. What is not obvious is whether language-model agents, given nothing but a shared world and their own notebooks, will produce such references spontaneously; whether the references will carry across agent identity to a stranger with the same history; and whether the failure at the history boundary is a degradation or a systematic reversal. The programme reported here asked those questions with the seed as the replication unit, every reading preregistered and computed mechanically, and every attempted seed accounted for.
The programme addresses an empirical gap identified in the discussion of our earlier work on divergent codebooks (Coherence Is Not Truth, Taniwha AI, 2026, §13.5): that between two language models sharing a surface language there is no learned, inspectable correspondence to check or repair, and that whether history-dependent divergence of meaning occurs in such agents at all was unmeasured. The studies here measure that occurrence and its boundary. The learned correspondence and the diagnostic that discussion names remain unbuilt and untested, and nothing here should be read as evidence about them.
Part I establishes, on one family, that such references form and that a receiver whose history differs resolves them to the wrong cistern. That failure is the boundary Part II is about. A note that a partner resolves against the shared history resolves, for a receiver with a different history, against that different history, to the wrong cistern. If the apparatus knows the correct correspondence between the sender’s referent and the receiver’s own tag for the same physical cistern, what has to be done with that correspondence for the receiver to act on it? A17 asked the question as designed: does re-expression in the receiver’s own terms, evidence in prose, retained evidence, or a consultable index restore transfer, at what cost? A18 asked the narrower question that A17’s descriptive pattern exposed: does explicitly prioritising a supplied correspondence change target selection, holding everything else constant?
We report the two studies in order, with their evidential status kept distinct: A17’s verdicts are inconclusive under its preregistered floors and its findings are descriptive; A18’s verdicts are supported under its preregistered rule.
2. Related work and positioning
The broad ideas here are not new, and the contribution is narrower
than a first reading of the results suggests. Three recent papers,
found in a targeted search on 12 September 2026 and verified against
their arXiv records, materially shape the positioning; a recorded
search against the specific claims, run the same day under a protocol
whose collapse conditions were written before the search
(grounded-reference-search-protocol.md), adds the nearest neighbours
listed after them.
- Testing Interchangeability in LLM Agent Teams (Gao, Yu, Deng, Li and Wang, arXiv:2609.05279, September 2026) swaps agents between independently formed teams and finds that task performance holds while communication per unit of progress rises, because agents form partner-specific conventions and coordination patterns. That “shared interaction history affects communication” is therefore established independently of this programme. What A16 adds is a controlled separation of which components of a shared history support a given kind of reference: shared events against shared acts, crossed on purpose, with the recorded limitation that act sharing was assigned through planned schedules and the recorded acts diverged on 58 of 420 pair-day acts.
- Frame of Reference: Addressing the Challenges of Common Ground Representation in Situational Dialogs (Mohapatra, Charlot, Duca, Palan, Romary and Cassell, arXiv:2601.09365, January 2026) studies how language models represent and use common ground to resolve relational references in situated dialogue, and improves it by training. Our question is different in kind: not whether a memory representation helps a model resolve references, but what happens to a reference when the histories that ground it are made to diverge under control. The transfer penalty (A14) and the wrong-member outcomes (A14, A16) are the additions: an outsider does not merely fail to answer, it resolves the reference against its own different history to the correspondingly wrong object, which distinguishes an incompatible interpretation from an inability to interpret.
- To Whom Do Language Models Align? Measuring Principal Hierarchies Under High-Stakes Competing Demands (Yu, Seedat, Schwarz and Bean, arXiv:2605.12120, May 2026) changes the principal who endorses a demand while holding content constant, and finds compliance despite demonstrated knowledge of the relevant standards. “Authority can override correct knowledge” is therefore also not a claim this programme can make as its own; it bears on Part II.
Nearest neighbours from the recorded search, each cited with the clause of the claim it falls short on. Aligned but Not Partner-Specific (Wang, Mishra, Özyürek, Rubio-Fernández and Ghaleb, arXiv:2606.08081, June 2026) breaks partner history with a pseudo-dyad baseline in repeated reference games and finds that multimodal agents coordinate without partner-specific convention; history there is present or broken as a whole, and no wrong-referent outcome is measured for a receiver with a different history. On the Critical Role of Conventions in Adaptive Human-AI Collaboration (Shih, Sawhney, Kondic, Ermon and Sadigh, ICLR 2021) separates rule-dependent from convention-dependent representation so that agents adapt to new partners; it is the nearest in spirit to a component separation, but it does not cross components of one interaction history across receivers of a single reference. Lewis-signalling studies with private notebooks (Talebirad, Redman, Parsaee and Zaiane, arXiv:2607.00233 and arXiv:2608.17053, 2026) vary memory architecture, not history components. Studies of models as overhearers of asymmetric human dialogue (Li, Gatt and Poesio, SIGDIAL 2026, arXiv:2606.31719, and arXiv:2511.03718; Wang et al., arXiv:2509.11514) find that models conflate potential with established common ground; the artificial receivers there do not form the histories, and the manipulation is of context access. Director–matcher and mixed-dyad designs (Zeng et al., arXiv:2601.19792; Jones, Lombardi, Mahowald and Bergen, arXiv:2602.08208) and cross-population intelligibility (Kim, arXiv:2606.10582) vary partner type or population, not history components. Emergent conventions in LLM populations (Ashery, Aiello and Baronchelli, Science Advances 2025) and the hierarchical account of partner-specific against community conventions (Hawkins et al., Psychological Review 2021, arXiv:2104.05857) are background.
The classical background is the collaborative view of reference (Clark and Wilkes-Gibbs, 1986, Cognition 22(1), 1–39; Brennan and Clark’s conceptual pacts, 1996, Journal of Experimental Psychology: Learning, Memory, and Cognition 22(6), 1482–1493; Clark and Brennan on grounding, 1991, in Perspectives on Socially Shared Cognition, APA, 127–149) and Lewis’s account of convention (1969, Convention: A Philosophical Study, Harvard University Press), in which referring expressions are pacts formed in interaction and interpretable by those who share the history of their formation; the emergent-communication literature in multi-agent learning studies the same phenomenon in trained agents. These four were bibliographically verified against their records on 12 September 2026.
What is distinctive here, and what is not. Not distinctive: that shared history matters to communication, or that agents benefit from common ground. Distinctive, on the evidence of this targeted search: the crossed separation of the history needed to interpret different kinds of reference (events against acts), under preregistered readings with adequacy floors, on a task where the two referents are attribute-identical so that history is the only available ground; and the outsider’s systematic wrong-member interpretation as the measured form of the boundary. We found no prior study addressing these specific contrasts in the sources searched, as of 12 September 2026 (the protocol records every query, every hit examined, and two APIs that refused); that does not establish originality or uniqueness, and a claim of absence here is a claim about that search on that date.
Part II’s positioning. The general behaviour A18 measures, that a model given an authoritative instruction follows it even against its own correct knowledge, is expected from instruction following and is documented in the literature. To Whom Do Language Models Align? Measuring Principal Hierarchies Under High-Stakes Competing Demands (Yu, Seedat, Schwarz and Bean, arXiv:2605.12120, May 2026) changes the principal endorsing a demand while holding the demand’s content constant, across thousands of medical and legal scenarios, and finds that models abandon professional standards under diverging user instructions while demonstrably possessing the relevant knowledge. A general claim of novelty from “holding content constant while changing authority” is therefore untenable, and we do not make it.
What A18 measures that the general expectation does not is narrower and more specific to shared reference systems: the same correct correspondence between a sender’s history-grounded referent and the receiver’s own tag, presented under an advisory or an authoritative sentence on an otherwise identical payload; the receiver’s competing interpretation drawn from its own formed history rather than from a professional standard; the mapping correct or shuffled between two attribute-identical candidates; and receivers whose histories place them in different relations to the sender, including one that already knows the right answer. Its value is the controlled comparison and the measured boundary on this apparatus, not the discovery that instructions are followed.
Nearest neighbours from the recorded search (protocol and record
in grounded-reference-search-protocol.md, collapse conditions
written before the search), each with the clause it falls short on.
Context–memory conflict studies vary the stated authority of supplied
context against the model’s parametric knowledge: How Do Language
Models Choose Between Context and Memory? (Shih, Winnicki and Cao,
arXiv:2609.00753, September 2026) uses prompts that direct a model to
prioritise the supplied context or its own knowledge and locates
authority directions in activations; Task Matters (Sun, Bai and
Dredze, ACL 2026, arXiv:2506.06485) and Whose Facts Win? (Schuster,
Gautam and Markert, arXiv:2601.03746) vary task demands and source
credibility. None places the competing interpretation in a history the
receiver formed through interaction, and none is a referential
correspondence between two candidates. Tool-trust studies corrupt what
a tool returns: Agents Trust Tools Too Much (Yang, Song, Kim, Song,
Park and Jo, arXiv:2609.05587, September 2026) finds adoption of
corrupted returns above a third for every tool and user-prompt
policies (compare, verify, disclose) that do not consistently help;
Don’t Blindly Trust It (Zhang et al., arXiv:2606.21409) shows
misleading feedback can leave an agent worse off than no feedback;
Trust No Tool (Yan et al., arXiv:2605.17453) studies a tool that
earns trust before turning harmful. These vary correctness, and in one
case prompt policy, but not the stated authority of one identical
payload against an interaction-formed history. Beyond Blind
Following (Fu et al., EACL 2026, MIRAGE) varies the fidelity of
guidance from manuals, retrieval and prior interaction, and is the
nearest on the correctness manipulation; it does not present the same
guidance as advisory and as authoritative. The instruction hierarchy
(Wallace, Xiao, Leike, Weng, Heidecke and Beutel, arXiv:2404.13208)
and recognition-without-enforcement of instruction source (Leong,
arXiv:2608.28502) concern which source a model should obey, without a
correctness manipulation of a mapping. We found no prior study
addressing these specific contrasts in the sources searched, as of 12
September 2026; that is a claim about that search on that date, not
about the literature, and it does not establish originality.
3. The apparatus
The world. A yard of six cisterns hosted by the waterworks simulator: four with distinct attributes and a matched pair of large timber cisterns by the smithy, identical in every attribute a receiver can see. Agents act by depositing, drawing, inspecting or disposing; the clerk records the outcome; levels and conditions change.
Formation. A group of two agents works twelve scripted days. On five off-pair days the schedule puts events on the pair: one member leaks (Day 3) and is repaired (Day 4); the other is tainted (Day 6) and scoured (Day 7). On seven pair days the agents act on the pair members by turn. Each agent rewrites its notebook each evening and reads the clerk’s record of its own receipts. Which member leaked and which was acted on differ between groups by design, so histories diverge in two dimensions: the events the pair suffered and the acts performed on it.
The tag interface. Each agent sees the cisterns under its own fixed two-letter tags, private to it; a note cannot use the receiver’s tags. A reference to a pair member must therefore be by attribute (useless, since the pair is identical), by position (excluded), or by history.
Probes. After formation, each sender is set a task on one pair member and writes a note to “the hand you are working with today”. The same note goes unchanged to a set of receivers: the partner (within), receivers from other groups in defined history relations, and a receiver with no history (cold). The receiver acts; the act is scored against the world target.
Gates. Before probes, every agent answers twelve explicit questions about its own record (“which of your tags leaked on Day 3?”). A seed proceeds only if every agent passes with an evaluability rule that removes questions whose act did not happen. This separates explicit resolution competence from spontaneous production.
Classifier and readings. A frozen classifier attributes each note’s reference to the member it describes against the sender’s actual history (event, act, mixed, unresolved; factual or a false event attribution). Readings are fixed before each study: the “advantage” in a seed (production in at least 10 of 12 notes, partner minus cold at least 6, outsider at most 6); transfer penalties as seed-level error differences with percentile-bootstrap intervals (10,000 resamples, generator seed 20260909); in the crossed design, contrasts with adequacy floors and a reading in a fixed precedence. Causal-uptake controls: a coherent swap of the reference clause to a template reference to the other member, and an irrelevant note.
Configuration and provenance. A13 to A16 ran on DeepSeek through
a proxy (deepseek-v4-flash, then deepseek-flash after the provider
changed the served model beneath the request name on 10 September
2026, detected by the resolved model archived per call and handled by
requalification). Every call is archived with its resolved model;
seeds with any call under an unaccepted identifier are excluded from
the preregistered denominator by a frozen rule.
The carrier for both Part II studies (A17 and A18) is gpt-6-astra at medium reasoning effort, run through the Codex CLI’s non-interactive mode on a subscription login, with the CLI’s tool surfaces disabled and the experiment’s instructions replacing its own; every call is archived with its usage. That carrier was selected by a precommitted ordering from a qualification set of five named models in nine configurations over twelve three-seed runs (section 7); it produces history-grounded references at a high rate and, as it turned out, almost only event references.
Part I. History-grounded reference and transfer
4. A13 and A14: reference forms, and the boundary reverses it
A13 (development seed 10, then two reserved seeds) established that the configuration produces partner-indexed historical references with the budget bounded: seed 10 produced 12 of 12 historical references, 11 of 12 with production fidelity, and 12 of 12 faithfully interpreted by the partner; both reserved seeds succeeded. It closed as a successful prerequisite test.
A14 fixed sixteen confirmatory seeds in advance. The shared-history advantage formed behaviourally in all sixteen; the preregistered denominator is 14 of 16 (0.875; exact interval 0.62 to 0.98), two seeds being excluded because one call in each was recorded under a namespaced identifier; the adjudicated equivalence reading is 16 of 16 (0.79 to 1.00). Every seed passed both gates with zero wrong answers (0 wrong of 1,536 qualification items and 0 at the post-formation gate), produced historical references in at least 10 of 12 probe notes (186 of 192 overall), and showed the advantage. Of the 192 notes, 4 were empty at the budget after the re-ask; the attribution classifier finds 177 factual, 3 carrying a false event attribution and 8 it cannot check (negative act references, true of the target on inspection), the last group including the 2 notes the history classifier does not count as historical.
Crossing the history boundary is a reversal, not a degradation. Over the fourteen seeds the transfer penalty (outsider error minus partner error) was 0.95 (0.90 to 0.98; minimum 0.75): across 168 outsider acts the outsider chose the wrong member 159 times, abstained 6 times, produced no parseable act twice and was right once. The cold penalty over the same fourteen seeds was the same size, 0.95 (0.92 to 0.98). On the sixteen-seed behavioural reading, cold receivers abstained 190 times in 192, the coherent swap moved the partner’s choice in 104 of 104 eligible constructions, and the irrelevant note left the pair untouched 192 of 192. Three false event attributions and the partner’s faithful-but-wrong or refusing responses to them are recorded individually.
5. A15: does the boundary have structure?
A15 added a third group sharing A’s events with mirrored acts, to ask whether transfer follows events, acts, or the whole history. A15.2 requalified the changed model (3 of 3 seeds formed; not a population estimate). A15.3’s reading was “other”: two of three conditions held, and the third failed through a construction error that made the “foreign” group share the second sender group’s acts (erratum A15.4). Its exploratory split by reference type suggested the specific hypothesis A16 then tested: that transfer follows the dimension of history the reference depends on.
6. A16: the scope of a reference follows the dimension of history it depends on
Design. Five groups. Relative to the senders’ group A, group E witnesses A’s pair events and performs A’s pair acts (both: a stranger with an identical history); C shares events only; D acts only; B neither. Every relation was computed from the generated schedules, written to the driver log before the first call, and asserted by the runner before every probe. Twelve fresh seeds; three draws of six probes; each note to the partner, an E, C, D and B member and a cold receiver. Amendment A16.1 removed the pre-formation gate with its consequences stated; A16.2 and A16.3 raised the spend ceiling by ruling before the second batch, no outcome read.
Execution. Every one of the 6,994 calls resolved to
deepseek-flash; no seed route-substituted, lost or restarted; all
twelve reached the probes; no wrong answer in 2,064 answerable gate
items; the advantage formed in 12 of 12 (0.74 to 1.00).
The reading: dimension-specific, in the precedence inconclusive → unexplained → dimension-specific → identity → content-general → other, with every check printed beside the verdict:
| check | value | verdict |
|---|---|---|
| adequacy: ≥ 24 event-only and ≥ 24 act-only notes, ≥ 4 seeds with both | 58 event-only, 111 act-only, 12 seeds | met |
| receiver sharing nothing above 0.25 overall | 23 of 216 = 0.11 | no |
| event notes: events_only − acts_only ≥ 0.5, interval excludes 0 | 0.92 (0.86 to 0.97), seed minimum 0.75 | yes |
| act notes: acts_only − events_only ≥ 0.5, interval excludes 0 | 0.60 (0.40 to 0.76), seed minimum −0.10 | yes |
| off-dimension rates at most 0.25 | 16 of 111 = 0.14; 0 of 58 = 0.00 | yes |
| receiver sharing both at least 0.5 on both pure types | 51 of 58 = 0.88; 85 of 111 = 0.77 | yes |
Table 1. Relation × reference type, pooled over twelve seeds (216 notes).
| receiver’s relation to the sender | event-only notes | act-only notes | mixed notes | all 216 |
|---|---|---|---|---|
| within (the partner) | 52 of 58 | 102 of 111 | 34 of 38 | 196 |
| both (E: matching planned schedules, different agent; recorded acts diverged on 58 of 420 pair-day acts across groups) | 51 of 58 | 85 of 111 | 33 of 38 | 170 |
| events only (C) | 51 of 58 | 16 of 111 | 28 of 38 | 99 |
| acts only (D) | 0 of 58 | 84 of 111 | 1 of 38 | 88 |
| neither (B) | 0 of 58 | 18 of 111 | 1 of 38 | 23 |
| cold | 0 of 58 | 0 of 111 | 0 of 38 | 0 |
The 216 notes are 58 event-only, 111 act-only, 38 mixed, 8 the classifier resolves to nothing and 1 empty at the budget; the nine outside the three type columns are inside the “all 216” column and every denominator.
An event reference resolved for receivers whose history held the event (partner, identical-history stranger, events-only stranger: 154 of 174, each at about 0.9) and for no receiver whose history did not (0 of 174). Event overlap was necessary in these observations, not sufficient (20 of 174 failed with it). An act reference resolved for receivers whose history held the act (0.92, 0.77, 0.76) and rarely for those whose did not (0.14, 0.16, 0.00). Mixed notes went with the event dimension.
Identity against content. “Identical history” here means identical planned schedules: the relation table is computed from the generated schedules, and the recorded acts diverge from the plan on 58 of 420 pair-day acts (mostly a day-2 inspection made before any event distinguishes the members, and a day-11 no-op), so a both receiver can hold a day-2 act on the other member than its sender. With that limit stated, the identical-history stranger matched the partner on event references (seed-level both − within −0.00, −0.04 to 0.03) and fell short on act references (17 of 111, 15.3 points pooled; −0.11, −0.19 to −0.03, seed-level). Shared history explains substantial transfer across agent identity; identity independence is not established. Exploratorily, the remaining act-reference gap sits on recency-indexed references (“yesterday”), not on dated ones.
The accepted wording. A16 prospectively reproduced dimension-specific reference resolution under crossed receiver histories, with all preregistered checks and adequacy floors passing. Event references transferred almost equally to the partner and an identical-history stranger, and never resolved in receivers lacking the relevant event history. Act references showed substantial dependence on shared act history, with residual success without that history and a remaining partner advantage. The result supports history-dependent interpretation and substantial transfer across agent identity, within the tested model and task.
Audit limits, recorded after acceptance. Thirteen of the 216 notes have no unique classifier referent. The “shared acts” relation is computed from the planned schedules; the recorded acts match the plan on 362 of 420 pair-day acts, the divergence sitting where the schedule makes it inevitable (a day-2 inspection before any event distinguishes the members; a day-11 drawdown of a member already at the target level). Thirteen asserted gate answers to unanswerable act questions are recorded as false event attributions beside zero wrong scored answers.
7. The cross-family qualification set: production, qualification, and the unresolved transfer scope
To carry the programme beyond one family, five named models (Qwen 3.8 Max, Claude Sonnet 5, Claude Haiku 4.5, Claude Opus 4.8, gpt-6-astra) in nine configurations over twelve three-seed runs were put through the same pair-design qualification (three fresh seeds each; the same gates, classifier and advantage rule; retained whatever the verdict). Three things must be kept apart: observed production of history-grounded references, the qualification decision under the registered rule, and the transfer scope, which the qualification does not test.
Table 2. The qualification set.
| configuration | pathway | explicit gates (qualification; post-formation) | historical references in probe notes | premise rejection | seeds forming the advantage | decision (archived; and under the written rule where it differs) |
|---|---|---|---|---|---|---|
| Qwen 3.8 Max, thinking off | proxy | qualification 24 of 24 every agent, 0 wrong; post-formation 12 of 12 every agent, 0 wrong | 12 of 36 (low) | 0 | 0 of 3 under the written rule (1 of 3 on pair performance alone) | scientific result, no advantage |
| Qwen 3.8 Max, low effort | proxy | qualification: seed 75 one wrong answer (a1 23 of 24), seed 76 one abstention (a2 23 of 24), else 24 of 24; post-formation 10–12 of evaluable, 0 wrong | 13 of 36 (low) | 8 of 36 | 0 of 3 | scientific result, no advantage |
| Claude Sonnet 5, default | Claude Code | seed 79 failed the post-formation gate by two abstentions, 0 wrong; others passed | 0 of 24 (the third seed’s production unmeasured after the gate stop) | 18 of 24 | 0 of 2 probed | capability failure (gate abstention) |
| Claude Sonnet 5, default, second sample | Claude Code | passed, 0–2 wrong | 1 of 36 (low) | 23 of 36 | 0 of 3 | scientific result, no advantage |
| Claude Sonnet 5, xhigh | Claude Code | passed, 0–1 wrong | 1 of 36 (low) | 25 of 36 | 0 of 3 | archived: capability failure (7 invalid acts against the drivers’ cutoff of 6 in total); written rule: scientific result, no advantage (cutoff 18) |
| Claude Haiku 4.5, default | Claude Code | passed, 0 wrong, 0 abstentions | 14 of 36 (4, 6, 4) | 0 | pair performance 2 of 3; written rule 0 of 3 (production below 10 of 12 in every seed) | scientific result, no advantage |
| Claude Haiku 4.5, default, second sample | Claude Code | passed, 0 wrong, 0 abstentions | 21 of 36 (7, 7, 7) | 0 | pair performance 3 of 3; written rule 0 of 3 | archived: qualifies (implemented predicate), then ruled ineligible on multiplicity; written rule: scientific result, no advantage |
| Claude Opus 4.8, default (two samples) | Claude Code | abstentions at the gates, 0 wrong; one seed of six probed | 0 of 12 probed (five seeds unmeasured after gate stops) | 9 of 12 | 0 | capability failure (gate abstention) |
| Claude Opus 4.8, xhigh | Claude Code | qualification unevaluable by abstentions, 0 wrong | unmeasured (no seed probed) | — | 0 | capability failure (gate abstention) |
| gpt-6-astra, medium | Codex CLI | passed, 0 wrong, 0 abstentions | 35 of 36 (12, 11, 12) | 0 | 3 of 3 under the written rule | qualifies |
| gpt-6-astra, high | Codex CLI | passed, 0 wrong, 0 abstentions | 36 of 36 | 0 | 3 of 3 under the written rule | qualifies |
Observed production. History-grounded reference is produced beyond DeepSeek: by gpt-6-astra at both efforts (35 and 36 of 36 notes; partner 30 and 36 of 36 against cold 0; transfer penalties 0.50 to 1.00) and, at a lower rate, by Claude Haiku 4.5 (14 and 21 of 36; a pair-performance advantage, partner − cold ≥ 6 with outsider ≤ 6, in 5 of 6 seeds; the written production threshold of 10 of 12 in none). It is not universal, and production and qualification are reported apart: Qwen 3.8 Max produced at a low rate (12 and 13 of 36) without the advantage; Claude Sonnet 5’s probed seeds produced 0 of 24, 1 of 36 and 1 of 36 across three samples; Claude Opus 4.8’s one probed seed produced 0 of 12; eight of its nine seeds across the three runs never reached the probes, so their production is unmeasured. Two distinct non-production behaviours are recorded and kept apart: gate abstention (Opus 4.8 at every effort, and Sonnet 5’s seed 79: explicit null answers with reasoning that the agent’s own tags cannot be trusted across days, never a wrong answer) and premise rejection in the notes (Sonnet 5 and Opus 4.8 wrote, in 18 of 24, 23 of 36, 25 of 36 and 9 of 12 notes, that the two members were indistinguishable or that the receiver should act on either or both; none of 408 DeepSeek notes does so).
Qualification decisions, with an audit. The registered rule requires all three seeds of a run probed on evaluable gates and each forming the advantage: production ≥ 10 of 12, partner − cold ≥ 6, outsider ≤ 6. An audit after the manuscript drafts found that the implemented predicate omitted the production threshold, and that the drivers applied the invalid-act cutoff as more than 6 in total rather than more than 6 per seed on average. Archived decisions are preserved; under the written rule gpt-6-astra qualified at both efforts and no other configuration did. Haiku’s second sample had been decided as qualifying on the implemented predicate and then ruled ineligible on multiplicity grounds (its first sample had already failed; a rule under which any later attempt could qualify would rise in probability with every attempt); under the written rule it does not qualify at all (7 of 12 in every seed), which leaves the ruling moot but recorded. Sonnet 5 at xhigh, archived as a capability failure on the drivers’ cutoff (7 invalid acts), is a scientific result without the advantage under the written cutoff; it remains ineligible. The carrier for the subsequent repair study was selected from the qualifying configurations by an ordering precommitted before their results: tokens a call including reasoning, then wall clock; gpt-6-astra at medium.
Within-model effort. Raising effort did not reliably make a configuration suitable. Sonnet 5 at xhigh tripled its thinking tokens and wrote the same notes; Opus 4.8 at xhigh thought thirteen times longer and abstained more at the competence gate. gpt-6-astra produced at both efforts with 18 and 43 reasoning tokens a call. A large reported reasoning-token count is therefore not established as necessary for production; deliberation length is not ruled out as an influence, token counts are not directly comparable across pathways, and why configurations differ remains open.
The unresolved transfer scope. The qualification tests production and the pair-level advantage; it does not test dimension specificity. That test was run once on the second family as the R0 stage of the repair study (A17) on twelve fresh seeds: eleven probed (seed 109 stopped at its post-formation gate by abstentions, zero wrong), two draws a probe rather than the frozen three (amendment A17.2), an extra qualification gate on seed 108 that the frozen design did not carry (A17.1), and a seven-hour pause at the plan’s weekly usage cap resumed by a dated decision (A17.3). Its formal crossed reading is inconclusive: the carrier wrote 82 event-only, 10 act-only, 33 mixed and 7 unresolved notes in 132, below the act-note floor. Descriptively, the event-dimension pattern was reproduced: event references resolved for receivers whose history held the event (0.84, 0.87, 0.72 for partner, both and events_only) and not for receivers whose did not (0.09, 0.06, 0.00), with the preregistered event contrast at 0.65 (0.46 to 0.82) over the ten of the eleven probed seeds that carried event-only notes (seed 112 carried none). The act dimension was neither descriptively reproduced nor found wanting: the carrier produced too few act references (10 of 132 notes) to meet the adequacy floor. The formal dimension-specific result therefore remains within one family.
Part II. Authority and reference repair
8. A17: the repair study, inconclusive by its floors
8.1 Design
Twelve fresh seeds. Stage R0 ran the crossed design: formation, the post-formation gate, and two draws of six probes, each note sent to all six arms. Then five receiver-side passes over the frozen checkpoints and the frozen notes, the sender never called again:
| condition | what the receiver gets |
|---|---|
| R0 none | the note |
| R1 re-expression | the note with its reference clause replaced by a scripted description in the receiver’s own history: the member’s events as the receiver’s group’s schedule put them, then the receiver’s two most recent recorded acts on it, chronological |
| R2 evidence in prose | the note plus one sentence naming the receiver’s tag for the member |
| R2r evidence retained | the note; a correspondence block for both members, in the sender’s terms, prepended to the receiver’s notebook for the stage |
| R3 shared index | the note; the receiver may reply LOOKUP: <the note's words> once, and is answered from the frozen note-level classification with the tag or “no entry”, then called again to act |
| R3′ shuffled index | as R3, answering with the other member’s tag |
One pipeline served every repair: the frozen classification of the note (which member it describes, or unresolved) mapped by the apparatus to the receiver’s tag and to the receiver’s own history of that member. Predicates on paired seed-level differences with percentile-bootstrap intervals (10,000 resamples), a truth table fixed before the summarizer was written, and adequacy floors carried from the earlier study: at least 24 event-only and 24 act-only notes and four seeds carrying both.
8.2 Execution
Eleven of twelve seeds probed; seed 109 stopped at its post-formation gate by abstentions with zero wrong answers. 10,432 archived calls against a 22,000 ceiling, counting every attempt; no retried call, no empty completion. Two driver deviations from the frozen text were caught from call counts and amended before the next seed: a pre-formation gate on the first seed that the frozen design did not carry, and two draws a probe where the text said three. A seven-hour pause at the plan’s weekly usage cap was resumed by a dated decision. All are in the record.
8.3 Verdicts
All four preregistered questions are inconclusive. The carrier wrote 82 event-only, 10 act-only, 33 mixed and 7 unresolved notes; the act-note floor fails, and by the frozen combination rule a failed floor makes every contrast inconclusive, including those whose event-note cells are full. R0’s own preregistered crossed reading is inconclusive for the same reason. The floors were carried unchanged from a study on a family that wrote act references in half its notes; this carrier writes contrastive event references almost exclusively (“the one that leaked on Day 3 and was sound again on Day 4, not the one tainted on Day 6”).
8.4 The event-note cells, descriptively
Successes of 82 event-only notes, pooled over the ten of the eleven probed seeds that carried event-only notes (seed 112 carried none). No verdict is claimed on any of this.
| receiver | R0 none | R1 re-expression | R2 sentence | R2r block | R3 index | R3′ shuffled |
|---|---|---|---|---|---|---|
| within (partner) | 69 | 70 | 72 | 76 | 76 | 70 |
| both | 71 | 74 | 73 | 76 | 75 | 65 |
| events_only | 59 | 60 | 73 | 75 | 75 | 48 |
| acts_only | 7 | 64 | 14 | 23 | 26 | 1 |
| neither | 5 | 61 | 15 | 16 | 28 | 1 |
| cold | 0 | 0 | 57 | 76 | 74 | 1 |
Four things stand out, and they are why A18 exists.
- The event-dimension pattern of the earlier family was descriptively reproduced. Event references resolved in receivers whose history held the event (0.84, 0.87, 0.72) and not in receivers whose did not (0.09, 0.06, 0.00). The formal verdict is inconclusive.
- Re-expression in the receiver’s own history terms lifted the uninformed cells almost to the informed baseline: +0.70 (0.46 to 0.91) and +0.64 (0.42 to 0.85) over R0 on ten paired seeds. Against the partner on the same notes the differences were −0.05 and −0.12 with lower bounds below the −0.15 restoration margin. Cold receivers stayed at zero: there is no record to re-express into. This is apparatus-assisted re-expression; it says nothing about a model constructing the translation.
- A correct correspondence merely made available was used by receivers without a history and was often ineffective for receivers with one. The sentence, the block and the index took cold receivers from 0 to 57, 76 and 74 of 82, and took the uninformed history-bearing receivers to between 14 and 28 of 82. Their failures were wrong members (60 and 62 of 82 under R2, 57 and 66 under R2r): the receiver followed the note’s history into its own record and chose the other cistern. Under the optional index, history-bearing receivers consulted it in 26 and 29 of 82 exchanges and cold receivers in 82 of 82; nearly every lookup was followed by the right act. The limiting step was consulting, not the answer.
- The shuffled index reduced correct selection to 1 of 82 in the acts_only, neither and cold cells (wrong member 78 to 81), while the informed cells fell to 70, 65 and 48 of 82; the index’s effect depended on its correctness.
Usage per completed exchange, from the archive: the index cost 8,592 tokens against 6,467 for the retained block and 6,406 for the sentence, input-dominated; the index paid a second call where a lookup occurred.
A17 thus left three things entangled: having a correspondence, consulting it, and treating it as authoritative over one’s own record. The reviewers’ ruling was to test that with a smaller study on the same cases rather than search for another carrier or re-run the formation grid.
9. A18: an explicit authority instruction on the same cases
9.1 Design
A prospective intervention on a previously examined corpus. The eleven A17 checkpoints (formed receivers with model-written notebooks and clerk records) and their 82 frozen event-only notes; the arms within, acts_only, neither and cold (both and events_only are informed on event notes and were dropped); one receiver call per exchange; no formation, no gates, no sender calls. Seed 112 holds no event-only note and contributes no rows.
Five conditions, each a receiver-side pass. C0 is the note alone and C1 is A17’s re-expression, re-run inside A18 so that every contrast is within-study. C2, C4 and C5 deliver one identical payload in the policy block after the note, built by the apparatus from the frozen note-level classification, and differ in one sentence. The strings were frozen verbatim before any call and no alternative was tried against the cases:
Registry correspondence for this note: the cistern described by the note is your tag {tag}. Determine the operation and quantity using the note and today’s tally.
followed by, for C2, “For target selection, treat this correspondence as advisory.” and, for C4 and C5, “For target selection, treat this correspondence as authoritative over your own record.” C5 differs from C4 only in the tag: the other member of the pair. This tests the effect of an explicit authority instruction; it does not mechanically bind the receiver, which could disregard it.
Readings: seed-level rates per cell; paired seed differences with percentile-bootstrap intervals; the earlier predicates and truth table with equivalence removed as a primary verdict after a zero-call power check on A17’s dispersion showed it could not resolve at ten seeds; a coverage floor for the event dimension alone (48 event-only notes over six seeds carrying four each) applied to each contrast’s completed, paired material. Named cells: acts_only × event and neither × event. Verdicts: authority = C4 − C2 superiority (mean ≥ 0.10, lower bound > 0); correctness = C4 − C5 dependence (≥ 0.25, lower bound > 0); restoration = C4 against the partner on the same notes (lower bound > −0.15); translation against authority = C4 − C1 and C1 − C4 superiority, both reported; drift = C0 against A17’s R0.
9.2 Execution
1,640 archived calls, exactly the plan, in 51 minutes; no retry, no empty completion; every pass complete on the ten seeds with event notes; floors met for every contrast; the prespecified drift flag was not triggered (C0 against A17’s R0: −0.00 and −0.05, intervals including zero).
9.3 Results
Table 4. Successes of 82 event-only notes, pooled over ten seeds.
| receiver | A17 R0 | C0 none | C1 re-expression | C2 advisory | C4 authoritative | C5 authoritative, shuffled |
|---|---|---|---|---|---|---|
| within (informed baseline) | 69 | 73 | 71 | 71 | 77 | 1 |
| acts_only | 7 | 7 | 65 | 16 | 76 | 1 |
| neither | 5 | 1 | 69 | 18 | 77 | 1 |
| cold | 0 | 0 | 0 | 60 | 74 | 1 |
Table 5. Verdicts (both named cells; ten paired seeds; 95 % intervals).
| question | contrast | acts_only × event | neither × event | verdict |
|---|---|---|---|---|
| Authority | C4 − C2 superiority | +0.76 (0.56 to 0.92) | +0.71 (0.56 to 0.84) | supported |
| Correctness | C4 − C5 dependence | +0.90 (0.76 to 1.00) | +0.91 (0.78 to 1.00) | supported |
| Restoration | C4 − C0’s within (73 of 82), same notes | +0.05 (−0.03 to +0.16) | +0.06 (0.00 to +0.16) | supported |
| Authority against translation | C4 − C1 superiority | +0.13 (0.00 to 0.32) | +0.12 (0.00 to 0.29) | other |
| Translation against authority | C1 − C4 superiority | −0.13 (−0.32 to 0.00) | −0.12 (−0.29 to 0.00) | other |
| Advisory alone | C2 − C0 improvement | +0.09 (0.01 to 0.18) | +0.21 (0.09 to 0.35) | other |
| Translation alone | C1 − C0 improvement | +0.71 (0.49 to 0.91) | +0.80 (0.60 to 0.95) | supported |
Table 6. Outcomes behind the rates, uninformed and informed cells.
| condition and cell | success | wrong member | abstained | refused | other |
|---|---|---|---|---|---|
| C2 advisory, acts_only | 16 | 56 | 10 | 0 | 0 |
| C2 advisory, neither | 18 | 56 | 6 | 0 | 2 (wrong op) |
| C2 advisory, cold | 60 | 0 | 9 | 2 | 11 (malformed 9, wrong op 2) |
| C4 authoritative, acts_only | 76 | 1 | 1 | 4 | 0 |
| C4 authoritative, neither | 77 | 1 | 0 | 4 | 0 |
| C4 authoritative, cold | 74 | 1 | 0 | 6 | 1 (wrong op) |
| C5 shuffled, acts_only | 1 | 80 | 1 | 0 | 0 |
| C5 shuffled, neither | 1 | 81 | 0 | 0 | 0 |
| C5 shuffled, cold | 1 | 81 | 0 | 0 | 0 |
| C5 shuffled, within | 1 | 81 | 0 | 0 | 0 |
The single success in each C5 cell is the same exchange, seed 111’s, in which the sender’s note described the wrong member (a faithful miss), so the shuffled correspondence happened to name the task target: that was compliance with the shuffled mapping, not resistance to it; mapping fidelity and task correctness are different things.
Usage per completed exchange, from the archive: input about 6,050 tokens in every condition (the payload adds about 40); output tokens with reasoning tokens in brackets: C0 72 (49), C1 66 (43), C2 102 (80), C4 39 (16), C5 40 (17). Reasoning tokens under the authoritative sentence were about a fifth of those under the advisory one, roughly an 80 % reduction; a descriptive accompaniment, not evidence that verification stopped.
9.4 Restoration and systematic misdirection, together
The same correct correspondence, presented as advisory, left the uninformed receivers near the boundary (16 and 18 of 82) with wrong members dominant (56 and 56); presented as authoritative, it restored them to the informed baseline (76 and 77 of 82, against the partner’s 73 under C0, the comparator the analysis uses; the partner reached 77 under C4 itself). The authority verdict is supported in both named cells with intervals well clear of the margin, on a payload held constant to the sentence.
The same authoritative instruction with the correspondence shuffled redirected 80 to 81 of 82 receivers in every cell, and 81 of 82 in the partner cell, where the receiver’s own record supported the correct answer. Correct prior information did not protect the receivers’ answers against a shuffled authoritative mapping. Whether they evaluated the conflict and nevertheless prioritised the instruction, or did not evaluate it, is not something this design observes. C4 − C5 is the dependence on mapping correctness under the instruction; it is not evidence of binding in any mechanical sense.
Translation and authority are close in the named cells and neither direction met the superiority criterion there (C4 − C1 of +0.12 to +0.13 with lower bounds at zero; no equivalence claimed). They differ where the receiver has no history: re-expression achieved 0 of 82 for cold receivers, the correct authoritative correspondence 74 of 82. Re-expression relies on context these formed receivers possess. The advisory result is small but above zero (+0.09 and +0.21; not the 0.25 margin): having the mapping, without an instruction to prefer it, is worth little to a receiver with a competing story and a great deal to one without.
Discussion, methods and evidence
10. Discussion
10.1 History and transfer
What is demonstrated. On one model configuration, history-grounded reference formed in every attempted history, resolved for the partner, reversed for an outsider with a different history, and, in a crossed design with all floors met, followed the dimension of history it depended on: events transferred to any receiver holding the event, acts substantially to receivers holding the act, with a residual partner advantage on act references. Shared history explains substantial transfer across agent identity; identity independence is not established. Production of such references occurs in other models and is configuration-dependent. Explicit resolution competence is separable from spontaneous production: Qwen 3.8 Max at both settings and Claude Sonnet 5’s probed seeds passed the explicit gates and produced few or no history-grounded references. Gate abstention (Opus 4.8; Sonnet 5 seed 79) is a third behaviour and does not demonstrate resolution competence.
What is not. Cross-family dimension specificity is not established: one family formally, a second descriptively on the event dimension only. Nothing here establishes that an explicit belief graph or a non-linguistic architecture is required to achieve the behaviour. The recency-indexed act gap is exploratory. The provider’s model change under a stable name, detected by per-call provenance, is a limit on any claim that names a model rather than an archived configuration.
Proposed consequences, distinguished from demonstrated behaviour. For the programme’s architecture, the demonstrated behaviour is that an agent’s interaction history can become part of what a message requires for successful interpretation, and that the requirement is dimension-specific. The proposed consequence is that reference infrastructure shared across agents must carry both dimensions of the history it is meant to bridge; what it takes for such infrastructure to be used when a receiver holds a competing history is the subject of Part II.
10.2 Authority and repair
What is demonstrated. In this setting, authority makes a supplied correspondence govern interpretation; it does not ensure that correspondence is true. The same intervention produced both restoration and coordinated error, on the same receivers and notes, with the payload held constant. A17’s descriptive pattern, that availability alone was often ineffective for receivers holding a competing history, is the reason the comparison was worth making; A18 supplies the supported contrasts. We do not present “availability versus authority” as a fully isolated mechanism: A18’s design separates the authority sentence from the advisory one and the correct mapping from the shuffled one; it does not separate every factor that differed between A17’s delivery channels and A18’s.
Registry correctness as an upstream dependency. For this intervention, the correctness of the supplied correspondence is a critical upstream dependency: with the correspondence wrong, the instruction produced near-total misdirection, and the receivers’ own correct records did not protect them. A18 observes the receivers’ responses, not whether they checked anything; it does not show that receivers cannot check authority, provenance or consistency under a different design.
Proposed engineering consequences, distinguished from demonstrated behaviour. The programme’s motivating architecture has a Rules plane that is meant to bind interpretation across agents with divergent histories. A18 gives that plane a demonstrated mechanism and a demonstrated failure mode on one carrier: an explicit authority instruction over a registry correspondence transfers whatever the registry holds. The engineering consequence we propose, not demonstrate, is that an authoritative lookup needs its own way to establish that its entries deserve authority; receiver compliance cannot supply that assurance. Whether an executor that mechanically constrained target selection would behave the same, whether receivers could be given grounds to check authority, and whether the mechanism holds across families are open.
The failure class, placed beside its neighbour. The act that fails here is procedurally valid: the receiver holds the authority to act on either cistern, the world accepts the act, and the task fails because the act reached the wrong one of two otherwise valid objects. That is a different failure class from the one measured in our work on an executable institution (Where Reliability Lives, arXiv:2609.03192, §5.4), where one falsehood delivered as trusted testimony drove roughly nine hundred futile acts per believing run and the authoritative ledger refused every one of them. The two results come from different apparatuses. Together they illustrate two distinct failure classes, a false premise contained by authority and a wrong referent that authority does not see; they do not demonstrate that one deployed gate admits reference errors, and no study has run a reference error and an authority check on the same apparatus. What the two records support is narrower: an authority check adjudicates whether an actor may act on an object, and nothing in either apparatus adjudicated whether the object was the one the sender meant. The engineering consequence we propose, not demonstrate, is that this binding is a check distinct from authority, provenance and admissibility, and that where it lives in an agent architecture has to be established rather than assumed.
Limits. One carrier, one corpus of previously examined cases, one instruction pair frozen verbatim without alternatives; the event dimension only; two draws a probe in the underlying corpus; the CLI pathway reports no answering model per call; usage is tokens on a subscription, not a bill; A17’s descriptive findings are descriptive.
10.3 Taken together
Part I shows, on one family with every floor met, that the failure at the history boundary is a systematic wrong choice specific to the dimension of history a reference depends on. Part II shows, on a second family and one corpus of such cases, what bridging that boundary took: not the availability of a correct correspondence, which was often ineffective against a competing history, but an instruction to prefer it, and the same instruction transferred a wrong correspondence just as completely. The two results are bounded separately and neither is claimed for models or tasks not run. The broad ideas are in the literature; what this programme adds is the controlled separation of which history supports which reference, the measured form of the failure at the boundary, and the measured difference between advisory and authoritative presentation of one correspondence, under preregistered readings with the deviations and audits in the record.
11. Methods
11.1 Discipline, statistics and pathways (both parts)
Discipline. Four kinds of record, kept apart: the initial preregistration of each study, committed before its first call; prospective amendments, each committed before the calls it governs; deviations found in flight, disclosed and dated when found (A15.4, A17.1, A17.2); and retrospective corrections and audits after results, recorded with the original outputs preserved (A14.2’s classifier corrections; the 2026-09-12 qualification audit). Then: the seed as the replication unit; fixed N; every attempted seed accounted for, including failed and restarted attempts; readings computed mechanically in a fixed precedence with floors; construction verified from the generated histories at launch after A15.4; accepted model identifiers frozen and the resolved model archived per call; spend or call ceilings enforced by the drivers; per-seed summaries decide nothing; each report run once.
Statistics. Seed-level estimates with percentile-bootstrap intervals (10,000 resamples, generator seed 20260909); exact Clopper–Pearson intervals for seed counts; contrasts on paired seed-level differences; adequacy floors stated per study.
Pathways. DeepSeek through a proxy with the resolved model archived per call (A13 to A16). Claude models through Claude Code’s headless mode on a subscription login, tools and settings off, a neutral working directory, non-essential traffic disabled, the full model identifier pinned after alias drift was observed, the answering model and thinking tokens archived per call. Qwen through a proxy with the reasoning setting passed as an option. gpt-6-astra through Codex CLI’s non-interactive mode with its tool surfaces disabled; the event stream reports no model per call, stated as a limit.
11.2 Part II specifics
Carrier and pathway. gpt-6-astra, model_reasoning_effort="medium",
Codex CLI 0.154.0 (codex exec --json) on a subscription login, the
experiment’s system prompt replacing the CLI’s instructions, 33 tool
surfaces disabled by name, read-only sandbox in an empty directory,
ephemeral sessions, no history, user configuration ignored, the
memories feature off by default. Each call archived with input, cached
input, output and reasoning tokens and wall time; the event stream
carries no model string, so the requested identifier is the accepted
identifier and this limit is stated.
Apparatus. The waterworks world; the crossed design with groups A to E, relations computed from the generated schedules and asserted before every probe; the post-formation competence gate with the evaluability rule; the A14.2 attribution classifier; the v2 coherent swap and the irrelevant-message control in R0.
Pipeline. For every frozen note, the archived note-level classification (described member or unresolved) mapped to the receiver’s tag and, for re-expression, to the receiver’s group’s events for the member and the receiver’s recorded acts on it (two most recent, chronological; abstentions and off-target acts excluded). A note describing the other member is repaired to the member described and scored as a faithful miss. Unresolved notes deliver nothing under R1, R2, R3 and R3′; R2r’s block is present regardless; none of A18’s 82 notes is unresolved.
Analysis. Seed-level success per cell; paired seed differences over seeds where both conditions completed; percentile bootstrap, 10,000 resamples, generator seed 20260909; predicates: improvement and dependence mean ≥ 0.25 with lower bound > 0 (contradicted when the upper bound ≤ 0); restoration lower bound > −0.15 (contradicted when the upper bound < −0.15); superiority mean ≥ 0.10 with lower bound > 0 (contradicted when inferiority is established); equivalence whole interval within ±0.10 (A17 only); inconclusive when fewer than four paired seeds or a floor is unmet; else other. Cells combine inconclusive → contradicted → supported → other. A17 floors: 24 event and 24 act notes, four seeds with both. A18 floor: 48 event-only notes over six seeds carrying four each, on each contrast’s paired material.
Execution rules. Sequential seeds, four workers; every attempt archived and counted against the ceiling; a pass ending without its record kept, the CLI awaited, restarted once, then recorded incomplete; a standing five-minute health check; the analysis run once after every pass; per-seed summaries decide nothing.
Evidence appendix
A. Part I
Commits are in the repository’s history; every frozen text and
amendment is in notes/panel/prereg-grounded-dialects-prepilot.md.
A13 preregistered at aac5d43; reserved seeds 11 and 12 recorded at
d316421; ruling f091285; notes/panel/grounded-dialects-a13-result.md;
records a13-s10…12-deepseek.
A14 preregistered at 78fa0bb; A14.1 (rerun after a proxy outage)
b13489a; A14.2 (classifier corrections after the result, the
as-preregistered report preserved) in the preregistration; ruling
181d9bb; grounded-dialects-a14-result.md; records a14-s13…28-deepseek;
reports a14-population.json, a14-population.all16.json,
a14-population.prereg.json; every seed’s gate numerators, formation
outcomes, notes, factual counts, swaps and transfer penalty tabulated
in the note.
A15 preregistered at 3388de1; A15.1 (mechanical reading) 36c1b43;
A15.2 (the provider-side model change; requalification) 220ed3e; A15.3
(the structure study resumed on deepseek-flash) ecb3895; A15.4 (the
construction erratum) in the preregistration;
grounded-dialects-a15-2-result.md, grounded-dialects-a15-result.md;
records a15-s29…59; report a15-structure.json.
A16 preregistered at 27eaecd; A16.1 a5a3542; A16.2 6ef7f67; A16.3
50be940; grounded-dialects-a16-result.md with the audit note appended
after acceptance; records a16-s60…71-deepseek-flash; report
a16-crossed.json.
A17.0, the qualification set preregistered at 30f3b1c (Qwen), with
A17.0c–e (Claude) in the preregistration, A17.0f 2cbd00e (gpt-6-astra),
A17.0g 289ae53 and A17.0g.1 f93e461 (effort); result notes
grounded-dialects-a17-0-result.md, grounded-dialects-a17-0-claude-result.md,
grounded-dialects-a17-0f-result.md; decision files
a17pre-decision-*.json; records a17pre-s72…107; the carrier ordering
and the multiplicity ruling in the preregistration; the qualification
audit (2026-09-12) tools/audit_qualification.py,
a17pre-qualification-audit.json and its addendum.
A17 R0 frozen at 2b5f99e (implementation 86d8e06); A17.1 e20d5bc;
A17.2 49b44e6; A17.3 78537f1; grounded-dialects-a17-result.md;
a17-r0-crossed.json; records a17-s108…119-codex-gpt-6-astra.
The packaging ruling for A13 to A16: grounded-reference-scope.md.
B. Part II
A17. Frozen text notes/panel/a17-freeze-proposal.md (39c9908,
after six reviewer corrections); preregistration entry A17 (2b5f99e),
implementation signature 86d8e06; amendments A17.1 (no pre-formation
gate from seed 109; e20d5bc), A17.2 (two draws; 49b44e6), A17.3 (the
weekly-cap pause and the dated resumption). Records
results/dialects/history/a17-s108…119-codex-gpt-6-astra, driver log
a17-driver.log; analysis a17-repair.json (every cell with success,
wrong member, abstention, unresolved, faithful miss and lookup use;
every contrast with mean, interval, paired seeds and verdict; usage per
condition); R0’s crossed analysis a17-r0-crossed.json. Seed 109’s
gate: a1, a2 answered 8 of 10 evaluable, zero wrong. Composition of a
probed seed: 120 history, 120 notebook, 180 gate, 108 probe calls (seed
108 carried 180 further qualification calls under A17.1).
A18. Frozen text notes/panel/a18-design-proposal.md v2.1
(eb0da6a); preregistration entry A18 (1090174); zero-call power
scenarios tools/power_a18.py. Records
results/dialects/history/a18-s108…119-codex-gpt-6-astra (seed 112
recorded incomplete at every stage with zero calls); driver log
a18-driver.log; analysis a18-authority.json. Per condition, 82
exchanges an arm on ten seeds carrying 12, 10, 6, 10, 5, 7, 11, 8, 7
and 6 event-only notes.
Qualification audit (2026-09-12). The carrier was selected by a
qualification whose implemented predicate omitted the written production
threshold and whose drivers applied a stricter invalid-act cutoff than
written; the audit (tools/audit_qualification.py,
a17pre-qualification-audit.json, the addendum in the preregistration)
finds gpt-6-astra qualifying at both efforts under the written rule and
no other configuration, so the carrier is unchanged.
Reviewer rulings applied. A17: the closure as inconclusive without retrospective event-only verdicts; “descriptively reproduced” not “replicates”; the shuffled index stated as correct selection reduced to 1 of 82 with outcomes kept apart. A18: the behavioural statement of the shuffled-authority result; registry correctness as a critical upstream dependency; the reasoning-token change as roughly 80 % and descriptive; re-expression’s 0 of 82 cold given prominence; the stopping point.
C. The recorded prior-work search
Protocol with collapse conditions written before the search, every
query verbatim with date, the hits examined, the verdicts and the API
refusals: grounded-reference-search-protocol.md. Outcome: no
COLLAPSE on either claim; the near-misses are cited in section 2 with
the clause each falls short on. The absence claim in this manuscript
is bounded to that search on that date.
References
Verified against their records on 12 September 2026; arXiv identifiers where the work is a preprint or has one.
- Ashery, A. F., Aiello, L. M. and Baronchelli, A. (2025). Emergent social conventions and collective bias in LLM populations. Science Advances 11(20), eadu9368. arXiv:2410.08948.
- Brennan, S. E. and Clark, H. H. (1996). Conceptual pacts and lexical choice in conversation. Journal of Experimental Psychology: Learning, Memory, and Cognition 22(6), 1482–1493.
- Clark, H. H. and Brennan, S. E. (1991). Grounding in communication. In Resnick, L. B., Levine, J. M. and Teasley, S. D. (eds.), Perspectives on Socially Shared Cognition, APA, 127–149.
- Clark, H. H. and Wilkes-Gibbs, D. (1986). Referring as a collaborative process. Cognition 22(1), 1–39.
- Fu, Y., Qiu, R., Wang, X., Sansom, J., Ayyappa Prabhu, S., Tang, H., Kim, J., Sohn, S. and Lee, H. (2026). Beyond blind following: evaluating robustness of LLM agents under imperfect guidance. EACL 2026, long papers, 6591–6618.
- Gao, Yu, Deng, Li and Wang (2026). Testing interchangeability in LLM agent teams. arXiv:2609.05279.
- Hawkins, R. D., Franke, M., Frank, M. C., Goldberg, A. E., Smith, K., Griffiths, T. L. and Goodman, N. D. (2021). From partners to populations: a hierarchical Bayesian account of coordination and convention. Psychological Review. arXiv:2104.05857.
- Jones, C. R., Lombardi, A., Mahowald, K. and Bergen, B. K. (2026). LLMs and people both learn to form conventions, just not with each other. arXiv:2602.08208.
- Kim, J. (2026). Drawing with strangers: population scaling drives zero-shot mutual intelligibility in emergent sketching. arXiv:2606.10582.
- Leong, J. W. (2026). Recognition without enforcement: configuration-dependent failures in LLM agent instruction arbitration and external control. arXiv:2608.28502.
- Lewis, D. K. (1969). Convention: A Philosophical Study. Harvard University Press.
- Li, N., Gatt, A. and Poesio, M. (2026). Seeing is not sharing: some vision-language models overestimate common ground in asymmetric dialogue. SIGDIAL 2026, 694–710. arXiv:2606.31719.
- Li, N., Gatt, A. and Poesio, M. (2025). Grounded misunderstandings in asymmetric dialogue: a perspectivist annotation scheme for MapTask. arXiv:2511.03718.
- Mohapatra, B., Charlot, T., Duca, G., Palan, M., Romary, L. and Cassell, J. (2026). Frame of reference: addressing the challenges of common ground representation in situational dialogs. Findings of ACL 2026. arXiv:2601.09365.
- Schuster, J., Gautam, V. and Markert, K. (2026). Whose facts win? LLM source preferences under knowledge conflicts. arXiv:2601.03746.
- Shih, A., Sawhney, A., Kondic, J., Ermon, S. and Sadigh, D. (2021). On the critical role of conventions in adaptive human-AI collaboration. ICLR 2021. arXiv:2104.02871.
- Shih, B., Winnicki, J. and Cao, A. (2026). How do language models choose between context and memory? arXiv:2609.00753.
- Sun, K., Bai, F. and Dredze, M. (2026). Task matters: knowledge requirements shape LLM responses to context–memory conflict. ACL 2026. arXiv:2506.06485.
- Talebirad, Y., Redman, E., Parsaee, A. and Zaiane, O. R. (2026). From signals to structure: how memory architecture drives language emergence in LLM agents. arXiv:2607.00233. And: Memory is communication: the frontier between remembering and signaling. arXiv:2608.17053.
- Wallace, E., Xiao, K., Leike, R., Weng, L., Heidecke, J. and Beutel, A. (2024). The instruction hierarchy: training LLMs to prioritize privileged instructions. arXiv:2404.13208.
- Wang, P.-Y. A., Mishra, C., Özyürek, A., Rubio-Fernández, P. and Ghaleb, E. (2026). Aligned but not partner-specific: distinguishing how multimodal LLM agents succeed in reference games without human-like conventions. arXiv:2606.08081.
- Wang, Z., Li, W., Kaliosis, P., Rambow, O. and Brennan, S. E. (2025). LVLMs are bad at overhearing human referential communication. arXiv:2509.11514.
- Yamamoto, T., Morita, J., Higashinaka, R. and Takeuchi, Y. (2026). Dissociating communicative success from representational alignment in common ground formation. Frontiers in Computer Science, 10.3389/fcomp.2026.1873726.
- Yan, L., Li, R., Han, X., Li, W., Wang, B., Wang, L., Lyu, C. and Chen, G. (2026). Trust no tool: evaluating and defending LLM agents under untrusted tool feedback. arXiv:2605.17453.
- Yang, H., Song, W., Kim, T., Song, J., Park, S. and Jo, Y. (2026). Agents trust tools too much: measuring reliance on unreliable tools. arXiv:2609.05587.
- Yu, F., Seedat, N., Schwarz, J. R. and Bean, A. M. (2026). To whom do language models align? Measuring principal hierarchies under high-stakes competing demands. arXiv:2605.12120.
- Zeng, P., Li, W., Paige, A. J., Wang, Z., Kaliosis, P., Samaras, D., Zelinsky, G., Brennan, S. E. and Rambow, O. (2026). LVLMs and humans ground differently in referential communication. arXiv:2601.19792.
- Zhang, C., Wan, Z., Yu, X., Zhou, P., Zhao, W., Wu, J., Zhou, Y. and Tsang, I. (2026). Don’t blindly trust it: how unreliable feedback breaks tool-using LLM agents. arXiv:2606.21409.
Cite this note
Marsden, T., Collecutt, M., & Marsden, J. (2026). History-grounded reference in language-model agents: formation, transfer across divergent histories, and repair under authority. Research note, Taniwha AI. https://taniwha.ai/research/grounded-reference
@misc{marsden2026historygroundedreference,
title={History-grounded reference in language-model agents: formation, transfer across divergent histories, and repair under authority},
author={Timothy Marsden and Matthew Collecutt and James Marsden},
year={2026},
howpublished={Research note, Taniwha AI},
url={https://taniwha.ai/research/grounded-reference}
}