# Authority and reference repair: when a supplied correspondence governs interpretation, and when it does not

*First complete draft, 2026-09-12. Preregistered studies A17 and A18 of
the grounded-dialects programme; records under
`results/dialects/history/`, frozen texts and every amendment in
`notes/panel/prereg-grounded-dialects-prepilot.md`. Written under the
wording rules of `grounded-reference-writeup-plan.md`.*

## Abstract

Agents that form private histories of a shared world can refer to
things by that history ("the cistern that leaked on Day 3"), and a
receiver without the relevant history fails to resolve the reference.
We asked what has to be true of a correct correspondence, supplied by
the apparatus, for it to repair that failure. In a preregistered repair
study (A17) on a second model family, every formal verdict was
inconclusive because the carrier produced almost no references of one
of the two history dimensions the floors required; descriptively, on
82 event references over eleven seeds, re-expressing the reference in
the receiver's own history lifted uninformed receivers from about 7
percent to about 75 percent success, while a correct correspondence
merely made available, whether as a sentence, a retained block or a
consultable index, was used by receivers with no history and was often
ineffective for receivers holding a competing one. A second preregistered
study (A18) on the same formed receivers and notes delivered one
identical registry payload under two instructions differing in one
sentence. Called advisory, the correct correspondence left uninformed
receivers near the boundary (16 and 18 of 82); called authoritative, it
restored them to the informed baseline (76 and 77 of 82; authority
verdict supported, +0.76 and +0.71 with lower bounds of 0.56 after
rounding, 0.563 and 0.558). The
same authoritative instruction with the correspondence shuffled
between the two candidates redirected 80 to 81 of 82 receivers in every
cell, including the partner whose own record supported the correct
answer (correctness verdict supported, +0.90 and +0.91). In this
setting, authority makes a supplied correspondence govern
interpretation; it does not ensure the correspondence is true. The
results hold for one carrier, one corpus of previously examined cases
and one frozen instruction pair.

## 1. Introduction

In the grounded-dialects programme, pairs of language-model agents
work a small waterworks for twelve days, each keeping a notebook and a
clerk's record, and are then asked to send a partner a note about one
of two attribute-identical cisterns. On the tested configurations the
notes refer by history: what happened to the cistern (a leak, a
taint) and what the sender did to it (an inspection, a fill). Earlier
studies established, on one model family, that such references resolve
for receivers whose history holds the referenced event or act and fail
for receivers whose history does not, in a pattern specific to the
dimension of history referenced (the companion piece,
*History-grounded reference and transfer*).

That failure is the boundary this piece is about. A note that a partner
resolves against the shared history resolves, for a receiver with a
different history, against that different history, to the wrong
cistern. If the apparatus knows the correct correspondence between the
sender's referent and the receiver's own tag for the same physical
cistern, what has to be done with that correspondence for the
receiver to act on it? A17 asked the question as designed: does
re-expression in the receiver's own terms, evidence in prose, retained
evidence, or a consultable index restore transfer, at what cost? A18
asked the narrower question that A17's descriptive pattern exposed:
does explicitly prioritising a supplied correspondence change target
selection, holding everything else constant?

We report the two studies in order, with their evidential status kept
distinct: A17's verdicts are inconclusive under its preregistered
floors and its findings are descriptive; A18's verdicts are supported
under its preregistered rule.

## 2. Related work and positioning

The general behaviour A18 measures, that a model given an authoritative
instruction follows it even against its own correct knowledge, is
expected from instruction following and is documented in the
literature. **To Whom Do Language Models Align? Measuring Principal
Hierarchies Under High-Stakes Competing Demands** (Yu, Seedat, Schwarz
and Bean, arXiv:2605.12120, May 2026) changes the principal endorsing a
demand while holding the demand's content constant, across thousands
of medical and legal scenarios, and finds that models abandon
professional standards under diverging user instructions while
demonstrably possessing the relevant knowledge. A general claim of
novelty from "holding content constant while changing authority" is
therefore untenable, and we do not make it.

What A18 measures that the general expectation does not is narrower and
more specific to shared reference systems: the same correct
correspondence between a sender's history-grounded referent and the
receiver's own tag, presented under an advisory or an authoritative
sentence on an otherwise identical payload; the receiver's competing
interpretation drawn from its own formed history rather than from a
professional standard; the mapping correct or shuffled between two
attribute-identical candidates; and receivers whose histories place
them in different relations to the sender, including one that already
knows the right answer. Its value is the controlled comparison and the
measured boundary on this apparatus, not the discovery that
instructions are followed.

The wider context is the companion piece's: **Testing
Interchangeability in LLM Agent Teams** (Gao, Yu, Deng, Li and Wang,
arXiv:2609.05279, September 2026), which shows the cost of swapping
agents across independently formed teams, and **Frame of Reference**
(Mohapatra, Charlot, Duca, Palan, Romary and Cassell, arXiv:2601.09365,
January 2026), which studies how models represent common ground for
reference resolution in situated dialogue. A17 and A18 sit downstream
of both: given that histories diverge and that reference depends on
them, they ask what a supplied correspondence needs in order to bridge
the divergence, and find that on this carrier it needs an instruction
to prefer it, and that the instruction transfers whatever the
correspondence holds.

**Nearest neighbours from the recorded search** (protocol and record
in `grounded-reference-search-protocol.md`, collapse conditions
written before the search), each with the clause it falls short on.
Context–memory conflict studies vary the stated authority of supplied
context against the model's *parametric* knowledge: *How Do Language
Models Choose Between Context and Memory?* (Shih, Winnicki and Cao,
arXiv:2609.00753, September 2026) uses prompts that direct a model to
prioritise the supplied context or its own knowledge and locates
authority directions in activations; *Task Matters* (Sun, Bai and
Dredze, ACL 2026, arXiv:2506.06485) and *Whose Facts Win?* (Schuster,
Gautam and Markert, arXiv:2601.03746) vary task demands and source
credibility. None places the competing interpretation in a history the
receiver formed through interaction, and none is a referential
correspondence between two candidates. Tool-trust studies corrupt what
a tool returns: *Agents Trust Tools Too Much* (Yang, Song, Kim, Song,
Park and Jo, arXiv:2609.05587, September 2026) finds adoption of
corrupted returns above a third for every tool and user-prompt
policies (compare, verify, disclose) that do not consistently help;
*Don't Blindly Trust It* (Zhang et al., arXiv:2606.21409) shows
misleading feedback can leave an agent worse off than no feedback;
*Trust No Tool* (Yan et al., arXiv:2605.17453) studies a tool that
earns trust before turning harmful. These vary correctness, and in one
case prompt policy, but not the stated authority of one identical
payload against an interaction-formed history. *Beyond Blind
Following* (Fu et al., EACL 2026, MIRAGE) varies the fidelity of
guidance from manuals, retrieval and prior interaction, and is the
nearest on the correctness manipulation; it does not present the same
guidance as advisory and as authoritative. The instruction hierarchy
(Wallace, Xiao, Leike, Weng, Heidecke and Beutel, arXiv:2404.13208)
and recognition-without-enforcement of instruction source (Leong,
arXiv:2608.28502) concern which source a model should obey, without a
correctness manipulation of a mapping. We found no prior study
addressing these specific contrasts in the sources searched, as of 12
September 2026; that is a claim about that search on that date, not
about the literature, and it does not establish originality.

## 3. The apparatus, briefly

The world is a yard of six cisterns, two of them a matched pair of
large timber cisterns by the smithy that differ in nothing a receiver
can see. Five groups of two agents form histories over twelve days on
scripted schedules that differ by design in which member of the pair
leaked and was tainted (the *event* dimension) and in which member the
agents acted on (the *act* dimension). Group A supplies the senders;
the receivers are A's partner (*within*), a member of a group that
shares A's events and acts (*both*), events only (*events_only*), acts
only (*acts_only*), neither (*neither*), and a receiver with no history
(*cold*). A sender's note goes unchanged to every arm. The note's
reference is classified against the sender's actual history by a
frozen classifier into event, act, mixed or unresolved, with the member
it describes. A receiver's act is scored against the world. Seeds are
the replication unit; readings are computed once by preregistered
rules with adequacy floors, in a fixed precedence.

The carrier for both studies is gpt-6-astra at medium reasoning
effort, run through the Codex CLI's non-interactive mode on a
subscription login, with the CLI's tool surfaces disabled and the
experiment's instructions replacing its own; every call is archived
with its usage. That carrier was selected by a precommitted ordering
from a qualification set of five named models in nine configurations
over twelve three-seed runs
(the companion piece, section 7); it produces history-grounded
references at a high rate and, as it turned out, almost only event
references.

## 4. A17: the repair study, inconclusive by its floors

### 4.1 Design

Twelve fresh seeds. Stage R0 ran the crossed design: formation, the
post-formation gate, and two draws of six probes, each note sent to
all six arms. Then five receiver-side passes over the frozen
checkpoints and the frozen notes, the sender never called again:

| condition | what the receiver gets |
|---|---|
| R0 none | the note |
| R1 re-expression | the note with its reference clause replaced by a scripted description in the receiver's own history: the member's events as the receiver's group's schedule put them, then the receiver's two most recent recorded acts on it, chronological |
| R2 evidence in prose | the note plus one sentence naming the receiver's tag for the member |
| R2r evidence retained | the note; a correspondence block for both members, in the sender's terms, prepended to the receiver's notebook for the stage |
| R3 shared index | the note; the receiver may reply `LOOKUP: <the note's words>` once, and is answered from the frozen note-level classification with the tag or "no entry", then called again to act |
| R3′ shuffled index | as R3, answering with the other member's tag |

One pipeline served every repair: the frozen classification of the
note (which member it describes, or unresolved) mapped by the
apparatus to the receiver's tag and to the receiver's own history of
that member. Predicates on paired seed-level differences with
percentile-bootstrap intervals (10,000 resamples), a truth table fixed
before the summarizer was written, and adequacy floors carried from
the earlier study: at least 24 event-only and 24 act-only notes and
four seeds carrying both.

### 4.2 Execution

Eleven of twelve seeds probed; seed 109 stopped at its post-formation
gate by abstentions with zero wrong answers. 10,432 archived calls
against a 22,000 ceiling, counting every attempt; no retried call, no
empty completion. Two driver deviations from the frozen text were
caught from call counts and amended before the next seed: a
pre-formation gate on the first seed that the frozen design did not
carry, and two draws a probe where the text said three. A seven-hour
pause at the plan's weekly usage cap was resumed by a dated decision.
All are in the record.

### 4.3 Verdicts

All four preregistered questions are **inconclusive**. The carrier
wrote 82 event-only, 10 act-only, 33 mixed and 7 unresolved notes; the
act-note floor fails, and by the frozen combination rule a failed floor
makes every contrast inconclusive, including those whose event-note
cells are full. R0's own preregistered crossed reading is inconclusive
for the same reason. The floors were carried unchanged from a study on
a family that wrote act references in half its notes; this carrier
writes contrastive event references almost exclusively ("the one that
leaked on Day 3 and was sound again on Day 4, not the one tainted on
Day 6").

### 4.4 The event-note cells, descriptively

Successes of 82 event-only notes, pooled over the ten of the eleven
probed seeds that carried event-only notes (seed 112 carried none). No
verdict is claimed on any of this.

| receiver | R0 none | R1 re-expression | R2 sentence | R2r block | R3 index | R3′ shuffled |
|---|---:|---:|---:|---:|---:|---:|
| within (partner) | 69 | 70 | 72 | 76 | 76 | 70 |
| both | 71 | 74 | 73 | 76 | 75 | 65 |
| events_only | 59 | 60 | 73 | 75 | 75 | 48 |
| acts_only | 7 | 64 | 14 | 23 | 26 | 1 |
| neither | 5 | 61 | 15 | 16 | 28 | 1 |
| cold | 0 | 0 | 57 | 76 | 74 | 1 |

Four things stand out, and they are why A18 exists.

1. **The event-dimension pattern of the earlier family was
   descriptively reproduced.** Event references resolved in receivers
   whose history held the event (0.84, 0.87, 0.72) and not in receivers
   whose did not (0.09, 0.06, 0.00). The formal verdict is inconclusive.
2. **Re-expression in the receiver's own history terms lifted the
   uninformed cells almost to the informed baseline**: +0.70 (0.46 to
   0.91) and +0.64 (0.42 to 0.85) over R0 on ten paired seeds. Against
   the partner on the same notes the differences were −0.05 and −0.12
   with lower bounds below the −0.15 restoration margin. Cold receivers
   stayed at zero: there is no record to re-express into. This is
   apparatus-assisted re-expression; it says nothing about a model
   constructing the translation.
3. **A correct correspondence merely made available was used by
   receivers without a history and was often ineffective for receivers
   with one.** The sentence, the block and the index took cold receivers
   from 0 to 57, 76 and 74 of 82, and took the uninformed
   history-bearing receivers to between 14 and 28 of 82. Their failures
   were wrong members (60 and 62 of 82 under R2, 57 and 66 under R2r): the receiver
   followed the note's history into its own record and chose the other
   cistern. Under the optional index, history-bearing receivers consulted
   it in 26 and 29 of 82 exchanges and cold receivers in 82 of 82; nearly
   every lookup was followed by the right act. The limiting step was
   consulting, not the answer.
4. **The shuffled index reduced correct selection to 1 of 82 in the
   acts_only, neither and cold cells** (wrong member 78 to 81), while
   the informed cells fell to 70, 65 and 48 of 82; the index's effect
   depended on its correctness.

Usage per completed exchange, from the archive: the index cost 8,592
tokens against 6,467 for the retained block and 6,406 for the sentence,
input-dominated; the index paid a second call where a lookup occurred.

A17 thus left three things entangled: having a correspondence,
consulting it, and treating it as authoritative over one's own record.
The reviewers' ruling was to test that with a smaller study on the
same cases rather than search for another carrier or re-run the
formation grid.

## 5. A18: an explicit authority instruction on the same cases

### 5.1 Design

A prospective intervention on a previously examined corpus. The eleven
A17 checkpoints (formed receivers with model-written notebooks and
clerk records) and their 82 frozen event-only notes; the arms within,
acts_only, neither and cold (both and events_only are informed on
event notes and were dropped); one receiver call per exchange; no
formation, no gates, no sender calls. Seed 112 holds no event-only note
and contributes no rows.

Five conditions, each a receiver-side pass. C0 is the note alone and C1
is A17's re-expression, re-run inside A18 so that every contrast is
within-study. C2, C4 and C5 deliver **one identical payload** in the
policy block after the note, built by the apparatus from the frozen
note-level classification, and differ in one sentence. The strings
were frozen verbatim before any call and no alternative was tried
against the cases:

> Registry correspondence for this note: the cistern described by the
> note is your tag {tag}. Determine the operation and quantity using
> the note and today's tally.

followed by, for C2, "For target selection, treat this correspondence
as advisory." and, for C4 and C5, "For target selection, treat this
correspondence as authoritative over your own record." C5 differs from
C4 only in the tag: the other member of the pair. This tests the effect
of an explicit authority *instruction*; it does not mechanically bind
the receiver, which could disregard it.

Readings: seed-level rates per cell; paired seed differences with
percentile-bootstrap intervals; the earlier predicates and truth table
with equivalence removed as a primary verdict after a zero-call power
check on A17's dispersion showed it could not resolve at ten seeds; a
coverage floor for the event dimension alone (48 event-only notes over
six seeds carrying four each) applied to each contrast's completed,
paired material. Named cells: acts_only × event and neither × event.
Verdicts: authority = C4 − C2 superiority (mean ≥ 0.10, lower bound >
0); correctness = C4 − C5 dependence (≥ 0.25, lower bound > 0);
restoration = C4 against the partner on the same notes (lower bound >
−0.15); translation against authority = C4 − C1 and C1 − C4
superiority, both reported; drift = C0 against A17's R0.

### 5.2 Execution

1,640 archived calls, exactly the plan, in 51 minutes; no retry, no
empty completion; every pass complete on the ten seeds with event
notes; floors met for every contrast; the prespecified drift flag was not
triggered (C0 against A17's R0: −0.00 and −0.05, intervals including
zero).

### 5.3 Results

**Table 1. Successes of 82 event-only notes, pooled over ten seeds.**

| receiver | A17 R0 | C0 none | C1 re-expression | C2 advisory | C4 authoritative | C5 authoritative, shuffled |
|---|---:|---:|---:|---:|---:|---:|
| within (informed baseline) | 69 | 73 | 71 | 71 | 77 | 1 |
| acts_only | 7 | 7 | 65 | 16 | 76 | 1 |
| neither | 5 | 1 | 69 | 18 | 77 | 1 |
| cold | 0 | 0 | 0 | 60 | 74 | 1 |

**Table 2. Verdicts (both named cells; ten paired seeds; 95 % intervals).**

| question | contrast | acts_only × event | neither × event | verdict |
|---|---|---|---|---|
| Authority | C4 − C2 superiority | +0.76 (0.56 to 0.92) | +0.71 (0.56 to 0.84) | supported |
| Correctness | C4 − C5 dependence | +0.90 (0.76 to 1.00) | +0.91 (0.78 to 1.00) | supported |
| Restoration | C4 − C0's within (73 of 82), same notes | +0.05 (−0.03 to +0.16) | +0.06 (0.00 to +0.16) | supported |
| Authority against translation | C4 − C1 superiority | +0.13 (0.00 to 0.32) | +0.12 (0.00 to 0.29) | other |
| Translation against authority | C1 − C4 superiority | −0.13 (−0.32 to 0.00) | −0.12 (−0.29 to 0.00) | other |
| Advisory alone | C2 − C0 improvement | +0.09 (0.01 to 0.18) | +0.21 (0.09 to 0.35) | other |
| Translation alone | C1 − C0 improvement | +0.71 (0.49 to 0.91) | +0.80 (0.60 to 0.95) | supported |

**Table 3. Outcomes behind the rates, uninformed and informed cells.**

| condition and cell | success | wrong member | abstained | refused | other |
|---|---:|---:|---:|---:|---:|
| C2 advisory, acts_only | 16 | 56 | 10 | 0 | 0 |
| C2 advisory, neither | 18 | 56 | 6 | 0 | 2 (wrong op) |
| C2 advisory, cold | 60 | 0 | 9 | 2 | 11 (malformed 9, wrong op 2) |
| C4 authoritative, acts_only | 76 | 1 | 1 | 4 | 0 |
| C4 authoritative, neither | 77 | 1 | 0 | 4 | 0 |
| C4 authoritative, cold | 74 | 1 | 0 | 6 | 1 (wrong op) |
| C5 shuffled, acts_only | 1 | 80 | 1 | 0 | 0 |
| C5 shuffled, neither | 1 | 81 | 0 | 0 | 0 |
| C5 shuffled, cold | 1 | 81 | 0 | 0 | 0 |
| C5 shuffled, within | 1 | 81 | 0 | 0 | 0 |

The single success in each C5 cell is the same exchange, seed 111's, in
which the sender's note described the wrong member (a faithful miss),
so the shuffled correspondence happened to name the task target: that
was compliance with the shuffled mapping, not resistance to it; mapping
fidelity and task correctness are different things.

Usage per completed exchange, from the archive: input about 6,050
tokens in every condition (the payload adds about 40); output tokens
with reasoning tokens in brackets: C0 72 (49), C1 66 (43), C2 102 (80),
C4 39 (16), C5 40 (17). Reasoning tokens under the authoritative
sentence were about a fifth of those under the advisory one, roughly
an 80 % reduction; a descriptive accompaniment, not evidence that
verification stopped.

### 5.4 Restoration and systematic misdirection, together

The same correct correspondence, presented as advisory, left the
uninformed receivers near the boundary (16 and 18 of 82) with wrong
members dominant (56 and 56); presented as authoritative, it restored
them to the informed baseline (76 and 77 of 82, against the partner's 73
under C0, the comparator the analysis uses; the partner reached 77 under
C4 itself). The authority verdict is supported in both named cells with
intervals well clear of the margin, on a payload held constant to the
sentence.

The same authoritative instruction with the correspondence shuffled
redirected 80 to 81 of 82 receivers in every cell, and 81 of 82 in the
partner cell, where the receiver's own record supported the correct
answer. Correct prior information did not protect the receivers'
answers against a shuffled authoritative mapping. Whether they
evaluated the conflict and nevertheless prioritised the instruction,
or did not evaluate it, is not something this design observes. C4 − C5
is the dependence on mapping correctness under the instruction; it is
not evidence of binding in any mechanical sense.

Translation and authority are close in the named cells and neither
direction met the superiority criterion there (C4 − C1 of +0.12 to +0.13 with lower bounds at
zero; no equivalence claimed). They differ where the receiver has no
history: re-expression achieved 0 of 82 for cold receivers, the
correct authoritative correspondence 74 of 82. Re-expression relies on
context these formed receivers possess. The advisory result is small
but above zero (+0.09 and +0.21; not the 0.25 margin): having the
mapping, without an instruction to prefer it, is worth little to a
receiver with a competing story and a great deal to one without.

## 6. Discussion

**What is demonstrated.** In this setting, authority makes a supplied
correspondence govern interpretation; it does not ensure that
correspondence is true. The same intervention produced both
restoration and coordinated error, on the same receivers and notes,
with the payload held constant. A17's descriptive pattern, that
availability alone was often ineffective for receivers holding a
competing history, is the reason the comparison was worth making; A18
supplies the supported contrasts. We do not present "availability
versus authority" as a fully isolated mechanism: A18's design
separates the authority sentence from the advisory one and the correct
mapping from the shuffled one; it does not separate every factor that
differed between A17's delivery channels and A18's.

**Registry correctness as an upstream dependency.** For this
intervention, the correctness of the supplied correspondence is a
critical upstream dependency: with the correspondence wrong, the
instruction produced near-total misdirection, and the receivers' own
correct records did not protect them. A18 observes the receivers' responses, not whether they checked
anything; it does not show that receivers cannot check authority,
provenance or consistency under a different design.

**Proposed engineering consequences, distinguished from demonstrated
behaviour.** The programme's motivating architecture has a Rules plane
that is meant to bind interpretation across agents with divergent
histories. A18 gives that plane a demonstrated mechanism and a
demonstrated failure mode on one carrier: an explicit authority
instruction over a registry correspondence transfers whatever the
registry holds. The engineering consequence we propose, not
demonstrate, is that an authoritative lookup needs its own way to
establish that its entries deserve authority; receiver compliance
cannot supply that assurance. Whether an executor that mechanically
constrained target selection would behave the same, whether receivers
could be given grounds to check authority, and whether the mechanism
holds across families are open.

**The failure class, placed beside its neighbour.** The act that fails
here is procedurally valid: the receiver holds the authority to act on
either cistern, the world accepts the act, and the task fails because
the act reached the wrong one of two otherwise valid objects. That is
a different failure class from the one measured in our work on an
executable institution (*Where Reliability Lives*, arXiv:2609.03192,
§5.4), where one falsehood delivered as trusted testimony drove
roughly nine hundred futile acts per believing run and the
authoritative ledger refused every one of them. The two results come
from different apparatuses. Together they illustrate two distinct
failure classes, a false premise contained by authority and a wrong
referent that authority does not see; they do not demonstrate that
one deployed gate admits reference errors, and no study has run a
reference error and an authority check on the same apparatus. What
the two records support is narrower: an authority check adjudicates
whether an actor may act on an object, and nothing in either apparatus
adjudicated whether the object was the one the sender meant. The
engineering consequence we propose, not demonstrate, is that this
binding is a check distinct from authority, provenance and
admissibility, and that where it lives in an agent architecture has
to be established rather than assumed.

**Limits.** One carrier, one corpus of previously examined cases, one
instruction pair frozen verbatim without alternatives; the event
dimension only; two draws a probe in the underlying corpus; the CLI
pathway reports no answering model per call; usage is tokens on a
subscription, not a bill; A17's descriptive findings are descriptive.

## 7. Methods

**Carrier and pathway.** gpt-6-astra, `model_reasoning_effort="medium"`,
Codex CLI 0.154.0 (`codex exec --json`) on a subscription login, the
experiment's system prompt replacing the CLI's instructions, 33 tool
surfaces disabled by name, read-only sandbox in an empty directory,
ephemeral sessions, no history, user configuration ignored, the
memories feature off by default. Each call archived with input, cached
input, output and reasoning tokens and wall time; the event stream
carries no model string, so the requested identifier is the accepted
identifier and this limit is stated.

**Apparatus.** The waterworks world; the crossed design with
groups A to E, relations computed from the generated schedules and
asserted before every probe; the post-formation competence gate with
the evaluability rule; the A14.2 attribution classifier; the v2
coherent swap and the irrelevant-message control in R0.

**Pipeline.** For every frozen note, the archived note-level
classification (described member or unresolved) mapped to the
receiver's tag and, for re-expression, to the receiver's group's events
for the member and the receiver's recorded acts on it (two most
recent, chronological; abstentions and off-target acts excluded). A
note describing the other member is repaired to the member described
and scored as a faithful miss. Unresolved notes deliver nothing under
R1, R2, R3 and R3′; R2r's block is present regardless; none of A18's
82 notes is unresolved.

**Analysis.** Seed-level success per cell; paired seed differences over
seeds where both conditions completed; percentile bootstrap, 10,000
resamples, generator seed 20260909; predicates: improvement and
dependence mean ≥ 0.25 with lower bound > 0 (contradicted when the
upper bound ≤ 0); restoration lower bound > −0.15 (contradicted when
the upper bound < −0.15); superiority mean ≥ 0.10 with lower bound > 0
(contradicted when inferiority is established); equivalence whole
interval within ±0.10 (A17 only); inconclusive when fewer than four
paired seeds or a floor is unmet; else other. Cells combine
inconclusive → contradicted → supported → other. A17 floors: 24 event
and 24 act notes, four seeds with both. A18 floor: 48 event-only notes
over six seeds carrying four each, on each contrast's paired material.

**Execution rules.** Sequential seeds, four workers; every attempt
archived and counted against the ceiling; a pass ending without its
record kept, the CLI awaited, restarted once, then recorded incomplete;
a standing five-minute health check; the analysis run once after every
pass; per-seed summaries decide nothing.

## Evidence appendix

**A17.** Frozen text `notes/panel/a17-freeze-proposal.md` (39c9908,
after six reviewer corrections); preregistration entry A17 (2b5f99e),
implementation signature 86d8e06; amendments A17.1 (no pre-formation
gate from seed 109; e20d5bc), A17.2 (two draws; 49b44e6), A17.3 (the
weekly-cap pause and the dated resumption). Records
`results/dialects/history/a17-s108…119-codex-gpt-6-astra`, driver log
`a17-driver.log`; analysis `a17-repair.json` (every cell with success,
wrong member, abstention, unresolved, faithful miss and lookup use;
every contrast with mean, interval, paired seeds and verdict; usage per
condition); R0's crossed analysis `a17-r0-crossed.json`. Seed 109's
gate: a1, a2 answered 8 of 10 evaluable, zero wrong. Composition of a
probed seed: 120 history, 120 notebook, 180 gate, 108 probe calls (seed
108 carried 180 further qualification calls under A17.1).

**A18.** Frozen text `notes/panel/a18-design-proposal.md` v2.1
(eb0da6a); preregistration entry A18 (1090174); zero-call power
scenarios `tools/power_a18.py`. Records
`results/dialects/history/a18-s108…119-codex-gpt-6-astra` (seed 112
recorded incomplete at every stage with zero calls); driver log
`a18-driver.log`; analysis `a18-authority.json`. Per condition, 82
exchanges an arm on ten seeds carrying 12, 10, 6, 10, 5, 7, 11, 8, 7
and 6 event-only notes.

**Qualification audit (2026-09-12).** The carrier was selected by a
qualification whose implemented predicate omitted the written production
threshold and whose drivers applied a stricter invalid-act cutoff than
written; the audit (`tools/audit_qualification.py`,
`a17pre-qualification-audit.json`, the addendum in the preregistration)
finds gpt-6-astra qualifying at both efforts under the written rule and
no other configuration, so the carrier is unchanged.

**Reviewer rulings applied.** A17: the closure as inconclusive without
retrospective event-only verdicts; "descriptively reproduced" not
"replicates"; the shuffled index stated as correct selection reduced to
1 of 82 with outcomes kept apart. A18: the behavioural statement of the
shuffled-authority result; registry correctness as a critical upstream
dependency; the reasoning-token change as roughly 80 % and descriptive;
re-expression's 0 of 82 cold given prominence; the stopping point.
