# Recorded prior-work search — protocol for the grounded-reference programme

*2026-09-12. Written before the search is run, so that the collapse
conditions cannot drift toward the results. The house rule from the
programme's earlier search appendix applies: a claim of absence is a
claim about these queries in these sources on these dates, never
about the literature in general. The search record (queries verbatim,
dates, hits, verdicts) is appended to this file as it is run; verified
citations only, each fetched and read from its record.*

## What the search supports if it comes back clear

Only this sentence, in both manuscripts: "We found no prior study
addressing these specific contrasts in the sources searched, as of
[date]." It does not establish originality or uniqueness.

## The two substantive claims under test, and their collapse conditions

Each collapse condition is written to catch **substantively
equivalent** experiments under any terminology. A study collapses a
claim if it does the equivalent thing, whatever it calls it; it does
not need the event/act apparatus, cisterns, or our instruction
wording.

**Claim 1 (history-and-transfer).** *Which component of a shared
interaction history a reference depends on was separated
experimentally, by varying components of history independently across
receivers whose candidate referents were otherwise indistinguishable,
and the failure at the boundary was measured as a systematic wrong
choice rather than an inability to answer.*

Collapses on any published study in which (a) two or more agents,
human or artificial, form histories through interaction; (b) at least
two distinguishable components of that history (for example, what
happened to an object versus what an agent did with it; observed
versus performed; perceptual versus procedural; or any other
partition) are varied independently across receivers, by design; (c)
a reference produced by one agent is interpreted by receivers in each
resulting relation; and (d) the outcome distinguishes a wrong referent
from no referent. Near-miss (cited, not a collapse): studies varying
shared history as a single quantity (amount, overlap, presence versus
absence), or varying components without independent crossing, or
reporting only success rates without the wrong-referent outcome.

**Claim 2 (authority-and-repair).** *For a receiver whose own
interaction history supports a competing interpretation, an identical
supplied correspondence between a sender's referent and the receiver's
own designation was presented under an advisory and under an
authoritative instruction, with the correspondence correct or wrong,
and the effect on the receiver's choice was measured.*

Collapses on any published study in which (a) a receiver holds a
competing interpretation grounded in its own prior interaction or
experience (not merely a stated professional standard or a
retrieved fact); (b) the same referential or mapping information is
presented with different stated authority; (c) the information is
also presented wrong, under the authoritative framing; and (d) the
receiver's choice among candidates is measured. Near-miss: principal-
or authority-hierarchy studies with factual demands rather than
referential correspondences (Yu et al. 2026 is the known one);
retrieval-trust and tool-output-trust studies without a competing
interaction history; instruction-following studies without a
correctness manipulation.

## Sources and queries

- **Sources:** arXiv (listing and API), Semantic Scholar (API if it
  answers; each refusal recorded), Google Scholar via web search, ACL
  Anthology, OpenReview, DBLP; the reference lists of the three papers
  already cited (Gao et al. 2026; Mohapatra et al. 2026; Yu et al.
  2026) followed one hop.
- **Query families, each run verbatim and recorded:** common ground /
  grounding / conceptual pacts with LLM agents or multi-agent; referring
  expressions private to interaction history; divergent or asymmetric
  histories, memories or experiences between agents; agent handoff,
  substitution, interchangeability, transfer between teams; emergent
  communication with history-dependent or partner-specific reference;
  "wrong referent" / misinterpretation under asymmetric knowledge;
  perceptual versus procedural (observed versus performed) memory in
  reference; authority, advisory versus mandatory, deference to
  supplied mapping or registry; retrieval or tool-output trust against
  parametric or experiential belief; instruction framing and
  correctness of injected context; shuffled or corrupted correspondence
  controls. Classical human-dialogue literature on referential pacts is
  cited as background and is not a collapse condition, since it does
  not vary history components by design across artificial receivers.
- **Verdict vocabulary:** COLLAPSE (a verified publication does the
  equivalent experiment), NEAR-MISS (a verified neighbour that falls
  short on a named clause, cited), CLEAR (nothing found for that query).

## Consequence of a match

A COLLAPSE on either claim changes the contribution claim for **that
result** to a replication or extension of the named prior work, with
the differences stated; it does not reclassify the other result or the
programme. A NEAR-MISS is cited in the related-work section with the
clause it falls short on.

## Record

*(appended as the search is run: date, source, query verbatim, hits
examined, verdict, verified citation where applicable)*

### Run of 2026-09-12 (13:40 to 14:50 NZST; 01:40 to 02:50 UTC)

**Sources answering.** Web search (the assistant's search tool over
the open web, which indexes arXiv listings, ACL Anthology, OpenReview,
DBLP, Frontiers, PMC and publisher pages); the arXiv listing search
page (`arxiv.org/search`); the arXiv abstract pages of every candidate
(each fetched and read); the HTML full text of the three cited papers
for their reference lists (one hop); the Semantic Scholar Graph API
for three of five queries.

**Sources refusing, recorded.** The arXiv API (`export.arxiv.org`)
returned HTTP 429 or timed out on all eight query families from two
hosts (two hosts) at 02:10 to 02:45 UTC; its results
are absent from this record and the arXiv listing page was used
instead. The Semantic Scholar API returned HTTP 429 on queries S2-2 and
S2-5 after five spaced retries each; S2-1 returned four irrelevant
records, S2-3 four records, S2-4 forty of 275. No key was used for
either API. The raw API responses and the query scripts are kept with
the session's scratch files, not in the repository.

**Web-search queries, verbatim, and the hits examined.**

| # | query | hits examined (verified against the record) | verdict |
|---|---|---|---|
| W1 | LLM agents "common ground" divergent interaction histories reference resolution experiment | Yamamoto et al. 2026 (Frontiers, 10.3389/fcomp.2026.1873726); Wang et al. arXiv:2606.08081; PhotoBook 1906.01530 | NEAR-MISS |
| W2 | multi-agent language models partner-specific referring expressions private interaction history transfer to new partner | Wang et al. 2606.08081; Zeng et al. 2601.19792 | NEAR-MISS |
| W3 | "conceptual pacts" large language model agents experiment referring | Hough et al. LREC-COLING 2024 (human-corpus pact models) | CLEAR for Claim 1 |
| W4 | asymmetric memory agents misinterpret reference wrong referent experiment language models shared history | Li, Gatt and Poesio 2606.31719 (SIGDIAL 2026); Talebirad et al. 2607.00233; Li et al. 2608.19564 | NEAR-MISS |
| W5 | emergent communication partner-specific conventions transfer new partner misinterpretation study 2025 2026 | Kim 2606.10582; Jones et al. 2602.08208; Hawkins et al. 2104.05857 (via the arXiv listing) | NEAR-MISS |
| W6 | observed versus performed events episodic memory reference resolution artificial agents experiment | human neuroscience and agent-memory surveys only | CLEAR |
| W7 | authoritative versus advisory instruction language model override own memory injected context experiment | Shih, Winnicki and Cao 2609.00753; Leong 2608.28502; Piehl et al. 2602.15344 | NEAR-MISS |
| W8 | "knowledge conflict" LLM context authority framing incorrect retrieved information deference instruction | Schuster, Gautam and Markert 2601.03746; Sun, Bai and Dredze 2506.06485 (ACL 2026); Zhang and Lin 2602.04918; surveys | NEAR-MISS |
| W9 | LLM agent trust tool output versus own experience conflicting mapping wrong tool output followed experiment | Yang et al. 2609.05587; Zhang et al. 2606.21409; Yan et al. 2605.17453; 2606.14476 | NEAR-MISS |
| W10 | agent handoff shared registry mapping identifiers across agents authoritative lookup experiment language models | registry specifications and vendor documentation only; Margalit et al. 2606.24535 | CLEAR |
| W11 | "context-memory conflict" language models 2025 2026 instruction prioritize context correctness | Sun, Bai and Dredze 2506.06485; Shih et al. 2609.00753 | NEAR-MISS |
| W12 | multi-agent LLM teams different histories coordination cost swapping agents conventions 2026 | Gao et al. 2609.05279 (already cited) | NEAR-MISS (cited) |
| W13 | "wrong referent" language model agents asymmetric knowledge misinterpretation confident incorrect resolution | Fenoglio 2607.28137 (position paper); 2407.11789 (misleading assistants) | CLEAR |
| W14 | advisory versus mandatory instruction language model compliance framing "advisory" experiment supplied mapping | framing-effect and pragmatic-influence studies (2603.19282, 2602.21223, 2608.27340); none with a supplied mapping | CLEAR |
| W15 | shuffled OR corrupted mapping control authoritative context LLM follows wrong mapping despite own correct knowledge experiment | Zhang and Lin 2602.04918; a medRxiv clinical-cue study | NEAR-MISS |
| W16 | LLM agents episodic memory divergent experiences same environment different agents interpret instruction differently experiment "shared experience" | memory surveys only | CLEAR |
| W17 | openreview language agents common ground interaction history referring expression interpretation partner swap | Wang et al. 2606.08081; Wang et al. 2509.11514 (overhearing) | NEAR-MISS |
| W18 | aclanthology 2026 LLM dialogue "shared history" reference resolution partner-specific asymmetric | Li, Gatt and Poesio, SIGDIAL 2026 | NEAR-MISS |
| W19 | dblp "common ground" "language model" agents interaction history reference 2026 | Mohapatra et al. 2601.09365 (Findings of ACL 2026; already cited) | NEAR-MISS (cited) |
| W20 | openreview.net authority framing instruction language model context memory conflict wrong correct 2026 | Sun, Bai and Dredze; CUB 2505.16518 | NEAR-MISS |
| W21 | perceptual versus procedural memory reference interpretation LLM agents "what happened" versus "what I did" history | agent-memory taxonomies only | CLEAR |
| W22 | agent follows supplied instruction over its own observation wrong map versus correct map embodied LLM agent experiment "own observation" deference | Fu et al., EACL 2026 (MIRAGE); Wang et al. 2604.17252 (belief inertia) | NEAR-MISS |
| W23 | LLM agent handoff context transfer between agents misinterpretation different memory "handoff" experiment 2026 arxiv reference | memory and context-management surveys | CLEAR |
| W24 | "private notebook" OR "private memory" agents Lewis signaling game receiver different history misinterprets signal wrong object arxiv | Talebirad et al. 2607.00233 and 2608.17053 | NEAR-MISS |

**arXiv listing-page queries** (`arxiv.org/search`, all fields, newest
first, 50 a page): `"common ground" "language model" agents history`
(1 result, Social-RAG 2411.02353: CLEAR); `"partner-specific" language
model` (4 results: 2608.08443, 2311.13061, 2104.05857, 2002.01510:
NEAR-MISS on the Hawkins papers); `authoritative advisory instruction
"language model" memory` (no results); `"knowledge conflict" authority
instruction context memory` (no results).

**Semantic Scholar queries, verbatim.** S2-1 "language model agents
shared interaction history reference interpretation wrong referent"
(4 records, none relevant: CLEAR); S2-2 "advisory versus authoritative
instruction supplied correspondence competing memory language model"
(refused, 429); S2-3 "partner-specific conventions transfer new partner
LLM agents referring expressions" (4 records: Wang et al. 2606.08081,
PhotoBook 2018, Hawkins et al. 2104.05857, one irrelevant: NEAR-MISS);
S2-4 "asymmetric memory agents misinterpret reference language model"
(40 of 275 records examined by title; 2609.05339 memory portability
across model upgrades and 2607.06157 deliberation under partial
observability fetched and read; neither varies history components or
interprets one reference across receivers: CLEAR); S2-5 "tool output
trust corrupted correct instruction framing agent knowledge conflict"
(refused, 429).

**One-hop reference lists.** Gao et al. 2609.05279 cite, on
conventions and histories: Wang et al. 2606.08081; Ashery, Aiello and
Baronchelli, Science Advances 2025 (arXiv:2410.08948); Ko and Geiping
2606.30571; Shih, Sawhney, Kondic, Ermon and Sadigh, ICLR 2021
(arXiv:2104.02871); the transactive-memory literature. Mohapatra et al.
2601.09365 cite the grounding literature (Clark and Schaefer 1989;
Traum and Allen 1994; Kruijt and Vossen 2022) and their own 2024
grounding-act papers. Yu et al. 2605.12120 cite Wallace et al. 2024
(arXiv:2404.13208), Sharma et al. ICLR 2024 on sycophancy, and
value-action-gap work. Each was checked against the collapse
conditions: NEAR-MISS or background, none a COLLAPSE.

**Verdicts by claim.**

*Claim 1 (history-and-transfer): CLEAR of COLLAPSE; the near-misses,
each with the clause it falls short on.*

- Wang, Mishra, Özyürek, Rubio-Fernández and Ghaleb, arXiv:2606.08081
  (June 2026): a pseudo-dyad baseline that breaks partner history in
  repeated reference games and finds multimodal agents coordinate
  without partner-specific convention. Falls short on clause (b):
  history is varied as present versus broken, not as independently
  crossed components; and on (d): no wrong-referent outcome for a
  receiver with a different history.
- Gao, Yu, Deng, Li and Wang, arXiv:2609.05279 (September 2026),
  already cited: swaps agents across formed teams and measures the
  coordination cost. Falls short on (b) and (c): history is a whole,
  and no reference is interpreted across relations.
- Talebirad, Redman, Parsaee and Zaiane, arXiv:2607.00233 (June 2026)
  and arXiv:2608.17053 (August 2026): Lewis signalling with private
  notebooks; memory architecture against channel capacity. Falls
  short on (b) and (c).
- Li, Gatt and Poesio, arXiv:2606.31719 (SIGDIAL 2026) and
  arXiv:2511.03718: models as overhearers of asymmetric human dialogue
  conflate potential with established common ground; a perspectivist
  annotation of speaker and addressee interpretations. Falls short on
  (a) and (b): the artificial receivers do not form the histories, and
  the manipulation is of context access, not of crossed history
  components; the wrong-versus-no-referent distinction is in the
  annotation, not in an artificial receiver's act.
- Zeng et al., arXiv:2601.19792 (January 2026) and Wang et al.,
  arXiv:2509.11514 (2025): director–matcher and overhearer designs
  with vision-language models. Falls short on (b) and (d).
- Jones, Lombardi, Mahowald and Bergen, arXiv:2602.08208 (February
  2026): convention formation in same-type and mixed dyads. Falls
  short on (b), (c) and (d).
- Kim, arXiv:2606.10582 (June 2026): zero-shot mutual intelligibility
  across independently trained populations. Falls short on (b) and (d).
- Hawkins, Franke, Frank, Goldberg, Smith, Griffiths and Goodman,
  arXiv:2104.05857 (Psychological Review, 2021) and arXiv:2002.01510:
  a hierarchical Bayesian account of partner-specific against
  community conventions, with human data. Background; falls short on
  (a) for artificial receivers and on (b).
- Shih, Sawhney, Kondic, Ermon and Sadigh, ICLR 2021
  (arXiv:2104.02871): separates rule-dependent from
  convention-dependent representation for adapting to new partners.
  Nearest in spirit to a component separation; falls short on (b) as
  an experimental crossing of history components between receivers of
  one reference, and on (c) and (d).
- Yamamoto, Morita, Higashinaka and Takeuchi, Frontiers in Computer
  Science 2026: communicative success dissociated from
  representational alignment in agent–agent common-ground formation;
  receiver modules fixed, no crossed histories. Falls short on (b)
  and (c).
- Ashery, Aiello and Baronchelli, Science Advances 2025
  (arXiv:2410.08948): emergent conventions in LLM populations. Falls
  short on (b), (c) and (d).

*Claim 2 (authority-and-repair): CLEAR of COLLAPSE; the near-misses,
each with the clause it falls short on.*

- Yu, Seedat, Schwarz and Bean, arXiv:2605.12120 (May 2026), already
  cited: principal changed, content constant, compliance despite
  demonstrated knowledge. Falls short on (a): the competing
  interpretation is a professional standard, not an interaction
  history; and on (c) as a referential correspondence.
- Shih, Winnicki and Cao, arXiv:2609.00753 (September 2026): prompts
  that direct a model to prioritise supplied context or parametric
  knowledge; authority directions in activations. Falls short on (a):
  parametric knowledge, not interaction history; (c) is not a
  correctness manipulation of a mapping; no candidate choice (d).
- Yang, Song, Kim, Song, Park and Jo, arXiv:2609.05587 (September
  2026): corrupted tool returns adopted at high rates; user-prompt
  interventions (compare, verify, disclose) that do not consistently
  help. Falls short on (a): parametric knowledge; and on (b): the
  prompt policies concern conflict handling, not the stated authority
  of one identical payload.
- Zhang et al., arXiv:2606.21409 (June 2026): faithful, misleading and
  absent feedback in a matched loop. Falls short on (a) and (b).
- Yan et al., arXiv:2605.17453 (May 2026): cognitive poisoning by a
  tool that earns trust. Falls short on (a), (b) and (d).
- Schuster, Gautam and Markert, arXiv:2601.03746 (January 2026):
  source credibility in inter-context conflicts. Falls short on (a)
  and (c).
- Sun, Bai and Dredze, arXiv:2506.06485 (ACL 2026): context–memory
  conflict across task types. Falls short on (a) and (b).
- Wallace, Xiao, Leike, Weng, Heidecke and Beutel, arXiv:2404.13208
  (April 2024): the instruction hierarchy; and Leong, arXiv:2608.28502
  (August 2026): recognition without enforcement of instruction
  source. Falls short on (a), (c) and (d).
- Fu, Qiu, Wang, Sansom, Ayyappa Prabhu, Tang, Kim, Sohn and Lee,
  EACL 2026 (MIRAGE): imperfect guidance from manuals, retrieval and
  prior interaction; robustness under incorrect guidance. Nearest on
  (c) in that guidance is varied in correctness; falls short on (a)
  as a competing interpretation formed in interaction and on (b),
  stated authority. Wang, Leong, Wang and Li, arXiv:2604.17252
  (belief inertia in embodied agents): falls short on (b) and (c).
- Zhang and Lin, arXiv:2602.04918; Piehl et al., arXiv:2602.15344
  (memory injection attacks): falls short on (a) or (b), and on (d).

**What the record supports.** In the sources searched, as of 12
September 2026, we found no prior study addressing these specific
contrasts: for Claim 1, components of interaction history crossed by
design across artificial receivers of one reference, with the wrong
referent distinguished from no referent; for Claim 2, one identical
supplied correspondence under advisory and authoritative instruction,
correct and wrong, against a competing interpretation formed in the
receiver's own interaction history. This is a claim about these
queries in these sources on this date, with two APIs refusing, and it
does not establish originality. The nearest neighbours above are cited
in the related-work sections with the clause each falls short on.

**Classical citations, bibliographically verified the same day.**
Clark and Wilkes-Gibbs (1986), Referring as a collaborative process,
Cognition 22(1), 1–39. Brennan and Clark (1996), Conceptual pacts and
lexical choice in conversation, Journal of Experimental Psychology:
Learning, Memory, and Cognition 22(6), 1482–1493. Clark and Brennan
(1991), Grounding in communication, in Resnick, Levine and Teasley
(eds.), Perspectives on Socially Shared Cognition, APA, 127–149. Lewis
(1969), Convention: A Philosophical Study, Harvard University Press.
