Research · Overview
When agents share a history, and when they don’t: reference, transfer and repair in a small world
Plain-language overview of the grounded-reference programme, September 2026. It accompanies the research note History-grounded reference in language-model agents: formation, transfer across divergent histories, and repair under authority, its two bounded reports, and their evidence package. Every number here is in the note and its archived records; nothing here goes beyond them.
The one-paragraph version
In these experiments, reference depended on which parts of an interaction history the receivers shared. Two language-model agents that worked twelve days together could refer to one of two identical cisterns by what had happened to it, and their partner understood. A receiver whose history was different did not simply fail to understand: it confidently picked the wrong cistern, following the reference into its own history. Which receivers understood depended on which component of history they shared, the events the cistern had suffered or the acts done to it. Explicitly instructing a receiver to prioritise a supplied correspondence could restore interpretation across different histories, but the same instruction with a wrong correspondence redirected even a receiver whose own record held the right answer. The results hold for the models and the task tested, and we say so at every step.
The setup
Picture a small yard with six water cisterns. Two of them, large timber cisterns by the smithy, are identical in every way a worker can see. Two agents, played by a language model, work the yard for twelve scripted days. One of the pair leaks and is repaired; the other is tainted and scoured. The agents inspect them, fill them, draw from them. Each agent keeps its own notebook and reads the clerk’s record of its own receipts, and each sees the cisterns under private tags that mean nothing to anyone else.
Then one agent is asked to send its partner a note about one of the two identical cisterns. It cannot point, it cannot use the partner’s tags, and the cisterns look the same. The only thing that can single one out is history: “the one that leaked on Day 3”, “the one you filled on Day 5”.
We send the same note, unchanged, to several receivers: the partner, strangers whose groups lived through some, all or none of the same history, and a receiver with no history at all. Then we watch which cistern each one acts on.
What we found
Reference formed, every time we tried. On the model configuration we studied first, the sender wrote history-grounded notes in every one of sixteen fresh histories, and the partner resolved them.
Crossing the history boundary reverses the reference. A stranger with a different history did not shrug. It resolved the note against its own history and acted on the wrong cistern almost every time: the seed-level transfer penalty was 0.95, and across 168 outsider acts the outsider chose the wrong member 159 times. A receiver with no history mostly declined to act.
Which history matters depends on the reference. In a design that crossed the two components of history on purpose, notes that referred to events (the leak, the taint) resolved in every class of receiver whose history held the event, about nine times in ten within each class (154 of 174), and for no receiver whose history did not (0 of 174). Notes that referred to acts resolved mostly for receivers whose history held the act. A stranger whose group had run the same planned schedule matched the partner on event references and fell about 15 points short on act references. So sharing a history transfers reference across agents to a substantial degree, but not completely, and what needs sharing depends on what the note refers to.
Other models produce it too, but not all of them. A second model produced history-grounded references in nearly every note and met the preregistered bar in two separate runs. A third produced them at a lower rate. Three others, at every setting we tried, passed explicit tests of their own records yet produced few or none, and two of them wrote instead that the cisterns were indistinguishable. Knowing your history and using it to refer are separable.
Repair: having the answer is not enough. For the second model, we asked what would let a receiver with the wrong history understand. If we rewrote the note in the receiver’s own history, uninformed receivers went from about 7 percent success to about 75 percent. If we simply supplied the correct answer, as a sentence, a retained note or a consultable index, receivers with no history used it, and for receivers with a competing history it was often ineffective: they followed their own history instead. These repair findings are descriptive: cross-family production was observed and event transfer was descriptively reproduced, but the model wrote too few act references to meet the preregistered coverage floor, which prevented a formal verdict.
Authority changes everything, in both directions. We then gave the same correct correspondence under two instructions that differed by one sentence. Marked advisory, it left uninformed receivers near their starting point: 16 and 18 successes of 82. Marked authoritative over the receiver’s own record, it brought them to 76 and 77 of 82, level with the informed partner. And when we shuffled the correspondence so that it named the wrong cistern and marked it authoritative, 80 or 81 of 82 receivers followed it in every group, including the partner who already knew the right answer from its own record. In this setting, authority makes a supplied correspondence govern interpretation; it does not ensure the correspondence is true.
What this does and does not show
It shows, on the tested models and this task, that an agent’s own interaction history can become part of what a message requires to be understood; that the requirement is specific to the kind of history referenced; that supplying a correct correspondence does little against a receiver’s competing history unless the receiver is told to prioritise it; and that the instruction is followed whether or not the correspondence is right.
It does not show a general theory of reference, or anything about models and tasks we did not run. The broad ideas, that shared history shapes communication and that instructions can override what a model knows, are already in the literature; the reports cite the closest prior work. What is specific here is the controlled separation of which history supports which reference, the measured form of the failure at the boundary, and the measured difference between advisory and authoritative presentation of the same correspondence.
Why it matters for building agent systems
If agents that worked apart are to hand work to each other, or to draw on a shared registry of what things are called, three things follow from these results. A registry that merely makes the right correspondence available may go unused by an agent that has its own story. An instruction to treat the registry as authoritative gets it used, and gets it used when it is wrong. So the correctness of what the registry holds becomes a critical dependency of the whole arrangement, and an authoritative lookup needs its own way of showing that its entries deserve authority. Those are engineering consequences we propose; the experiments demonstrate the behaviour, not the fix.
How to check it
Every study was preregistered before its first call, with every amendment, deviation and audit dated in the record. Every reading was computed once by a fixed rule from archived records, and the evidence package lets a reader regenerate every figure from the committed data with one command per study. Where a rule was implemented differently from how it was written, the audit says so, keeps the original decisions, and shows the readings under both.
The note, its two bounded reports, the recorded prior-work search, the evidence package and the manifest are at taniwha.ai/research/grounded-reference.
Published 2026 — Taniwha AI. This is the companion read; the note itself, with its two bounded reports, the recorded prior-work search, the evidence package and the manifest, is at History-grounded reference in language-model agents.