# Related-work search appendix — protocol and record (2026-08-26)

*This appendix is the contemporaneous search record behind every
"closest found" and "no prior found" claim in "Where Reliability
Lives." It was produced while the results were organised as a
two-paper pair; the internal labels are preserved as written rather
than edited after the fact — "Paper 1" is now Part I of the unified
paper, "Paper 2" is now Part II, and section numbers inside those
labels refer to the pair's drafts.*

This document is the search appendix the paper cites. It records the
protocol under which the related-work sections of Paper 1 ("Epistemic
channels of an executable institution…", v0.4) and Paper 2 ("Durable
institutions, replaceable minds", v0.1) were built, and scopes every
"no prior found" or "closest found" claim to the searches actually
conducted. The house rule applies: a claim of absence is a claim about
these queries on this date, never about the literature in general.

## Protocol

- **Date of all searches:** 2026-08-26.
- **Method:** four parallel research passes (web search + page fetch),
  one per area group. **Verification rule: a citation enters either
  paper only if its abstract/landing page was fetched and its metadata
  (title, authors, year, venue, identifier) read from the page** — no
  citation from memory, no unfetched entries. Candidates that failed
  verification are recorded as such, not silently dropped.
- **Selection rule:** 4–7 strongest works per area; relevance judged
  against the papers' actual claims, not keyword overlap.
- **Query logging and accounting:** this appendix reproduces **85
  distinct query strings verbatim** (Group 1: 12; Group 2: 28;
  Group 3: 19; Group 4: 20; Group 5: 6). Additional follow-up
  searches performed during adjudication (Groups 1 and 3) were not
  contemporaneously enumerated; accordingly **no total count of
  search executions is claimed** — the verbatim lists, sources,
  dates, and closest-found adjudications are the search surface, not
  a count. The two concepts are kept distinct throughout: the
  enumerated lists are *verbatim query strings preserved in the
  contemporaneous agent record*, not a tally of executions. Group 2's
  list, restored in full after an external reviewer noted the earlier
  condensed listing, carries a line-level provenance receipt —
  timestamp, record-line SHA-256, and an independently re-run
  mechanical occurrence check (28/28) — published as Receipt C of the
  accompanying verification receipts.

## Area groups

1. **Paper 1 corpus verification + LSU no-prior search** — exact
   verification of the six work-families already cited informally
   (PBRC; agent-society validity; LLM network epistemology; Zollman;
   LSU/LSU-E and the two identifiers carried since v0.2; selective
   prediction), plus the registered no-prior-computational-replication
   search for Linear Sequential Unmasking with full query log.
2. **LLM-agent robustness/evaluation + tool-use verification and
   execution-grounded agents** (Paper 2 areas 1–2), with the novelty
   probe: any work experimentally separating cognition from an
   enforcement layer and substituting/ablating cognition while scoring
   pre-declared institutional invariants against an authoritative
   world.
3. **Constitutional/governance architectures + institutional
   approaches to multi-agent control** (Paper 2 areas 3–4), including
   the classic normative-MAS / electronic-institutions line, with the
   probe: any cognition ablation/substitution under a fixed
   institution with scored institutional properties.
4. **Externalised memory / durable task state + multi-agent
   coordination under unreliable cognition** (Paper 2 areas 5–6), with
   the probe: any deliberate corruption of agent beliefs/context whose
   containment is measured at an enforcement boundary rather than by
   output quality.

## Findings

### Group 1 — Paper 1 corpus verification + LSU no-prior search

**Verification outcomes.** All six existing work-families verified by
page-fetch, with three substantive discrepancies found and fixed in
Paper 1 v0.5:

1. **arXiv:2511.14098 descriptor mismatch.** The paper is Jain,
   Krishnamurthy & Zhang (2025), "Collaborative QA using Interacting
   LLMs…" — hallucination contagion through LLM networks; it never
   uses "conformity" or "wrong-but-sure cascades". Those descriptors
   belong to Han, Tan, Yu, Zheng & Tang (2026), "Conformity Dynamics
   in LLM Multi-Agent Systems…", arXiv:2601.05606. Fix: cite both,
   each for what it shows.
2. **arXiv:2603.00113 characterisation.** Li & Tao (2026), "AI Agents
   Alone Are Not (Yet) Sufficient for Social Simulation." The draft's
   "mechanistic, counterfactual, collectively scored evaluation" is
   not supported by the abstract; softened to the abstract-supported
   reading (role-play plausibility ≠ behavioural validity; outcomes
   dominated by environment/protocol/information mechanisms, which
   must be explicit and auditable).
3. **Identifier resolution.** PII S2589871X22000018 = Quigley-McBride,
   Dror, Roy, Garrett & Kukucka (2022), the LSU-E casework *tool*
   paper (FSI: Synergy 4:100216) — distinct from LSU-E itself (Dror &
   Kukucka 2021, FSI: Synergy 3:100161). DOI 10.1126/science.aat8443 =
   Dror (2018), "Biases in forensic experts," Science 360(6386):243.
   Both now cited in their own roles.

Verified without discrepancy: Zollman 2007 (Phil. Sci. 74(5):574–587),
2010 (Erkenntnis 72(1):17–35), 2013 (Phil. Compass 8(1):15–27); Dror
et al. 2015 (J. Forensic Sci. 60(4):1111–1112); Alqithami 2026
(arXiv:2604.15558, PBRC); Geifman & El-Yaniv 2017 (NeurIPS 30);
El-Yaniv & Wiener 2010 (JMLR 11:1605–1641). Access notes: philpapers
and Springer block fetchers; paywalled records verified via the
CrossRef API, arXiv, JASSS, JMLR, papers.neurips.cc, ACL Anthology,
OJP and Europe PMC.

**Gap-fill additions (verified):** Windrum, Fagiolo & Moneta 2007
(JASSS 10(2):8 — ABM empirical validation); Wen et al. 2025 (TACL
13:529–556 — LLM abstention survey; arXiv:2407.18418); Dror 2026 (FSI:
Synergy 12:100657 — LSU-E boundary conditions); and, on the
evidence-stream axis, O'Connor & Weatherall 2018 (EJPS 8:855–875 —
provenance-conditioned evidence discounting) and Weatherall, O'Connor
& Bruner 2020 (BJPS 71(4):1157–1186 — selective sharing/publication
manipulating the evidence stream without touching topology). Nothing
was found implementing admissibility gating as an *enforced
institutional rule* in a collective-epistemology model; nearest is
PBRC.

**LSU no-prior search — verdict: zero computational replications
found.** Closest hits, each verified and adjudicated: Whitehead,
Williams & Sigman 2022 (Forensic Chemistry 29:100426 — LSU embedded in
a decision-theoretic lab workflow, executed by three human analysts);
Cuellar, Mauro & Luby 2022 (JRSS-A 185(S2):S620–S643 — probabilistic
simulation of bias propagation, no countermeasure implemented
in-silico); Dror 2026 (analytical, no simulation); Rojas Alfaro et al.
2025 (FSI: Synergy 10:100569 — human field implementation); Stewart &
Kukucka 2025 (Behav. Sci. 15:1094 — "simulated" = tasks for N=149
human participants); Hefetz 2025 (FSI: Synergy 11:100645 — review).

**Query log (Group 1, verbatim, 2026-08-26):**
`"linear sequential unmasking" simulation` · `"linear sequential
unmasking" computational model` · `"sequential unmasking" agent-based`
· `forensic "contextual bias" simulation agent` · `"contextual
information management" computational forensic model` · `LSU forensic
"agent-based model"` · `"linear sequential unmasking" "LLM" OR "large
language model" agents` · `"sequential unmasking" OR "evidence lineup"
in-silico institution simulation ground truth scoring` · `forensic
examiner "agent-based" simulation cognitive bias procedure` ·
`"linear sequential unmasking" replication study 2024 OR 2025 OR 2026
simulation` · `"network epistemology" LLM agents conformity cascade
arXiv` — plus, for the admissibility-axis gap-fill: `agent-based model
manipulating evidence "admissibility" OR "provenance" collective
epistemology network` and follow-ups.

### Group 2 — LLM-agent evaluation + tool-use verification (Paper 2 areas 1–2)

**Verified roster (state-scored evaluation):** τ-bench (Yao, Shinn,
Razavi & Narasimhan 2024, ICLR 2025, arXiv:2406.12045 — database
end-state scoring, pass^k) and τ²-bench (Barres et al. 2025,
arXiv:2506.07982); WebArena (Zhou et al. 2023, ICLR 2024,
arXiv:2307.13854); OSWorld (Xie et al. 2024, NeurIPS D&B,
arXiv:2404.07972); Agent-Diff (Pysklo, Zhuravel & Watson 2026,
arXiv:2602.11224 — state-diff contracts); AgentDojo (Debenedetti et
al. 2024, NeurIPS D&B, arXiv:2406.13352); InjecAgent (Zhan et al.
2024, ACL Findings, arXiv:2403.02691); AgentPoison (Chen, Xiang,
Xiao, Song & Li 2024, NeurIPS, arXiv:2407.12784). AgentBench (Liu et
al., ICLR 2024, arXiv:2308.03688) verified but not selected
(capability-focused).

**Verified roster (verification/enforcement layers):** ToolEmu (Ruan
et al. 2023, ICLR 2024 spotlight, arXiv:2309.15817); CaMeL
(Debenedetti et al. 2025, arXiv:2503.18813); Progent (Shi et al.
2025, arXiv:2504.11703); AgentSpec (Wang, Poskitt & Sun 2025, ICSE
2026, arXiv:2503.18666); VeriGuard (Miculicich et al. 2025,
arXiv:2510.05156); FORGE (Palumbo et al. 2026, arXiv:2602.16708);
SWE-agent's guardrail ablation (Yang et al. 2024, NeurIPS,
arXiv:2405.15793); the specification/verification/enforcement survey
(Dantas, Cordeiro, Nowroozi & Tihanyi 2026, arXiv:2608.14590 — incl.
the "verifier tax" finding).

**Probe verdict: no verified work combines** (i) a persistent
authoritative world with adjudicated acts and typed refusal, (ii)
pre-declared institutional invariants including duty recovery / zero
duplicate accepted work / zero false completions, and (iii)
systematic cognition ablation, kill-reset, frozen-LLM substitution
and testimony corruption scored against those invariants. Ranked
closest: **Gode & Sunder 1993** (J. Political Economy 101(1):119–137
— random-bid programs substituted for human traders in a fixed
double-auction; allocative efficiency shown to derive from
institutional structure: the substitution-under-fixed-institution
schema with one aggregate outcome and none of (i)'s machinery);
runtime policy enforcement for MCP agents (Wang, Zhu & Li 2026,
Electronics 15(13):2829 — scripted replay eliminating model-side
nondeterminism as a controls move); CaMeL/FORGE (invariance by
construction/proof, not by experiment); AMELI (architecture, not
experiment); Chupilkin 2026 (arXiv:2608.04020 — the exact transpose:
agents fixed, institution varied); Concordia (Vezhnevets et al. 2023,
arXiv:2312.03664 — adjudication layer exists but is itself an LLM);
Generative Agents (Park et al. 2023, UIST — ablates cognition
components, scores believability); AgentRewind (Zhuang et al. 2026,
arXiv:2608.14380 — recovery by restoring the cognition's own
checkpoint, the contrast to duty recovery from institutional history
by a fresh cognition).

**Query log (Group 2, verbatim, 2026-08-26 — all 28, restored in full
from the pass's own record after an external reviewer noted the
earlier condensed listing):**
`tau-bench benchmark tool-agent-user interaction database state evaluation` ·
`benchmark LLM agents scored against final environment state execution-based evaluation 2024 2025` ·
`prompt injection benchmark LLM agents AgentDojo InjecAgent 2025` ·
`LLM agent memory poisoning corruption attack benchmark AgentPoison` ·
`runtime enforcement LLM agent tool calls policy refuses invalid actions structured errors` ·
`formal methods verification LLM agent action safety guardrails execution layer 2025` ·
`separate planner from enforcement layer LLM agent substitute cognition invariants preserved` ·
`persistent multi-agent LLM simulation society institutions append-only ledger world state adjudication` ·
`electronic institutions AMELI middleware enforce institutional rules regardless of agent internal architecture` ·
`replace LLM agent with scripted or random policy ablation environment invariants still hold evaluation harness` ·
`"frozen" LLM substituted cognition layer experiment institutional invariants persist simulation` ·
`SWE-agent agent-computer interface guardrails linter rejects invalid edits error messages` ·
`"institutional invariants" LLM agents evaluation benchmark world state` ·
`agent killed restarted recovers pending obligations from event log durable execution LLM benchmark` ·
`Ding 2026 long-running agents restart after crash resume state arXiv` ·
`ToolEmu "Identifying the Risks of LM Agents" ICLR 2024 spotlight OpenReview` ·
`AgentDojo NeurIPS 2024 Datasets and Benchmarks track accepted` ·
`OSWorld NeurIPS 2024 accepted benchmark Xie` ·
`"SWE-agent" NeurIPS 2024 accepted Yang Jimenez` ·
`"Generative Agents: Interactive Simulacra of Human Behavior" UIST 2023 ACM proceedings` ·
`Esteva Rodriguez-Aguilar Rosell Arcos "AMELI" AAMAS 2004 electronic institutions middleware` ·
`tau-bench ICLR 2025 OpenReview accepted Sierra` ·
`WebArena ICLR 2024 accepted Zhou realistic web environment` ·
`CaMeL "Defeating Prompt Injections by Design" accepted venue security conference` ·
`"Progent" "privilege control" AI agents paper accepted conference venue` ·
`"cognition substitution" OR "substituting cognition" agents institutional layer invariants` ·
`security properties hold regardless of LLM behavior model-agnostic guarantees agent system enforcement` ·
`Gode Sunder 1993 "zero-intelligence traders" "market as a partial substitute for individual rationality" Journal of Political Economy`.

### Group 3 — constitutions + institutional multi-agent control (Paper 2 areas 3–4)

**Verified roster (constitutions/governance):** Constitutional AI (Bai
et al. 2022, arXiv:2212.08073); Collective Constitutional AI (Huang et
al. 2024, FAccT, arXiv:2406.07814); NeMo Guardrails (Rebedea et al.
2023, EMNLP demo); TrustAgent (Hua et al. 2024, EMNLP Findings,
arXiv:2402.01586); AgentSpec (above); deontic runtime governance
(Joshi, Finin, Joshi & Kagal 2026, IEEE Agentic Services,
arXiv:2606.19464); Law Informs Code (Nay 2022/2023, arXiv:2209.13020).

**Verified roster (institutional tradition):** electronic institutions
(Esteva, Rodríguez-Aguilar, Sierra, Garcia & Arcos 2001, LNCS 1991);
AMELI middleware (Esteva, Rosell, Rodríguez-Aguilar & Arcos 2004,
AAMAS, DOI 10.1109/AAMAS.2004.10060 — verified via DBLP; IIIA PDF
unparseable); normative MAS (Boella, van der Torre & Verhagen 2006,
CMOT 12:71–79); MOISE+ (Hübner, Sichman & Boissier 2002, LNCS 2507);
MAIA (Ghorbani, Bots, Dignum & Dijkema 2013, JASSS 16(2):9); nADICO
(Frantz, Purvis, Nowostawski & Savarimuthu 2013, LNCS 8291);
Institutional AI / Cournot governance graphs (Bracale Syrnikov et al.
2026, arXiv:2601.11369 — append-only governance log; explicitly
descends from Esteva/AMELI/Boella); GovSim (Piatti et al. 2024,
NeurIPS, arXiv:2404.16698). Available spares: Ye & Steinhardt 2026
(arXiv:2607.09766); He et al. 2024 (COINE@AAMAS, arXiv:2403.16517).

**Probe verdict: not found.** Closest, with reasons: GovSim
substitutes many LLMs under fixed rules but its environment enforces
nothing and its scored outcomes collapse with cognition (the
complement of the paper's claim); Governance-as-a-Service (Gaurav,
Heikkonen & Chaudhary 2025, arXiv:2508.18765) injects adversarial
agents under a fixed enforcement wrapper but scores
enforcement-performance metrics, not invariants of an authoritative
world; Institutional AI varies the governance regime, not the
cognition (confirmed by targeted full-text fetch: no
weaker/random/corrupted-agent arms, no institution-side properties
scored); Chupilkin 2026 is the transpose; the EI-era tradition
asserted cognition-independence architecturally and never ablated it
experimentally.

**Query log (Group 3, verbatim, 2026-08-26):** 19 queries as run,
including: `Constitutional AI: Harmlessness from AI Feedback Bai 2022
arXiv` · `NeMo Guardrails programmable rails LLM Rebedea arXiv
2310.10501` · `runtime enforcement safe LLM agents AgentSpec rule
enforcement 2025` · `LLM agent governance runtime constitution policy
layer 2025 arXiv` · `Collective Constitutional AI public input
language model behavior FAccT 2024` · `"Law Informs Code" Nay legal
informatics approach aligning AI` · `AMELI middleware execution
electronic institutions Esteva AAMAS 2004` · `normative multiagent
systems Boella van der Torre Verhagen introduction survey` · `MOISE+
organisational model multiagent systems Hubner Sichman Boissier
deontic specification` · `MAIA Ghorbani framework agent-based
modelling institutional analysis Ostrom JASSS 2013` · `Frantz nADICO
nested grammar of institutions computational modelling ADICO` · `norm
enforcement LLM multi-agent systems institutions sanctions 2025 arXiv`
· `GovSim Cooperate or Collapse sustainable cooperation society LLM
agents NeurIPS 2024` · `TrustAgent agent constitution safe trustworthy
LLM agents arXiv` · `"electronic institutions" LLM agents revival
norms regimentation 2024 2025` · `Esteva "on the formal specification
of electronic institutions" Springer` · three probe queries
(`experiment vary agent architecture fixed institution norm compliance
electronic institutions simulation evaluate` · `"same institution" OR
"fixed institution" substitute different LLMs ablate agent reasoning
score institutional properties compliance` · `normative multi-agent
systems compare agent populations norm-compliant violating agents
enforcement effectiveness simulation study`).

### Group 4 — durable state + unreliable cognition (Paper 2 areas 5–6)

**Verified roster (externalised memory / durable state):** MemGPT
(Packer et al. 2023, arXiv:2310.08560); agent-memory survey (Zhang et
al. 2024, arXiv:2404.13501); Mem0 (Chhikara et al. 2025,
arXiv:2504.19413); Durable Functions semantics (Burckhardt et al.
2021, OOPSLA/PACMPL, DOI 10.1145/3485510); Tango shared-log data
structures (Balakrishnan et al. 2013, SOSP); LogAct (Balakrishnan et
al. 2026, arXiv:2604.07988 — agents as deconstructed state machines
over a shared log with pre-execution gating); SagaLLM (Chang & Geng
2025, PVLDB, arXiv:2503.11951). Failed-verify (excluded; cite by hand
if wanted): Kreps "The Log" (404), Helland "Immutability Changes
Everything" (403/unparseable).

**Verified roster (unreliable cognition):** MAST failure taxonomy
(Cemri et al. 2025, arXiv:2503.13657); correlated LLM errors (Kim,
Garg, Peng & Garg 2025, ICML, arXiv:2506.07962); LLM conformity (Weng,
Chen & Wang 2025, ICLR oral, arXiv:2501.13381); faulty-agent
resilience (Huang et al. 2024, arXiv:2408.00989); BFT for LLM MAS
(Zheng et al. 2025, AAAI, arXiv:2511.10400); Prompt Infection (Lee &
Tiwari 2024, arXiv:2410.07283); CaMeL (above); GuardAgent (Xiang et
al. 2025, ICML, arXiv:2406.09187); Multi-Agent Risks (Hammond et al.
2025, arXiv:2502.14143).

**Probe verdict: not found.** The literature splits into a corruption
side that measures attack success (AgentPoison; MINJA, Dong et al.
2025, arXiv:2503.03704; PoisonedRAG, Zou et al. 2025, USENIX Security,
arXiv:2402.07867; 2026 defences measuring filtering/screening — incl.
an explicit negative result, 0/360 poisoned memories rejected,
arXiv:2608.21230) and an enforcement side that measures
attack-success reduction under policy checks (Progent, CaMeL,
GuardAgent, AgentDojo). **No verified work corrupts beliefs upstream
and then measures containment as a flood of resulting invalid acts
refused against ground truth at an institutional boundary while the
agents keep believing the lie** — the E5 cell is unoccupied on these
searches.

**Query log (Group 4, verbatim, 2026-08-26):** 20 queries as run,
including: `MemGPT LLMs as operating systems virtual context arXiv` ·
`survey memory mechanism LLM-based agents arXiv 2024 2025` · `durable
execution LLM agents workflow crash recovery arXiv` · `SagaLLM
transaction guarantees multi-agent LLM planning arXiv` · `"Why Do
Multi-Agent LLM Systems Fail" taxonomy arXiv` · `Byzantine fault
tolerance LLM multi-agent systems resilience malicious agents arXiv` ·
`error propagation cascading failures LLM multi-agent systems arXiv` ·
`conformity LLM agents multi-agent debate sycophancy arXiv 2025` ·
`AgentPoison memory poisoning LLM agents knowledge base attack arXiv`
· `CaMeL defeating prompt injections by design capability security
policy enforcement arXiv` · `Progent privilege control LLM agents
enforce policy block tool calls arXiv` · `correlated errors across
large language models monoculture arXiv` · `stateless LLM agent
externalized state checkpointing survive restarts arXiv` · `Burckhardt
durable functions semantics stateful serverless OOPSLA` · `distributed
systems principles applied to LLM agents "distributed systems"
perspective position paper arXiv 2025` · `AgentDojo dynamic
environment attacks defenses LLM agents NeurIPS arXiv` · `MINJA memory
injection attack LLM agents arXiv 2025` · `"Multi-Agent Risks from
Advanced AI" correlated failures cascades Cooperative AI arXiv` ·
`memory poisoning defense agent "invalid actions" refused enforcement
ground truth evaluation` · `GuardAgent safeguard LLM agents guardrail
agent arXiv`.

### Group 5 — adversarial falsification response (2026-08-26, same day)

An external reviewer (GPT) ran an independent falsification pass
(reported as 102 queries; that pass's own query log was not exported
and is not part of this record — no claim in the paper relies on it;
only the candidates it produced, each independently fetch-verified
below, are relied on) and
proposed six papers as threats to the absence claims. All six were
fetch-verified by a fresh verification pass (6 queries, 13 direct
fetches, logged below): **all exist; the reviewer's descriptions were
accurate in substance.** Exact first-public dates: Doran JASSS 1(1):3
published 3 Jan 1998; Praetor (arXiv:2604.26274) v1 29 Apr 2026; ECA
"Hallucination as Exploit" (arXiv:2605.19192) v1 18 May 2026; TIDE
"When Truth Is Distributed" (arXiv:2608.03421) v1 4 Aug 2026; POLIS
"Multi-Agent AI Safety as an Institutional Design Problem"
(arXiv:2608.09828) v1 10 Aug 2026; AgentFlow (arXiv:2608.22868) v1
24 Aug 2026.

**Adjudication of the claimed falsification.** The reviewer ruled one
broad claim broken: "no prior work has shown corrupted/false model
beliefs being contained by an external enforcement boundary." That
sentence does not appear in either paper; the published claims were
already the narrow conjunction. Dimension-grid verification (deliberate
corruption / retention / ongoing cost / refusal-level accounting vs
authoritative world / trust-condition comparison / persistent-world
invariants) across all six papers found: **(c) ongoing behavioural
cost of a retained falsehood and (e) causal trust-condition comparison
are missed by all six**; the corruption papers (TIDE, Doran) have no
containment mechanism and the containment papers (ECA, POLIS,
AgentFlow, Praetor) corrupt no beliefs (ECA gates the model's own
per-call hallucinations; nothing is retained); POLIS varies
institutional treatments only, with a per-episode-resetting world and
cognition never manipulated; and **no paper in the set performs
kill-reset or frozen-model substitution as an experimental treatment
at all**. Both narrow absence claims survive; all six papers were
folded into Paper 2 §8 as prominently-cited near-misses under the
citation-relation discipline (chronology-precise verbs; Praetor and
ECA predate our frozen E-designs and are classified prior-in-time
discovered retrospectively; TIDE/POLIS/AgentFlow are contemporaneous).

**Discrepancies caught by verification:** search snippets mis-state
POLIS's ID as 2608.09857 (an unrelated robotics paper; on-page
identifier is 2608.09828); "AgentFlow" is a collided name on arXiv
(2607.01640, 2605.27466 are different papers); 2604.26274's abstract
shows an unexpanded `\codename` macro (system name Praetor per full
text); Doran's "behavioural consequences measured" is one
collective-level figure and it is a *benefit* of misbelief, not a
cost accounting.

**Query log (Group 5, verbatim, 2026-08-26):** `arXiv 2605.19192
"Hallucination as Exploit" evidence-carrying multimodal agents` ·
`POLIS "Multi-Agent AI Safety as an Institutional Design Problem"
arXiv` · `"When Truth Is Distributed" misinformation LLM multi-agent
arXiv 2608.03421` · `"AgentFlow" flow-centric policy language securing
LLM agent systems arXiv` · `"behavioral firewall" "benign
trajectories" structured-workflow AI agents arXiv 2604.26274` · `Doran
"Simulating Collective Misbelief" JASSS 1998 pseudo-agents`. Direct
fetches (13): the six abstract pages (plus 2608.09857 as a
discrepancy check), the five arXiv HTML full texts, and
jasss.org/1/1/3.html (metadata, characterisation, datelines).

## Summary for the papers

Four probes, four "not found" verdicts, each scoped to the queries
above on 2026-08-26 — plus a fifth, adversarial pass (Group 5) that
verified a reviewer's six falsification candidates and confirmed both
narrow absence claims survive, while adding the six as must-cite
near-misses. The bracketing structure the sweep establishes:
Gode & Sunder (1993) ran cognition-substitution under a fixed
institution three decades early with one aggregate outcome; Chupilkin
(2026) runs the exact transpose (agents fixed, institution varied);
the electronic-institutions tradition asserted cognition-independent
enforcement architecturally without ever ablating cognition; the
modern enforcement stack proves per-query security by construction
without persistent institutional invariants; and the corruption
literature measures attack success rather than institutional
containment. The pair's claimed cell — pre-declared institutional
invariants of an authoritative persistent world, scored under
systematic ablation, reset, frozen substitution and corruption of the
cognition layer — is unoccupied within these searches.

---

## Pre-submission recheck (2026-08-27)

**Status: COMPLETE — no replication, no supersession; three
neighbours and two brackets added to §9.** Conducted immediately
before arXiv submission: an independent reviewing model checked the live
arXiv cs.MA and cs.AI listings and ran targeted searches around the
paper's novelty bundle; every candidate it surfaced was then
verified at its primary arXiv record (title, authors, v1 date,
abstract) before any citation was added. The recheck's query strings
were not contemporaneously enumerated, so — per the accounting
discipline of the record above — no query count is claimed for it;
the 85-string enumeration above is the sealed record of the original
sweep and is unchanged.

Findings, with v1 dates:

- **When Do Institutions Beat Intelligence?** (Han,
  arXiv:2608.11357, v1 2026-08-11 — contemporaneous, missed by the
  original sweep): names the capability-vs-institutional
  intervention distinction and crosses it over four
  information-mechanism failure modes; the outcome is collective
  task performance in separate synthetic ecologies. Anticipates part
  of the framing; no persistent authoritative world, no standing
  invariants, no cognition destruction. Cited as [76].
- **When Agents Evolve, Institutions Follow** (Fei, Guo & Xiao,
  arXiv:2604.27691, v1 2026-04-30 — earlier; found only in this
  recheck): seven historical political institutions (spanning four
  canonical governance patterns) as executable multi-agent
  architectures, compared across three models and two task suites. The nearest model-by-institution comparison; it
  measures performance, not invariants under a persistent
  enforcement boundary. Cited as [77].
- **Beyond Memory: A Transactional Continuity Kernel for Long-Lived
  AI Agents** (He & Yu, arXiv:2608.11632, v1 2026-08-12 —
  contemporaneous): authorised state lineage with typed terminal
  dispositions and at-most-once effect identity, verified by bounded
  model checking (≈2.8M reachable states, zero invariant
  violations). The formal substrate counterpart: specification and
  mechanical verification of such a boundary, with no live cognition
  experimentally manipulated above it. Cited as [78].
- **AID-Guard: Stateful Authorization for Delegated Agent Effects**
  (Tong, Dai & Guo, arXiv:2608.21159, v1 2026-08-21 —
  contemporaneous): at-most-one provider effect across retry,
  ambiguity, recovery and successor execution under proposer
  compromise; episodic authorisation, no persistent society, duty
  reconstruction or cognition substitution. Cited as [79].
- **When "Must" Becomes "Maybe"** (Sun et al., arXiv:2608.24569, v1
  2026-08-25 — contemporaneous, one day before the original sweep):
  constraint text survives workflow handoffs while its
  action-binding force is lost; downstream verification still blocks
  the forbidden actions. A supporting bracket for the distinction
  between semantic availability and institutional enforcement.
  Cited as [80].

Adjudication against the novelty probes, re-run at claimed width: no
candidate combines a persistent authoritative world with an
adjudicated accepted history; none runs the ablation → kill/reset →
frozen-substitution → testimony-corruption sequence; none scores
singular accepted reality, typed refusal, duty recovery, duplicate
prevention and false completions together; none measures the
continuing behavioural cost of an admitted falsehood while an
institution contains its effects. Consequential wording changes in
§9: the POLIS superlative narrowed to "one of the closest
contemporaneous institutional experiments our sweep found", and the
occupied cell restated precisely (a persistent authoritative
institution, pre-declared standing invariants, and systematic
ablation, interruption, substitution and corruption of the cognition
acting above its enforcement boundary).
