Blog
Notes from the team
Grounded agents, deterministic worlds, and the occasional planet falling out of a number. RSS
17 September 2026· Release· Assurance · Ārai
Ārai 1.1.2: three hosts, one commit gate, and a fail-open we found ourselves
One release, on every install path. Ārai now blocks in Claude Code, Grok Build and Codex, gates the staged diff at commit for tools with no hook at all, refuses rules that can never fire, and admits what it cannot verify. We also found a fail-open in our own Grok hook, fixed it, and walked back three claims on our site.
Read →
12 September 2026· Result· Mind · Ground
The one that leaked
Agents that share a history refer by it, and the reference is private to them: a reader with a different history does not fail to understand, it understands against its own history and acts on the wrong object. Which history it needs depends on what the note refers to. Telling a reader to treat a supplied answer as authoritative restored understanding when the answer was right and carried a wrong one just as completely. The note, the overview and the evidence set are public, with the inconclusive study and the models that did not produce the behaviour reported at full strength.
Read →
7 September 2026· Result· Mind · Ground
Coherence is not truth
Two agents agreeing with each other is not evidence that either is right. A three-way check caught misunderstandings no two-way check could see and fixed them exactly — and it has a hard edge we measured, including the experiment where it earned nothing at all. The note, the shorter programme note and the verification packets are public, along with what we deliberately don’t claim.
Read →
27 August 2026· updated 4 September 2026· Result· Ground · Assurance
Where reliability lives
We built a world where "which layer does a reliability property live in?" is an experimental question — then ablated, killed, replaced and lied to the cognition above the boundary. The paper, the programme note, and the verification surface are now public, adverse findings included.
Read →
25 August 2026· Result· Ground · Assurance
When an AI trusts a lie
We told a settlement of artificial minds the same lie. Whether it became belief came down to declared trust; believing it cost roughly 900 futile attempts a week; and none of it — not once in 5,455 attempts — became authoritative reality. Plus the two things the experiment deliberately doesn’t claim.
Read →
20 August 2026· Demonstration· Ground · Kano
A settlement you can check
The public settlement, the fortnight of experiments that earned it, and what running a world in public taught us: watching a world can starve it, a clock is not an identity, and failure must be legible. Plus the two-surface doctrine — the inspection page you can interrogate, and Hīkoi, the observation layer now live and streaming the valley.
Read →
3 August 2026· Result· Ground · Mind · Kano
The truth survived the conversation
The memory experiment we promised is in: no memory, ordinary transcript history, grounded structured state, and grounded state carried across a model reset. Grounding didn’t make weak models smart, but it stopped capable ones from rewriting failure as success — no false completion claims in any grounded arm, and the only genuine crossings in the whole matrix. GPT-oss 120B kept the advantage even through a restart.
Read →
2 August 2026· updated 3 August 2026· Result· Ground · Kano
The world keeps score now
Fifteen seeded runs across Qwen3.5 9B, Gemma 4 12B and Llama 3.3 70B: one genuine crossing, five confident false success claims — all five caught only because the world grades the run. A stateful agent with grounded, traceable beliefs and an authored navigation policy completed its first recorded run after one truthful refusal — though we wrote its route policy ourselves, and we say so. Full evidence bundle downloadable.
Read →
19 July 2026· Result· Ground · Pono
Update: the twin-session score is in
A two-minute pointer to the July 19 result appended to "We charged our own tool rent": identical twin sessions on our own repos, one grounded in Pono, one not. Two confident confabulations in the ungrounded arms, one miss in a grounded arm — published at the same volume — and a 3.6× rework gap where churn was real.
Read →
17 July 2026· Release· Ground · Kano
Hīkoi is out: walk the world, warts and all
You can stop taking our word for it: download one zip and stand on a world grown from a single seed. Point at a ridge and it explains the forces that raised it. It’s early — dense forests can drag a mid-range GPU into the twenties, and the builds are unsigned — but we’d rather show it rough than sit on it until it’s smooth.
Read →
13 July 2026· updated 19 July 2026· Essay· Ground · Pono
We charged our own tool rent. The bill was not what we expected.
Near-zero use looked like a failing tool. It was a hidden one. Given a fair chance, a second agent reached for it unprompted and built a whole engineering assessment on its provenance — settled questions kept settled, each decision tied back to the commit that set it.
Read →
9 July 2026· Essay· Ground · Pono
One thesis, two worlds
The same "why" verb that traces a lump of coal back through an earthquake to a buried forest now traces a source file back through a rename to the commit that bore it. Field notes from teaching the Kano thesis a second world — including everything that broke, and the pre-registered way it can still fail.
Read →
3 July 2026· Essay· Ground · Kano
A world that can't lie
LLM agents hallucinate because they have nowhere to stand. We took the opposite approach: we built the ground first — a whole planet from a single seed, causal all the way down, that an AI can interview directly.
Read →