Working notes · v2 · August 2026

A systems approach to building with AI

Every layer here exists to fix a failure that had already happened, which is the only thing that earns a layer its place. Most of what's left to fix sits at the seams between layers, and the rest is ordinary maintenance debt inside a layer that hasn't been paid down yet.

Shashank Mehrotra 15 August 2026 About 10 minutes

One

Goals

Using AI more was never the goal. The point was to fix one repeated failure: a good session depended on the agent already knowing my conventions, already knowing what I was working on, and being able to pick up work another machine had started an hour earlier. No prompt fixes that. Plumbing does.

The actual goals were narrower and more mechanical than they sound. One canonical set of instructions, no matter which tool or which machine reads them. Memory and task state that survive a session ending. Work that moves between devices without me carrying it by hand. Real output at the end, not just conversation.

This is built for one person working at high velocity across many project areas, on two laptops and a phone. Not a team, not a product. That explains choices which would look like gaps in a team context, no permissions system and no multi-user anything, because those problems don't exist here.


Two

Core requirements

Six constraints shaped every choice below, and they're independent of each other rather than a sequence. Each one is a test that later architecture decisions have to pass.

  1. Neither laptop is primary. Anything that must survive a machine switch has to live in a git remote, never only on disk.
  2. No credential is ever written to a tracked file. A password manager plus a launcher script, because every project having its own remote only stays safe if nothing in any repo is a secret.
  3. Nothing may block a session. A continuity system that can wedge you mid-work is worse than no continuity system at all.
  4. State has to be hand-editable without an agent in the loop. A plain file store, not a database.
  5. Instructions and state change on different clocks and can't live in the same place, or rewriting one risks disturbing the other.
  6. The record has to be honest about what's verified and what's judgement. A hook writes only what it can mechanically check, and a human-triggered command writes the narrative.

Requirement six is the one that matters most and is easiest to skip. It's tempting to let an automated hook infer intent, because intent is the useful part. But a hook that guesses is a hook that lies confidently, and a confident lie is worse than a gap you notice.


Three

How it was actually built

Version one was one instruction file, symlinked to two harnesses, Claude Code and Codex. Everything below is additive to that single move.

The continuity problem surfaced next. Two laptops meant re-explaining context by hand, or worse, re-running a dead end the other machine had already ruled out. A pair of hooks fixed it: one that writes a handoff when a session ends, one that surfaces it on the other machine, plus a command for the narrative half no hook can infer.

I split state from instructions once the two proved to change at different rates, for different reasons, and memory became its own store, versioned separately from the instruction file. A separate attempt to formalise "what to work on" as its own tracked system didn't earn its keep, so I shelved it. I name it here as a real dead end rather than pretend every layer has stuck.

The system had no name, and that cost real time. I was using three informal names for the same thing, so nothing could be referred to as a whole. Naming it fixed a coordination problem rather than a technical one: a map file now answers exactly one question, where does the rule I need live.

A third harness arrived almost for free. Google's Antigravity, one more symlink onto the same instruction file. That's the clearest evidence the earlier layering decisions were right: adding a fourth or fifth harness should cost about the same.

Work started needing more than one agent per session. A named bench of specialists arrived, each with its own model and remit, alongside a delivery playbook distilled from a real project's post-mortem and an orchestration layer for fan-out and adversarial-verify patterns. The deploy target grew the same way, from "push and it's live" into a small registry, once Cloudflare's real constraints bit: a project can't be renamed once created, only recreated, which destroys its secrets; attaching a hostname is a dashboard-only action with no command-line route; and there is exactly one legal place in the whole zone for a redirect to live. A convention became a registry the moment a mistake there got expensive.

Model tiering went from an unstated habit to a named rule once the bench got big enough that "which model for this" needed an actual answer instead of a guess.


Four

The architecture, as layers

What the system looks like today, laid out bottom to top. This is one diagram rather than three, because what crosses the layers matters more than the layers on their own. Three real flows are traced through the stack: resuming work on the other laptop, shipping an essay, and adding a fourth harness. The third is the cheap case: it should touch almost nothing. It sits against the one seam that isn't cheap yet, explained in the next section.

owned and changed here depended on, not owned resume on the other laptop ship an essay add a fourth harness
%%{init: {"theme":"base","flowchart":{"htmlLabels":false,"useMaxWidth":true,"curve":"basis","nodeSpacing":28,"rankSpacing":52},"themeVariables":{"background":"#ffffff","primaryColor":"#fbf4e2","primaryTextColor":"#1a1917","primaryBorderColor":"#8a6f2e","lineColor":"#8f8b83","secondaryColor":"#eaf1f8","tertiaryColor":"#f8f7f5","clusterBkg":"#f8f7f4","clusterBorder":"#ddd9d0","fontSize":"12.5px","edgeLabelBackground":"#ffffff"}} }%%
flowchart TB
  subgraph L0["layer 0 · foundation"]
    direction LR
    PER["Shashank"]
    MB1["MacBook one"]
    MB2["MacBook two"]
    PH["Phone"]
  end

  subgraph L1["layer 1 · harnesses"]
    direction LR
    CC["Claude Code"]
    CX["Codex
only runs inside a trusted git repo"] AG["Antigravity
alone reads project-scoped rules"] end L2["layer 2 · canonical instructions
CONTEXT.md, symlinked three ways"] L3["layer 3 · state
the memory store, typed and indexed"] subgraph L4["layer 4 · execution"] direction LR BEN["agent bench
8 specialists"] PLB["delivery playbook"] WF["workflow orchestration"] end L5["layer 5 · transport
git and GitHub"] subgraph L6["layer 6 · external services"] direction LR API["model APIs"] MCP["MCP servers"] KC["credential store"] CF["Cloudflare Pages"] end %% invisible backbone: forces the vertical rank to match the layer numbers. %% every real edge below is also written to point downward, 0 toward 6, so %% the backbone and the real structure agree instead of fighting each other PER ~~~ CC CC ~~~ L2 L2 ~~~ L3 L3 ~~~ BEN BEN ~~~ L5 L5 ~~~ API linkStyle 0,1,2,3,4,5 stroke-width:0,opacity:0 PER --> MB1 & MB2 & PH MB1 --> CC & CX & AG MB2 --> CC & CX & AG PH --> CC CC --> L2 CX --> L2 AG --> L2 L2 -.- L3 L2 --> L4 L3 --> L4 L1 --> L6 L4 --> L6 L3 --> L5 L4 --> L5 L5 --> L6 MB1 -.->|"1 handoff"| L3 MB2 -.->|"1 resume"| CC linkStyle 27,28 stroke:#e8c15a,stroke-width:2.5px linkStyle 20,21,25,26 stroke:#b06a4f,stroke-width:2.5px linkStyle 18 stroke:#5b7d3e,stroke-width:2.5px classDef core fill:#fbf4e2,stroke:#8a6f2e,stroke-width:1.5px,color:#4a3c12 classDef ext fill:#eaf1f8,stroke:#3f6f9f,color:#1e3f61 classDef person fill:#f7efd8,stroke:#8a6f2e,stroke-width:1.5px,color:#4a3c12 class PER,MB1,MB2,PH person class CC,CX,AG,L2,L3,BEN,PLB,WF,L5 core class API,MCP,KC,CF ext
Fig. 1 Gold traces resuming work on the other laptop: a handoff commits state through git, the other laptop resumes straight from it, instructions never move because both machines already have them. Rust traces shipping an essay: execution reads instructions and state, pushes through transport, lands on Cloudflare. Green traces adding a fourth harness: one edge, into layer 2, nothing else moves. Layer 6, at the bottom, is reached directly by more than one layer above it, not only the one drawn next to it. This is a dependency map grouped by concern, not a strict enforced layering where calls only cross adjacent levels.

Five

What's working, what isn't

Held against the six requirements from section two, three things clearly work, three are gaps, and two are open questions rather than failures.

Working

A fourth harness costs one symlink

Configured, then checked. I made both Codex and Antigravity read back a fact from the instruction file cold, unprompted, which proved the wiring works rather than merely exists.

Working

The continuity loop holds its boundary

Verified handoffs record facts. Hand-written narrative records intent. The two haven't been confused with each other yet, which is the exact failure mode requirement six exists to prevent.

Working

No secret has ever needed a tracked file

Across every project having its own remote, the credential rule has held without exception.

Partly working

Memory now flags itself, doesn't prune itself

A weekly script checks every memory store for a fact a script can verify: a broken link, an orphaned file, a path that no longer exists, an entry untouched for months. It writes a report and stops there. Nothing decides yet whether a flagged memory is actually wrong, that's still a human read. This sits entirely inside layer three, not at a seam.

Not working

A handoff is per project, not per thread

Two unrelated pieces of work in one repo flatten into a single note. This is inside the continuity mechanism, not a seam either.

Not working, the real seam

No behavioural test across harnesses

Three harnesses now read one file, and nothing yet proves they agree once they've read it. Identical bytes in is not identical behaviour out, and this is exactly the layer one to layer two seam the opening claim is about.

Open, not broken

Antigravity's project rules

It reads project-scoped rules the other two harnesses structurally can't. That asymmetry hasn't been decided on, only noticed, at the same layer one to layer two seam.

Open, not broken

The phone is a steering wheel

Good for redirecting a running session, poor for starting one from scratch. Still unclear whether that's a tooling limit or the correct division of labour.


Six

What's next

Ordered by the cost of leaving each one unfixed, cheapest first. This is a short list, not a wishlist.

  1. A verdict on the memory flags: read a few weeks of the review script's findings, then decide whether staleness ever gets judged automatically, or stays a human read by design.
  2. A way to split a handoff by thread of work, not just by project.
  3. A minimal behavioural check across harnesses, so "identical bytes" stops being taken on faith.
  4. A real decision on the Antigravity asymmetry: either give the other harnesses an equivalent, or state plainly why one harness having more local context is fine.
  5. A decision on the phone, not just a question: run one real task phone-only for a week and see what actually breaks.

Seven

What the wider field suggests

Not a survey. Each comparison below attaches to one specific gap from the last two sections rather than standing as a generic parallel.

Spotify's own account of building an opinionated, documented "golden path" through its engineering organisation1 makes one argument worth stealing directly: a golden path only works as maintained, docs-as-code, with a named owner. A diagram drawn once and left alone stops being true within a quarter. The map file needs the same treatment: an owner and an update cadence rather than a one-time artefact.

Section three already reads like an informal log of decisions: symlink, hooks, memory split off, a dead end shelved, a name chosen. Michael Nygard's original essay on architecture decision records2 argues for making that log explicit: one short, numbered file per decision, with its own status and consequences, rather than folded into essay prose that only gets rewritten when someone happens to revisit it.

The memory-pruning gap almost got its answer from Anthropic rather than from building one. Social posts and a fan-maintained compilation describe a scheduled review pass for Claude Code, nicknamed /dream, that runs a subagent over past sessions and memory and drops what no longer holds.3 Checked against Anthropic's own changelog, current as of this piece, no such command has shipped. Citing it anyway would have broken requirement six from section two before this section even finished making its point.

So the actual answer is a script, not a citation. A weekly pass now walks every memory store and reports three plain facts: a link to a file that's gone, a file a memory names that no longer exists, an entry untouched a long time. It writes a report and stops, exactly the shape requirement six asks for, a hook records what it can verify, a human decides what a flag means. First run found five real findings on the first pass, including a memory in another project still pointing at a path this piece's own history moved: the exact failure mode this section opened with.

Two smaller parallels, worth a sentence each rather than a section. Dotfile managers like chezmoi and GNU Stow solved the literal layer one to layer two problem years before this piece needed it: one canonical file, applied consistently across machines. And the standard argument for hexagonal architecture, that adding a new adapter against a fixed contract should cost about the same as the last one, is exactly the claim the "add a fourth harness" flow in section four is making.

None of this changes the shape of the system. It sharpens four specific decisions already sitting in section six.


Notes
  1. Gary Niemen, "How We Use Golden Paths to Solve Fragmentation in Our Software Ecosystem", Spotify Engineering, 17 August 2020.back
  2. Michael Nygard, "Documenting Architecture Decisions", Cognitect, 15 November 2011. The essay that defined the now-standard ADR format.back
  3. /dream is described at a fan-maintained, source-linked compilation, howborisusesclaudecode.com, which traces it to Anthropic's own public posts. Anthropic's official changelog does not list it as a shipped command as of this piece's publication, so it's named here as an unconfirmed report, not a citation. A related, more specific claim about deleting instruction files on a six-month cycle could not be independently verified anywhere and is not repeated at all.back
  4. Drafted with Claude, structured in D4, reviewed by a six-role editorial panel and a separate research pass before drafting. A first version of this piece ran in August 2026 as a strict C4 (context, container, component) essay; this version replaces that structure after the underlying system grew past what three C4 levels could hold cleanly.