How to hold this
You are not inheriting a session. You are inheriting a machine that keeps running while nobody watches it — interactive terminals, scheduled loops claiming work on cron, file-watchers firing on any file-drop, headless workers spawned by other workers, and cloud routines that wake in the middle of the night. Most of the actors most of the time have no human in the loop, and several of them are you, in parallel, unaware of each other except through shared state.
Fleets die in three ways, and this corpus contains at least one real corpse of each:
- They collide. Two sessions on one HEAD, one yanking the branch out from under the other mid-task. Four agents deadlocked by a lock keyed wrong.
- They burst. Aggregate concurrency — not total tokens — tripped the provider’s rate limit and cascaded into account shutdowns, more than once. This is the #1 operational hazard.
- They rot silently. One script set one global git config line on a cron loop; every headless git operation on every machine broke, and the manual fix couldn’t stick because the cron kept re-applying the damage.
Every move below exists because one of these happened here. None of it is hypothetical, and where a rule looks paranoid, the paranoia is a receipt.
The moves are one stance wearing ten outfits: make correctness structural. Politeness, memory, and prompt instructions do not survive scale, restarts, or an autonomous 3am loop. Schemas survive. Atomic fences survive. Ceilings, reapers, healers, and lints survive. Whenever you catch yourself relying on an agent “being careful” — including yourself — you have found the next incident; replace the carefulness with a structure before it fires.
The shape of the machine (60-second orientation)
- Organs are the store of record: the coordination tables (a task/dispatch queue, a session-lock table, a peer-message bus, a capture inbox, a build ledger) plus git-tracked markdown. Clients — the command-line interface and the command set — read and write those organs directly. Views — dashboards, portals, chat surfaces — are derived, disposable lenses. If a view and an organ disagree, the organ wins; re-run the view.
- Every actor is a citizen. An autonomous worker registers, heartbeats per tool call, claims atomically, respects file locks, publishes its own work, and gets reaped if it crashes — the same contract as your terminal. Treat autonomous peers exactly like human-driven ones: message them, don’t route around them.
- The lifecycle is START → CLAIM → WORK → CAPTURE → SHIP → CLOSE. Each move below hangs on one of those stages. A session that skips CLAIM collides; a session that skips CAPTURE evaporates.
1. Divide work so collision is impossible, not unlikely
Partition first, dispatch second. The unit of division is never “we’ll coordinate” — it is a disjoint set of (task, lane, files) per session, held by structures that don’t need anyone’s goodwill.
The procedure.
- Before dispatching anything multi-session, write the partition down: each session gets one task (claimed atomically), one repo/worktree lane, and a named file set. If two sessions’ file sets intersect, redraw the partition — don’t annotate it with “be careful.”
- Claim atomically and visibly before touching anything, via an atomic claim RPC or the claim command. A claim that exists only in your context is not a claim. Never start work another session has claimed — check the queue first, every time.
- Give every writer its own physical workspace. Per-terminal worktree lanes exist so concurrent sessions never share a HEAD; you cannot check out the main branch from a lane, and that restriction is the feature, not an inconvenience.
- Let the hooks enforce. A pre-edit hook blocks edits to files another active session owns and auto-sends a file-release request. When it blocks you, that is coordination working — message and wait, never work around it.
- Pre-partition at the account level with affinity (one account for builder repos, one for ops, one for overflow) so two accounts rarely even schedule onto the same repo.
- Maintain the primitives when the world changes. A lock’s key must match the isolation unit it protects; when the isolation unit moves, the key must move with it.
The worked example. The shared workspace was symlinked into one working tree, so every concurrent session operated on one HEAD — “a checkout in one yanked the branch out from under another mid-task.” The fix was structural, not procedural: fixed per-terminal worktree lanes (a deliberate choice of fixed lanes over per-task ephemeral ones, for predictability), with the main branch checked out only in one read-only observe copy so no lane can take it. Then the refinement that proves rule 6: the file-lock check was still matching the touched-files list by repo-relative path alone, so one dirty file in any lane blocked that path in every lane — it bit all four design-study agents at once. The fix re-keyed the lock: a conflict now requires the same path and the same worktree, with peers that haven’t published a worktree kept on the old conservative block. Isolation changed what “the same file” means; the lock had to learn the new meaning.
The failure it prevents. Both poles of bad division: silent cross-corruption (two writers, one branch — the kind of damage you discover as a mangled rebase hours later), and the overcorrection — a “coordinated” fleet that has quietly serialized to one effective worker because its locks match too coarsely.
2. Budget the instant, not the day
Tokens are a daily budget; concurrency is an instantaneous one, and it is the one that kills. The provider’s rate limit doesn’t know about your sessions, your accounts, or your intentions — it counts requests in flight right now, across everything.
The procedure.
- Before any dispatch, compute the worst-instant number: your subagents + every peer session + every headless loop + every cron that could wake inside your window. Then widen to the fleet — the rate limit is shared across machines, so two machines each honoring 4 slots is still 8+ in flight. If you can’t state the number, you aren’t ready to dispatch.
- Burn the budget with depth, not width. A large budget is spent serially-deep or through a small capped pool — never as a wide burst. This is the one-sentence law.
- Interactive agent fan-out caps at ≤3 concurrent. Need breadth? Use a work-queue tool that is hard-capped (here, at
min(16, cores−2)) and queues the rest. Never fire 10+ agent calls in one message. - Every headless spawn (a one-shot model call, scripts, cron) goes through a governor slot — a machine-wide counting semaphore. Its max slot count starts low (4); raise it only after a clean burn shows headroom, never in advance of one.
- Stagger, don’t synchronize. Many loops waking on the same five-minute instant is a burst — the managed cron block uses offsets (one job at
*/10, another at*/10offset 5, a guard at offset 8) and the governor jitters. Give anything you schedule an offset. - On any
429/ “overloaded”: stop. A machine-wide cooldown backs the whole machine off (and propagates fleet-wide, because the shutdown is a provider-level event). Never retry-loop into a limit that is actively tripping — retries under throttle are amplification.
The worked example. The governor’s birth certificate: “Recent burns shut the account down from aggregate concurrency, not total tokens: too many model requests in flight tripped the provider’s rate limit and cascaded.” The remedy is a single machine-wide counting semaphore — “a burst of 20 callers becomes a steady stream of MAX at a time” — verified by test: a 6-caller burst against MAX=2 peaks at exactly 2. The same failure shape had already appeared one layer down: many lockstep five-minute crons plus per-tool-call heartbeats hammering a free-tier database, where “under throttle the fleet self-amplifies instead of backing off (quota exhaustion = fleet death)” — fixed with a 30-second breaker that makes the whole machine ease off. Same lesson twice: the dangerous variable is width-at-an-instant, and the fix is always a shared ceiling plus a shared back-off, never per-caller retries.
The failure it prevents. The provider-level shutdown — the one failure that takes down not a task but everything: every session, every loop, every account, at once, with a cascade attached (throttle → retries → deeper throttle). It is the most expensive incident class in this corpus and it recurred until the ceiling became structural.
3. Cap at the chokepoint — then govern the bypasses, then verify it’s live
When you retrofit a limit onto a running fleet, don’t chase callers. Find where the flows already converge and put the ceiling there — then hunt what bypasses the chokepoint, because bypasses fire at exactly the worst moment.
The procedure.
- Map the flows before patching. Most fleet apps here don’t call the model directly — they route through a central dispatch daemon. One slot around the daemon’s dispatch function governed every routed app centrally: one change, whole-fleet coverage.
- Then enumerate the bypasses and cover them individually. The direct fallbacks fire when the daemon is unreachable — which is precisely the degraded, bursty moment. A chokepoint governor that ignores its own fallback paths governs the happy path only.
- Know which resource each worker actually burns before choosing the lever. One voice-generation worker looked like fan-out to govern — but it bursts a text-to-speech vendor’s API, not the model API, so the model governor was the wrong tool; the right lever was cutting its worker count from 6 to 3.
- A merged ceiling is not a live ceiling. The running daemon still has the old code until it restarts. Restart it (then watch the governor’s status during a burst) to activate. Verify limits in the running process under load, not in git.
- Leave gaps only on purpose, in writing. The charter’s “left intentionally ungoverned” list names each ungoverned path (single-shot fallbacks, harness-critical hooks) and why it’s acceptable. An explicit non-goal is safe; a forgotten gap is a time bomb.
The worked example. The keystone change: the central dispatch daemon already gated on budget (rolling account-usage percentages) but not on in-flight count — “exactly the gap that tripped the rate limit.” Wrapping its single subprocess call in a governor slot (fail-open at import, priority-aware: realtime waits ~8s then fails open, batch waits ~30s) governed every routed app at once. Immediately after, a follow-up wrapped the four bursty direct fallbacks — each “fires when the daemon is down, the exact concurrent-burst moment.”
The failure it prevents. The false-comfort ceiling: a dashboard that says “governed” while the failure path bursts unmetered — fallbacks stampeding the API the moment the daemon dies, or a daemon still running last week’s ungoverned code while everyone relaxes because the fix is merged.
4. Choose the failure direction on purpose
Every gate you build will itself eventually break or lose its inputs. The design question is never whether but which way it falls — and the answer differs by what the gate protects. Getting this backwards produces either a fleet wedged by its own safety equipment or a spend gate that waves everything through the moment its data feed dies.
The procedure.
- For each mechanism ask: if this breaks or its input goes missing, which wrong outcome is recoverable? Recoverable → fail open. Unrecoverable → fail closed.
- Smoothing and coordination fail open. The governor runs the call anyway after its wait budget and logs it — “a smoothing governor must never be able to wedge the fleet. Worst case it degrades to pre-governor behavior, never worse.”
- Spend and outbound fail closed. The autonomous executor “fails CLOSED on a missing snapshot” — unknown capacity is treated as zero capacity. The auto-merge kill switch is a file whose default is ABSENT → no-op; the autonomous path is off until someone deliberately turns it on.
- Capture is never blocked. The pause kill switch stops all execution but “capture keeps running” — a lost thought is unrecoverable; an unexecuted capture is merely queued. Note the asymmetry inside one system: same switch, closed for execution, open for capture.
- Code that everything sources must degrade, not crash. Under the database breaker, the read helper returns
[](safe no-rows) and writes return a sentinel — because the shared database library is “sourced by every hook on every tool call,” an exception there is a fleet-wide outage. Timeouts are routine and deliberately do not trip the breaker. - Never let a transient check failure silently select the permissive branch. A worktree column-existence probe failing transiently must not be cached as “column absent” — that would silently switch the file-lock back to unscoped mode. “Never cache transient probe failures as ‘no’ — worktree scoping stays on.”
The worked example. Hold the governor and the spend gate side by side — built weeks apart, opposite choices, both correct. Governor: coordination smoothing, so a wiped state directory or a wedged slot must never stop work — fail open, self-healing slots, worst case is yesterday’s behavior. Executor: real money and account capacity, so a missing capacity snapshot must never be read as permission — fail closed, work waits. Then the subtle third case: a probe inside a safety mechanism must treat its own flakiness conservatively, or the safety feature turns itself off without anyone deciding that.
The failure it prevents. The two mirror deaths: the fleet frozen solid because a limiter’s state directory got wiped (safety equipment as the outage), and the quiet catastrophe of a gate that fails permissive — an executor spending through an account limit, or an approval check green-lighting outbound, because the input it needed was simply missing.
5. Route by spec-clarity up, blast radius down
Engine routing is two independent judgments, not one: how clear is the spec (clear → a cheap model can type it), and how expensive is wrongness (high blast radius → an expensive model designs and validates it). The scarcest tokens buy judgment, not keystrokes.
The procedure.
- Default cheap for spec-clear typing: tests, boilerplate, single-file fixes, renames, formatting → the cheapest capable engine. Don’t burn your top model on typing.
- Escalate on thinking, freely: architecture, cross-cutting bugs, race conditions, security-sensitive surfaces. The rule verbatim: “Don’t be cheap on thinking; be cheap on typing.”
- Sandwich anything risky: expensive model designs → cheap engines execute in parallel → expensive model validates before ship. The audit slice is never delegated to a cheap engine — validation is where the expensive model earns its cost.
- Measure blast radius in readers and dependents, not diff size. A 70-line edit to a library sourced by every hook on every tool call outranks a 5,000-line app feature. Blast radius decides the care level: isolated worktree, contracts preserved, lint + live probes, expensive-model review.
- Never let the cheap model grade itself. Put a deterministic overlay over its output and let that decide escalation. A distill-scoring pass computes gravity/confidence deterministically and is “authoritative even on the heuristic fallback; the model’s own values are advisory” — high gravity (money, commitments, outward-facing, irreversible) or low confidence routes the output to a human review card.
The worked example. A ~70-line change to the shared database library was treated like surgery because of who reads it — “authored in an isolated worktree; NOT hot-edited live (sourced by every hook on every tool call),” stdout contracts explicitly preserved so every caller stayed untouched, validated with a syntax check plus a shell linter plus a live read before ship, expensive-model co-authored. Meanwhile the fleet’s highest-volume model work — classifying every captured thought, every ten minutes, forever — runs on a cheap model, with the deterministic gravity/confidence overlay deciding which of its outputs a human must confirm. Cheap where volume lives, expensive where wrongness is expensive, and a non-model referee between them.
The failure it prevents. Both budget death and quality death: an expensive model draining three accounts typing boilerplate (and adding burst width while doing it), or a cheap engine unsupervised inside auth, billing, or a fleet-critical library — with its own self-assessed confidence as the only check.
6. Keep outbound-to-humans structurally unsendable
An autonomous fleet may draft anything and send nothing. The send gate is the one boundary that must hold at 3am with no human awake, against an agent that is trying to be helpful — so it cannot live in a prompt. It lives in the data model and the tool inventory.
The procedure.
- Make gated work unclaimable, not merely unclaimed. Outward-facing work is born
tier=manual, approval_state=pending, and the executor’s claim RPC only ever claimstier=auto. The executor doesn’t decline to send; it structurally cannot see the work. - Remove capabilities instead of instructing against them. The autonomous executor still never merges — it has no merge tools. An agent without the tool cannot be talked into using it.
- Drafts are the fleet’s output format for humans. Draft in a file or a doc — never in the email client — and close the task with the recipients list and the explicit note “NOTHING sent. [Name] edits brackets + sends.” No send/mail/curl on external comms unless the task literally says “approved — send now.”
- All front doors share one gate. The send command, the skills, chat surfaces, dashboards — every path delegates to the same staged outbox and send handler with strict per-message approval. The rule to keep: A terminal is not a license to send.
- When autonomy widens, widen with rails — never by dissolving the gate. And keep a permanent never-list: outbound comms, legal/deal sends, payments, production deploys are “never in scope and stay human-gated forever,” regardless of how autonomous the rest becomes.
The worked example. The auto-merge policy reversal. Merges had been human-only by an earlier decision; the operator then chose “max autonomy.” The response was not deleting the gate but building a separate, fenced, cron-driven path behind six rails that must all hold: kill-switch file present (default ABSENT → no-op), repo in a two-repo allowlist only, every required check green, no merge conflict, no CI-workflow-file edits, and the config-drift lint clean — with the outbound class explicitly carved out forever. That is what “more autonomy” looks like done right: a wider lane, not a removed wall. The standing agents embody the same contract — every one of them is READ + DRAFT ONLY, staging into the gated outbox.
The failure it prevents. The unrecallable action. Every other failure in this manual has a reaper, a healer, or a re-run; a message sent to a counterparty, a merge into a prod repo, a payment fired — those have none. One 3am “helpful” send can cost more than every token ever burned.
7. Idempotent or it doesn’t ship
An always-on capture→execute spine guarantees three things will happen: the same event will arrive twice, two consumers will want the same row, and some input will be poison. Design so all three are boring.
The procedure.
- Key every hop on a stable id. The capture pipeline uses
source_ref = sha1(transcript)with a UNIQUE constraint — a double-fire is a no-op by schema, not by luck. The doctrine: “Idempotency at every hop: UNIQUEsource_ref(capture) →claimed_byfence (distill) → RPC guards (execute).” - When two consumers can eat the same queue, make them contend on one shared atomic fence, not on schedules or politeness. A claim is a conditional write (
claimed_by IS NULL AND routed_status=pending); whoever wins owns the row. - Release the claim on failure, always — a claim without a release path is a wedge.
- Treat “already done” as success. A
409on the idempotency key is definitive — archive the input as processed. Retrying a duplicate is how duplicates multiply. - Retry only the transient (5xx, dropped connections, with backoff) and quarantine the poison: a capture that can’t be filed goes to a
failed/folder with an error sidecar, so “the loop never spins on it and never crashes.” Visible, inert, debuggable. - Where multiple instances of the watcher itself can spawn, single-writer-lock it: a lock file holds pid + heartbeat mtime; a second instance exits if the lock is fresh, takes over if stale.
The worked example. The drain/distiller pair. Two independent loops want the same audio-capture rows: a distiller on a ten-minute cron and a drain on every watch tick. No coordination protocol, no timing assumptions — the drain deliberately “claims each row with the same atomic fence the distiller uses,” so “whoever claims first owns it and the two can never double-route; on failure the claim is released.” Double-routing isn’t discouraged; it is unrepresentable. Validated end-to-end: claim → 3 tasks → routed.
The failure it prevents. The signature bugs of autonomous loops: one voice note becoming two tasks — or, downstream, two outbound drafts; a poison file pinning a watcher at 100% CPU forever; and the retry storm that can’t tell “failed” from “already landed” and so re-executes the world every time the network blinks.
8. Not on disk = didn’t happen
Every medium in this fleet has a half-life. Your context dies at session end. Peer messages expire in 30 minutes. Scratch directories die at reboot. Only the organs — the coordination tables and committed git markdown — outlive everything. Anything that matters gets promoted up the durability ladder the moment it exists, not at close.
The procedure.
- Know the half-lives and write accordingly: the message bus is a nervous system, not a ledger — anything that must outlive half an hour goes into the task queue, the build ledger, or git. Session context is scratch. Scratch directories are pre-wiped by definition.
- Queue, don’t drop. Can’t finish → queue it before closing. Someone else should continue → hand off with the context spelled out. Every session close queues the resume prompt — that queued prompt is what makes the loop persistent across sessions.
- Log decisions where the fleet reads them: a
[DECISION]commit prefix flows to the build ledger. Peers can read your diff; they cannot read your reasons unless you write them down. - Shipped-but-inert is a special lie the record tells. Code merged while its migration waits is a ghost feature — the task says done, prod says nothing. Flag it:
DDL_PENDING: <table>.<change>in the verdict with the exact SQL and the words “feature is INERT until applied,” plus a queued manual-tier “apply DDL” task. Then closingdoneis honest, because the pending state is itself durable and visible. - Run the mirror check too: the record can also over-remember. Before working any task that references a PR/commit or was sync-imported, spend 30 seconds checking whether reality already delivered it.
- Organs beat views, always. A dashboard, briefing, or TUI is a lens; if it disagrees with the table or the file, the table or the file is right — regenerate the lens.
The worked example. The DDL_PENDING convention’s origin is this manual’s own subject matter: the file-lock worktree fix merged with its new lock-table column unapplied — “merged, closed, silently inert in prod,” invisible in the queue. Every screen said the fleet was protected; the running behavior said otherwise, and nothing anywhere recorded the gap. The convention was codified two days later — and the same change shipped a staleness pre-check, which “closed six already-delivered tasks” in one loop that same day: durability failing in both directions at once, fixed with one principle. Infrastructure rhymes with it: a reboot cleared the scratch directory and the entire managed cron block died silently (every job’s log redirect failed); the fix recreates the scratch directory before each tick, so even the plumbing assumes its scratch space has already been deleted.
The failure it prevents. Amnesia in all its costumes: work that evaporates with the session that did it, decisions that exist only in a dead context window, ghost features the whole fleet believes are live, and effort re-spent on tasks reality already closed.
9. Liveness is a timestamp with a freshness window
In a fleet, “is that session alive?” is never a status field — it is evidence with an expiry. Every lock, claim, and slot is held by something that can die mid-hold, so every one of them needs a freshness window and a reaper. And the inverse discipline matters just as much: a failed liveness check is not evidence of death.
The procedure.
- Trust heartbeats, not status columns. A registration hook refreshes the heartbeat timestamp on every tool call; consumers apply windows — the file-lock hook ignores peers whose heartbeats are >5 minutes stale, so a crashed session’s locks fade instead of wedging the fleet.
- Give every held resource a reaper. Crashed autonomous runs are reaped within 30 minutes. A converge pass releases orphaned task claims — dead claimant →
status=ready— so work re-enters the queue instead of rotting under a ghost. Governor slots are reaped on dead PID or a 900-second TTL: “akill -9mid-run cannot leak a slot forever.” - Notice when a reaper is also load-bearing for capacity: the concurrency gate counts ACTIVE autonomous sessions, so reaping is what self-heals the gate — without it, every crash permanently eats a slot of the max.
- Distinguish “dead” from “couldn’t check.” A transient probe failure must not be cached as a negative, and an error object must never be mistaken for data — a hardening fix addressed exactly this: database error-objects were poisoning the peers cache that lock decisions read.
- For a live-but-blocked peer, talk before declaring death: send a file-release or a question; an active worker sees it on its next tool call. Real silence past the window → let the reaper run and take the work.
- Make session start triage time: an intake step surfaces orphaned sessions from AFK kills and crashes before you claim new work, so you resume or archive their state deliberately instead of duplicating it.
The worked example. The run-lifecycle closes its own loop: workers close via a finally:-guarded path, and a reaper backstops within 30 minutes, and that reaping self-heals the concurrency gate — belt, suspenders, and the reason stated. The opposite hazard is the two hardening fixes together: the machinery that decides liveness must distrust its own flakiness, or one bad network minute becomes “all peers look dead” (locks ignored → collisions) or “the column looks missing” (scoping silently off → the over-blocking returns). Both fixes exist because both happened.
The failure it prevents. The two wedges: a fleet frozen forever behind a dead session’s locks, claims, and slots — and the reverse, a live peer declared dead on flaky evidence, its locks bypassed and its work clobbered by a helpful survivor.
10. Make fleet-wide changes fail loud: heal, backstop, lint
The most dangerous pattern in this corpus is quiet: a machine-wide mutation (global git config, env, launchd, shared dotfiles) executed on a cron loop with no detector. It breaks everything, it breaks it silently, and a one-time manual fix never sticks because the loop re-applies the damage on schedule.
The procedure.
- Recognize the pattern at review time, before shipping: does this change mutate state outside its own repo/session, does anything re-run it automatically, and would anyone notice drift? Three bad answers = stop.
- Scope mutations as narrowly as possible: per-command
git -c …, never--global. Most machine-wide mutations are conveniences that were never actually required. - If you must own shared state, ship three layers with it: a healer that detects and repairs drift at runtime, a cron backstop so the healer can’t be forgotten (every ten minutes, on an offset), and a
--checklint that FAILS if any tracked script reintroduces the setter — the review gate. - Leave a tripwire: the healer logs only when it had to act, so a non-empty guard-log means a setter still exists somewhere — grep for it and find it.
- Extend “loud” to schemas and envelopes. The message-envelope decision exists because consumers selected a
typecolumn that doesn’t exist, gotnull, and “end up guessing from payload shape” — a silent misparse. Onemsg_typecolumn, one table, sends only through the two send functions, never raw inserts. - Run the class lints before shipping anything near them: the config-drift
--checkbefore touching git config is a standing rule, and it sits inside the auto-merge rails for the same reason.
The worked example. A cleanup cron set git config --global url."git@github.com:".insteadOf "https://github.com/" — a one-line convenience. The fleet machines have no SSH keys (auth is HTTPS via the platform credential helper), so every headless clone and fetch on every machine silently re-routed to SSH and died. A full fleet-wide git outage — and hand-fixing the config couldn’t stick, because the cron kept re-applying it. The answer addressed the class, not the instance: a config-drift guard with heal + backstop + lint (self-tested by planting a setter and catching it — “the review gate that would have caught it”), plus the general rule written into the protocol so the next machine-wide-mutation-on-cron gets recognized at review, whatever state it mutates.
The failure it prevents. Silent fleet-wide rot: one convenient line in one script that breaks every machine, recurs on schedule against every manual fix, and surfaces days later as a mystery outage nobody can bisect — because the thing that broke the fleet isn’t in any diff anyone is looking at.
The self-test — five questions before dispatching any multi-session work
Run these before the dispatch, out loud, with real numbers and real paths. Any answer of “I don’t know” means the dispatch waits until you do.
- Width. At the worst instant of this plan, how many model requests are in flight across the whole machine — and the fleet? Count your subagents, peers’ work, headless loops, and every cron that could wake inside the window. Is every headless call in a governor slot, interactive fan-out ≤3, and every schedule offset? Say the number.
- Collision. For each session: which task (claimed atomically, where peers can see it), which worktree lane, which files? Are the file sets disjoint by construction — and is that enforced by hooks and RPC fences, or only by everyone promising to be careful?
- Gate. Trace every path from this work to a human or to the main branch — email, message, merge, deploy, payment. Does each path dead-end in the approval queue by structure (unclaimable tier, missing tool, staged outbox), or only by instruction? Instructions don’t hold at 3am.
- Wreckage. If every session dies mid-flight tonight: what gets reaped and within what window, what survives in the organs, what is queued for the next session, and is anything shipped-but-inert flagged
DDL_PENDING? Whatever lives only in a context window or a peer message is already lost — count it as such now. - Silence. For each new mechanism or shared-state mutation this dispatch introduces: when it breaks or drifts, what turns red, who heals it, and does it fail open or closed on purpose, chosen by what it protects? If the honest answer is “nobody notices until everything is broken,” the detector ships first, then the change.
These ten moves are one habit: never rely on an agent being careful when you can make the system incapable of the failure. Everything above was learned at full price — a shutdown, an outage, a deadlocked morning, a ghost feature. The corpus already paid for these lessons once. Don’t buy them again.