Resume: bind the transcript base after the Runner accepts
Freezes on merge; later changes supersede by citation, never rewrite. Linear: RIG-4297. Parent:
docs/designs/agent/compass-agent-session-persistence/design.md§ T4 and § T6. Sequenced with RIG-3107. This record adds to the parent’s rebase model; it does not rewrite it.
Problem / Intent
Section titled “Problem / Intent”Both resume paths bind the transcript base before the Runner accepts the
Start. service.startResumeSession (go/server/service.go) and
lifecycleService.resumeSession (go/server/lifecycle.go) call
Store.BindLifetime and only then call Hub.StartResume
(go/internal/runnerhub/resume_start.go). BindLifetime
(go/internal/store/agent_transcripts.go) sets
agent_sessions.base_entry_seq to the session’s current MAX(entry_seq).
Store.AppendTranscriptEntry reads that one base on every frame and writes
the row at base + entry_seq.
A frame does not say which lifetime sent it. TranscriptEntry
(proto/compass/v1/agent.proto) carries entry_json, checkpoint, and the
agent-stamped entry_seq (from 1 per agent process). The durable commit,
Hub.CommitConversationFrame (go/internal/runnerhub/relay_comms.go), is
keyed by the logical session id, and a resume reuses that id
(agentHost.Start in go/internal/runner/host.go).
So a refused resume still moves the base of a live session:
- Lifetime L1 is live. It has written k frames at base B1.
- A second resume calls
BindLifetime. The base becomes B2 = B1 + k. agentHost.Startrefuses witherrAlreadyRunning(RUNNER_ERROR_CODE_ALREADY_RUNNING,go/internal/runner/dispatch.go).- L1’s next frame, k + 1, lands at B2 + k + 1. The rows B1 + k + 1 to B2 + k never exist.
TestWakeAgentLiveElsewhereIsRefusedByRunner
(go/server/lifecycle_wake_pgtest_test.go) witnesses the moved base. The gap
harms reads: Hub.ReconstructSessionBody (go/internal/runnerhub/reconstruct.go)
numbers each archived segment line as MinEntrySeq + i, so a segment that
spans a gap gets wrong seqs and merges out of order against the PG tail.
The wake path narrows it: lifecycleService.WakeAgent skips an agent that
Hub.CachedSessionForAccount reports live, so a wake needs the gap between
the Runner’s Start result and the map write in Hub.promoteSession.
service.startResumeSession has no live check. An authorized client that
resumes a session already live reaches the race today, even with one Server.
RIG-3107 (multi-instance) widens the wake path too: a Server whose cache does
not hold a session that another instance promoted will resume it.
Matt ruled Opt 3 + 1 on RIG-4297. Opt 3: accept the race now and document it. Opt 1: design bind-after-accept, sequenced with RIG-3107. Frames stay without a lifetime id. This record is the Opt 1 design for resume. The Opt 3 comments ship as their own code PR.
Reload has a related bug that this record does not fix (see Open
Questions): agentHost.reloadLocked relaunches under the same session id with
no rebind and no resume file, so the new process’s frame 1 hits
ErrConflict on the row the old process wrote.
Approach
Section titled “Approach”The Runner binds a resume’s lifetime, not the Server. The bind moves to the one
point that already knows the Start is accepted and no frame of the new
lifetime exists yet: inside agentHost.Start, under lockContainer, after
the errAlreadyRunning checks and before link.StartAgent.
- New unary
RunnerService.BindLifetime(container_name, session_id), Runner to Server, on the internalproto/compass/v1/runner.proto. It returns an empty response. It is the same Runner-initiated unary shape asFetchSecrets. The Server still gains no inbound route. - The handler authorizes like
Handler.FetchSecrets:Hub.AccountForContainerfor the authenticated Runner and the container. A resume already depends on that binding:Hub.promoteSessionreads it after Start, and without it the resumed session gets no session binding and its comms calls fail closed. So the bind adds no new precondition. The Runner door sets no tenant, so the handler then resolvesStore.AccountTenantunderstore.WithSystemRoleand binds onstore.WithTenant(store.WithoutSystemRole(ctx), tenant), aslifecycleService.wakeCtxdoes. The bind matches only a session whoseagent_account_idis that account. A foreign container, another account’s session, and an unknown session all return the samePermissionDenied. - A bind error fails the Start before the agent runs. The container lock is
still held, so no second Start can interleave. The error maps to
RUNNER_ERROR_CODE_INTERNAL, neverNOT_FOUND:resumeSessiontreatsNOT_FOUNDas a missing container and would reprovision. One visible change: a session deleted betweenstartResumeSession’s authz and the bind now reaches the client asInternal, not today’sNotFound. The window is a concurrent delete, and both codes fail the resume. - A fresh Start skips the bind.
service.StartAgentSessionwrites the session row only after Start returns, so a bind there would find no row and fail every fresh Start. The default base 0 is correct. service.startResumeSessionandlifecycleService.resumeSessiondrop theirStore.BindLifetimecall. Authz,ReconstructSessionBody,StartResume, and the reprovision retry are unchanged. The retry binds once: the registry miss returnserrSessionUnknownbefore the bind runs.
The invariant this relies on is at most one live lifetime per session, and
today single-Runner placement enforces it. The container lock serializes
binds across any number of Servers, because one Runner owns the container. It
does not serialize across Runners: a stale placement on Runner B can still
bind a session live on Runner A. Multi-Runner placement must close that before
it ships (deferral below). Hub.AccountForContainer reads the in-memory
containerAccounts of the instance that relayed Provision. Under RIG-3107 the
bind must reach that instance or read the durable placement; that is the same
dependency FetchSecretsByContainer already has.
Alternatives considered
Section titled “Alternatives considered”Bind on the Server after StartResume returns
Section titled “Bind on the Server after StartResume returns”The agent is already running when the result comes back, and its first frames can commit before the bind. They land at the old base, on rows the previous lifetime wrote. Closing that gap needs the Server to hold every commit for the session until the bind, which is the same coordination with more moving parts.
Restore the old base when the Runner refuses
Section titled “Restore the old base when the Runner refuses”L1’s frames between the bind and the restore already landed at B2 + seq. A compare-and-set restore cannot move those rows back.
A lifetime id on every frame
Section titled “A lifetime id on every frame”The store could key the rebase per lifetime. It is a wire change on every frame, and the ruling keeps frames without one. The parent record rejected it for the same reason (§ T4).
Derive the lifetime from the idempotency-key nonce
Section titled “Derive the lifetime from the idempotency-key nonce”The agent mints each key as <nonce>-<n> with one nonce per process
(createSocketFrameSink, packages/compass-agent/src/transport/frame-sink.ts).
The Server could group by that prefix. It is a lifetime id in disguise, which
the ruling excludes, and it turns an opaque dedup key into a parsed field.
Session-level advisory lock on the Server
Section titled “Session-level advisory lock on the Server”It serializes Servers but still binds before the Runner decides. A refused resume still moves the base.
Global Constraints
Section titled “Global Constraints”- Frames carry no lifetime id.
TranscriptEntryis unchanged. - The new RPC is internal (
runner.proto). No public proto changes. - The bind runs under the session’s tenant, resolved from the account. Never the door ctx, never the system role for the write.
- Fail closed: an unauthorized bind is
PermissionDeniedand matches an unknown session byte for byte. - A bind failure never wraps
errSessionUnknown. - Lands after the accepted-race comment PR for RIG-4297, whose comments T1 deletes.
- Tests advance on observed state (channels,
testing/synctest), never sleeps. - Land before RIG-3107 enables a second Server instance.
- Move the resume bind to the Runner (one PR). Proto and generated code,
Handler.BindLifetime,Hub.BindLifetime, the account-scoped store bind,ServerLink.BindLifetime, and the call inagentHost.Start. The same commit removes the Server-side calls inservice.goandlifecycle.go, so exactly one bind runs per lifetime, and it rewrites the Server pgtests that assert the Server-side bind. Split, the RPC would merge with no caller.
-
T1 — Runner binds a resume under the container lock (
implement-go). Interfaces:rpc BindLifetime(BindLifetimeRequest) returns (BindLifetimeResponse); request{container_name string, session_id string}, empty response.func (h *Handler) BindLifetime(ctx context.Context, req *connect.Request[compassv1internal.BindLifetimeRequest]) (*connect.Response[compassv1internal.BindLifetimeResponse], error).func (h *Hub) BindLifetime(ctx context.Context, runnerID, containerName, sessionID string) error: authz, tenant resolve, bind.- New Hub seam, set by
func (h *Hub) SetLifetimeBinder(b LifetimeBinder); nil fails the RPCUnavailable, likeerrTranscriptsUnavailable:type LifetimeBinder interface { AccountTenant(ctx context.Context, account store.AccountID) (store.TenantID, error); BindLifetime(ctx context.Context, sessionID string, account store.AccountID) (uint64, error) }.*store.Storesatisfies it. Store.BindLifetimebecomesfunc (s *Store) BindLifetime(ctx context.Context, sessionID string, account AccountID) (uint64, error); the sqlcBindLifetimequery addsAND agent_account_id = $2. Missing or foreign row isErrNotFound. The base return stays for the store tests that assert it. Every caller and test moves in the same commit.func (l *ServerLink) BindLifetime(ctx context.Context, containerName, sessionID string) error, called fromagentHost.Startonly whenreq.GetResumeSessionId() != "".- Every test
RunnerServiceHandlerfake that serves a resume Start (capturePublish,recordingRelay,w3Relayingo/internal/runner) implementsBindLifetimeand records calls, or the embeddedUnimplementedhandler fails those Starts. - Production wiring:
hub.SetLifetimeBinder(st)ingo/server/sinks.gobesidehub.SetTranscriptStore(st). The leg-five e2e resume (go/e2e/legfive_test.go) runs against the real stack, so it fails if the wiring is missing.
Tests:
- Handler pgtest: a tenant-B session binds through the handler. A foreign
container, an account mismatch, and an unknown session return identical
PermissionDenied. - Host tests: a resume Start on a live container returns
errAlreadyRunningand the fake handler records zero binds. A fresh Start records zero binds. A bind error fails Start beforeStartAgentand maps toRUNNER_ERROR_CODE_INTERNAL. - Server pgtests: delete the
boundBaseassertions inTestWakeAgentPriorSessionResumes,TestWakeAgentLiveElsewhereIsRefusedByRunner(go/server/lifecycle_wake_pgtest_test.go), andTestStartAgentSessionResumeKeyedOnStableLogicalIdAcrossResumes(go/server/service_resume_pgtest_test.go). Their fake Runner never binds, so a base check there cannot fail. The refused-resume witness moves to the host test above.
Comment sweep: the accepted-race comments in
service.goandlifecycle.go, the bind-ordering comment onHub.StartResume, theBindLifetimedoc ingo/internal/store/agent_transcripts.go, the leg-five comment ingo/e2e/legfive_test.go, and the named-entrypoint list in thestore.WithSystemRoledoc, which gains the bind’s tenant lookup.
Open Questions
Section titled “Open Questions”- Non-load-bearing deferral (out of scope; RIG-4451): what is a
Reload’s transcript?
reloadLockedrelaunches withh.agentEnv(handle)and noResumeSessionFile, so the new process starts a new SDK session (cli.tscallsmanager.setSessionFileonly with a resume file). Today its frame 1 fails withErrConflictand the tee fails the session. Rebinding alone would make two SDK sessions share one logical transcript. Options: (a) Reload is a resume: materialize a reconstructed body, then bind (recommended); (b) Reload mints a new logical session id; (c) Reload keeps the id and the new process opens with a checkpoint. Tracked on RIG-4451 (human-action); a follow-up record designs the chosen option. Reload now calls the bind beforeStartAgent(option (a)’s bind half, RIG-4418). - Non-load-bearing deferral: in-flight commit across a rebind. A dead
process can have one
CommitConversationFramestill in the Server after the Runner sees its call cancelled: the Gateway commits on the agent’s request ctx, and the append reads the base and inserts in two statements. A bind that lands between them collides the old frame with the new lifetime’s frame 1. Reload under (a) or (c) hits it, and so does a resume of an ERRORED session, whichagentHost.Startadmits. Today’s Server-side bind has the same window for that resume, so this record does not regress it. The fix is a detached commit ctx plus a drain, or row locks between bind and append. - Non-load-bearing deferral: cross-Runner fence. Multi-Runner placement
must refuse a bind when the durable
session_bindingsrow names another Runner, before more than one Runner can hold an agent. Reload treats any bind denial as “no row yet”; that must split from a placement miss by then.