What survives when an AI worker disappears?
A useful AI workflow often outlasts the process that begins it. The worker might restart, a provider might respond late or a reviewer might return tomorrow. If the only account of the work lives in the worker's memory, the next process inherits a puzzle rather than a task. It sees a pending job but cannot tell which inputs were used, which actions happened or which result deserves review.
Durability therefore requires more than storing a status field. A task needs a recoverable account of its authority, inputs, attempts, artifacts and unresolved effects. My design question is simple: what should another authorized worker be able to establish without trusting the vanished worker's narrative?
I examine that question through a synthetic research workflow. It collects three public source passages, prepares a comparison memo and checks each substantive claim against the selected evidence. The workflow may save a local draft artifact. It does not publish the memo, contact anyone or access private company materials. This is an illustrative architecture exercise, not an account of a system I have deployed at Ethen or elsewhere.
Give the job a stable identity
The logical job is the owner's request for a comparison memo under a particular scope. Its identity should remain stable across attempts. A worker identifier cannot serve this purpose because workers change. A generated model response cannot serve it either, because the workflow may produce several responses before the owner accepts one.
The job record should identify its owner, permitted scope, requested output and input revision. It should also record the policy under which work was admitted. If the request changes, the system needs to represent that change explicitly. Otherwise an apparently continuous job can quietly move from comparing public sources to drawing conclusions from materials the owner never authorized it to use.
In the example, job J refers to a memo comparing three named public documents. Attempt A reads the documents and prepares a draft. Attempt B later resumes verification. Both belong to J, but their authority and completed steps should remain separately inspectable. This prevents a recovery worker from presenting all work as its own fresh execution.
A stable identity is useful for deduplication, but it does not settle the meaning of repeated requests. Two identical-looking prompts may represent the same retried submission or two intentional jobs. The caller and service need an explicit submission identity and a documented scope for reuse. Text similarity is an unreliable substitute for that agreement.
A recovery record for each step
| Step | Evidence retained | Boundary a replacement worker must check |
|---|---|---|
| Fetch sources | URL, retrieval time, permitted snapshot and input revision | Access remains authorized; missing sources remain visible |
| Prepare memo | Artifact ID, artifact revision and input-bundle revision | The draft corresponds to the recorded inputs |
| Review memo | Reviewer or verifier identity, rubric and exact reviewed revision | Approval applies to this revision, not an earlier one |
| Complete job | Known effect receipts, current ownership and resolved required claims | No unknown dispatch or stale review is hidden by a completed label |
The table is an original synthetic workflow. It describes evidence to retain rather than a universal completion API. In particular, durable state does not grant permission to retain a confidential source indefinitely or expose it to a replacement worker. The record must carry the relevant access and retention boundary as well as the source identity.
Store the evidence boundary
The memo should reference source identities and the actual passages used, rather than only remembering the documents' current URLs. A public page may change after the job runs. A future reviewer who reads the new page could reach a different conclusion while believing they inspected the same evidence.
A suitable evidence record can include the fetched URL, retrieval time, content fingerprint and selected passage boundaries. Where permitted, retaining a restricted snapshot helps reconstruct the review. A fingerprint helps detect whether a saved item changed; it does not prove that the item was accurate, complete or lawfully acquired. Those are separate questions.
For this public-source example, the job should also record failed fetches and unsupported formats. If one of three sources was unavailable, the comparison cannot honestly imply equal coverage of all three. A recovery worker should see that limitation before generating another confident summary from the two documents that happened to be accessible.
The W3C PROV primer provides a vocabulary for describing entities, activities and agents in provenance records. That vocabulary is useful background for relating inputs, transformations and outputs. My proposed job record does not claim conformance to PROV, and the existence of a provenance graph would not independently certify the memo's factual quality.
A checkpoint must say what completed
“Running” is too broad to be a useful checkpoint. The worker could be fetching a source, waiting for a model or saving an artifact. After a restart, those states demand different next actions. A checkpoint should identify the completed step and the evidence that supports treating it as completed.
For the comparison job, the source-collection checkpoint references three fetch records. The draft checkpoint references a specific artifact revision and the input bundle used to generate it. The verification checkpoint references a claim review for that same revision. A completion checkpoint should require the evidence specified by the task, rather than accepting whichever artifact happens to be newest.
Saving a checkpoint and saving its artifact can fail separately. If the artifact is durable but the checkpoint is not, another worker may repeat the step. If the checkpoint is durable but the artifact is missing, the next worker may trust a result it cannot inspect. The design must address those gaps through the storage system's actual guarantees or through explicit reconciliation.
One practical approach is to persist an artifact reference only after the artifact is available, then validate that reference when resuming. This does not eliminate every failure window. It does give the recovery path something concrete to inspect. The important property is that a checkpoint is supported by recoverable evidence, not merely by the previous worker's intention.
Worker ownership must expire safely
A job service often needs temporary ownership so that two workers do not act on the same task simultaneously. An ownership record must account for pauses and crashes. If ownership never expires, a vanished worker can strand the job. If expiration immediately grants unrestricted authority to another worker, the original worker may return and continue acting at the same time.
A proposed solution is to associate ownership with a monotonically advancing generation. A new acquisition receives a new generation. Writes from an earlier generation are rejected at the authoritative boundary. The local companion fixture illustrates this narrow rule: generation four cannot update a record after ownership has advanced to generation five.
That fixture does not prove that a distributed implementation enforces the rule. Enforcement depends on where the check occurs, how it is made atomic with the write and which destinations participate. If a worker can directly call an external service that does not recognize the generation, protecting the job database alone does not prevent a stale external action.
Worker ownership and user authority are also different things. Acquiring the current generation grants the right to work on the job according to its contract; it does not expand the owner's permissions. A fresh worker must still respect revoked access, changed policy and the boundary against publishing the memo. Recovery cannot turn administrative ownership into unrestricted task authority.
Preserve the dispatch boundary
Before a network operation, a worker knows it has not dispatched that operation. After dispatch, a missing response can leave the result unknown. The job record should preserve that distinction. A recovered process should not assume that the absence of a completed-step record proves the absence of an external effect.
Suppose the comparison draft was saved through a remote artifact service, but the worker crashed before recording the returned artifact identifier. The next worker knows that dispatch occurred and that acknowledgment is missing. It needs a reconciliation procedure, perhaps based on a stable operation identity or a destination lookup. Starting a new creation request without investigation could create a duplicate.
The HTTP semantics specification supplies useful background on request semantics. Application recovery still depends on the destination's documented behavior. An idempotency mechanism may have limits involving retention, scope or request matching. Those limits belong in the recovery design rather than in a footnote that nobody consults during an incident.
In this example, an unresolved dispatch prevents automatic repetition of an effectful step. The job can continue with independent read-only checks if they do not rely on the missing result. It can also pause with an actionable reconciliation record. Pausing is a valid state when continuing would require pretending that an unknown outcome is known.
Cancellation is a state transition
An owner might cancel the comparison while a provider request is in flight. The system can stop scheduling future work, but it may be unable to withdraw a request already received by the provider. A cancellation interface should not promise that no remote processing occurred unless the system can establish that fact.
The job record should distinguish cancellation requested from cancellation acknowledged at the relevant boundaries. A late response can be retained or discarded according to the approved retention policy, but it should not automatically reopen the task. The owner's decision to stop remains part of the authority record even when the worker later receives a usable answer.
A recovered worker should check cancellation before dispatching another step. It should also check again at the boundary where the next effect is authorized. A check performed only when the job was created cannot account for a cancellation that arrived during a long source fetch. Timing matters because authority changes while work is running.
Cancellation does not erase completed actions. The memo draft may already exist, and source records may already have been saved. The system should describe those effects accurately and apply the actual retention rules. Silently deleting evidence to make the task look as though it never ran would damage both diagnosis and the owner's understanding of what happened.
Verification must follow the artifact revision
The worker's final prose is not sufficient evidence that the memo is ready. Each substantial factual claim should be connected to a supporting passage or marked unresolved. Interpretive claims should be recognizable as interpretation. The review should identify contradictions and distinguish a source's stated position from an independently established result.
The verification record must refer to the artifact revision it checked. If another worker edits the memo after review, the old verification result should not automatically attach to the new version. Some edits are cosmetic, but determining that is itself a policy decision. Treating every saved revision as equivalent makes it impossible to know what the reviewer actually approved.
A synthetic record with draft revision two and review revision one should therefore fail the completion predicate. A record with revision two reviewed but an unresolved required claim should also fail it. The companion fixture checks both cases. It illustrates a completion boundary rather than measuring the effectiveness of a semantic verifier.
The public Ethen durable-job article is relevant contextual reading about recovery infrastructure. Its implementation narrative belongs to its publisher. This essay instead proposes an evidence-centered memo workflow and completion predicates. Public descriptions do not establish my personal authorship of the service or independently demonstrate its behavior in production.
Handoffs need a useful summary
The next worker should receive a concise account of what is established, what remains uncertain and what action is permitted next. A raw event stream can support investigation, but it is a poor default instruction. A handoff that says “continue the task” may encourage the worker to repeat work or skip unresolved checks.
For job J, a useful handoff identifies the three source snapshots, the latest draft revision, the missing acknowledgment from artifact creation and the prohibition against publication. It recommends reconciling the artifact operation before creating another draft. If source three could not be fetched, it says so explicitly rather than relying on the recovery worker to infer the gap from an error buried in logs.
The summary should be derived from authoritative records and link back to them. It should not become a second, independently editable truth store. Otherwise the checkpoint can say one thing while the handoff says another, and workers will choose whichever is easier to follow. Disagreements should be visible rather than resolved by silently trusting the most recent prose.
An owner-facing handoff has a different purpose. It should explain the available artifact, the unresolved issue and any concrete decision needed. It should not expose confidential traces or require the owner to understand database generations. Good internal records make that concise explanation possible without hiding the real state of the task.
Retention is part of the architecture
Evidence has a lifecycle. Keeping every source, prompt, artifact and event indefinitely is not automatically better than keeping too little. A system needs a policy explaining which records support recovery, how long they remain available and who can inspect them. Those decisions affect whether a later audit can reconstruct a completed task.
For a public research memo, retaining source fingerprints and selected passages may be useful. For private material, the same retention scheme could be inappropriate. Access controls and deletion obligations must follow the material's actual requirements. This article does not propose a universal retention period or permission to copy private company repositories into an evidence bundle.
Removing an artifact can also weaken a completion record. A receipt may still show that verification occurred, but another reviewer can no longer inspect the verified object. The system should represent that difference. A historical statement that a check passed is weaker than a still-inspectable package containing the exact checked revision and its supporting evidence.
Retention tests should include expired references and revoked access. The job must not keep presenting a missing artifact as available or encourage a worker to reacquire private content without authorization. Recovery should respect present access conditions while preserving an honest account of earlier activity.
Rehearse the owner handoff
Before deployment, a team can give a reviewer only the durable job package and ask for the permitted next action. The reviewer should not need the vanished worker's chat history or a developer's explanation. If the package leaves them guessing whether an artifact exists, the recovery design is incomplete at a practical boundary.
For the synthetic memo, the rehearsal should yield a specific answer: inspect the unresolved artifact operation, preserve the current draft revision and keep publication prohibited. If the reviewer instead starts a fresh creation request, the package may be missing a visible dispatch record or a clear reconciliation instruction.
This rehearsal tests comprehensibility rather than distributed correctness. It complements failure injection by checking whether the records communicate the state they are meant to preserve. A technically sound checkpoint that nobody can interpret during an incident can still lead to unsafe recovery choices.
Test failures between the steps
The most informative tests interrupt the workflow at boundaries: after a fetch but before its checkpoint, after artifact dispatch but before acknowledgment, and after review but before completion. Each interruption asks whether a fresh worker can discover the correct next action from durable records alone.
The illustrative exercise covers stale ownership, unknown dispatch outcomes and review revision mismatch. It uses synthetic records, not real queues or providers. A complete implementation test would need the actual storage and destination services, realistic concurrency and controlled failure injection. The distinction prevents a few illustrative assertions from becoming an exaggerated durability claim.
Operational evaluation should track unresolved jobs, reconciliation time, duplicate artifacts and completion rejected for missing evidence. A low crash rate does not establish good recovery. The important question is what happens when a crash actually occurs and whether the owner receives an accurate description of the result.
Give a replacement worker the job record without its predecessor’s chat history. Can it identify the latest reviewed revision, the unresolved effect and the permitted next step? If not, improve the recovery record before adding another worker. Related architecture commentary appears in projects and writing.
Sources
- W3C — PROV Primer — primary methodological or standards source.
- RFC 9110 — HTTP Semantics — primary methodological or standards source.
- Ethen — Inside the Durable Job Service — company-authored context; not independent implementation evidence.