A fallback is a new execution decision

The first provider times out. The second provider is healthy. A routing dashboard offers a comforting answer: send the request somewhere else. For a text completion with no sensitive inputs and no external effects, that might be an acceptable recovery. For a workflow that reads a protected document, calls tools or changes a shared system, availability is only one part of the decision.

A fallback can change where data travels, which capabilities are available, what an approval covers and whether an operation happens twice. Treating it as an invisible transport detail makes those changes difficult to see. My proposed design starts with a stricter rule: every fallback must satisfy the original task contract and account for what the previous attempt may already have done.

The example here is deliberately synthetic. An assistant reads an internal incident report and prepares a structured summary for its owner to review. It may create a draft artifact, but it may not send that artifact to another person. There are no real customer records, provider requests or production measurements behind the example. The purpose is to examine recovery decisions before they acquire real consequences.

Write the contract before choosing a route

The request needs more than a prompt. It needs a contract describing permitted data destinations, required inputs, expected output, allowed tools and the boundary between preparation and action. In this example, the model must accept the report format, return a summary matching a defined structure and operate through a destination approved for that report. Sending the summary remains outside the contract.

These conditions should not disappear when the preferred route fails. A cheaper alternative with no support for the report's input format cannot satisfy the task merely because it can produce fluent text. A provider that accepts the format but violates the destination restriction is also unsuitable. Ranking those alternatives together hides the fact that some should never enter the ranking.

Consider three invented routes. Route A meets the input, destination and output requirements. Route B accepts only plain text, while the report includes diagrams needed for interpretation. Route C supports diagrams but sends the report through an unapproved destination. Only A is eligible. If A becomes unavailable, the result is no eligible route, even if B and C advertise excellent latency.

That refusal is useful information. It tells the caller that fulfilling the task now requires changing something the owner actually cares about. The caller can wait, ask for an authorized transformation of the input or seek a new destination decision. It should not silently reinterpret the request as an easier task and return an answer under the original success label.

Separate request failure from task failure

A failed network request does not always establish that the remote operation failed. The connection might break before dispatch, while the provider processes the request or after the response was generated. Those possibilities have different recovery implications. An application that collapses all of them into a timeout loses the information needed to make a careful next decision.

For the summary example, failure before dispatch means the report was not sent by that attempt. Failure after dispatch means the report may have reached the provider even if no response was received. If the workflow can also create an artifact, a lost response could leave a draft that exists but is not recorded locally. Recovery has to account for both disclosure and duplicate creation.

The HTTP semantics specification distinguishes method semantics, including idempotency. It does not make an arbitrary application workflow safe to repeat. A provider's endpoint, a tool operation and the surrounding job each need their own documented behavior. Using the same HTTP method twice is not an application-level proof that only one draft was created.

A practical state record should distinguish at least not dispatched, dispatched with a known result and dispatched with an unknown result. These are proposed application states, not a claim about a particular provider's API. Where reliable reconciliation is available, an unknown result can be investigated. Where it is unavailable, the workflow should preserve the uncertainty instead of manufacturing a clean failure.

Recovery decisions for the synthetic summary task

Table 1: Synthetic worked example
State of the previous attempt What is known Permitted next step under this design
Not dispatched This attempt sent no report Recheck route eligibility and current authorization before trying again
Completed with a receipt The named effect is recorded Continue from that effect instead of recreating it
Dispatched, outcome unknown Disclosure or artifact creation may already have occurred Reconcile with the destination; pause if the effect cannot be established
New route changes the task Original conditions are no longer satisfied Seek authority for the changed contract or stop

This is a proposed decision table, not an assertion that a particular provider implements these receipts. It deliberately leaves an unresolved effect unresolved. A successful second response cannot retroactively explain the first attempt's missing outcome.

Make the attempt visible

Each execution attempt should have an identifier linked to the logical task. The task identifier means “prepare this owner's incident summary.” The attempt identifier means “try this route using this input revision and policy state.” If the first attempt is uncertain and the second starts, those identifiers prevent the audit record from presenting the second attempt as the only thing that happened.

The record should describe the selected destination, the resolved model identity, the relevant contract version and the dispatch state. It should also record the reason for a fallback. A route change prompted by a capability mismatch has different implications from a change prompted by a temporary connection failure. Grouping both under a generic retry label makes later diagnosis harder.

Recording an attempt does not require saving every confidential prompt in a broadly accessible log. The log can reference a restricted input record and include the evidence needed for authorized review. Access rules, retention and redaction still matter. An audit trail that copies the report into multiple unrestricted systems solves one observability problem by creating a disclosure problem.

The identity of the input revision matters as well. If the owner updates the incident report while an attempt is in flight, the next attempt must choose deliberately between continuing the old task and beginning a new one. A fallback should not quietly mix an earlier approval, a later document and a response generated from an unidentified version.

Recovery needs its own admission check

A recovery decision should rerun the relevant admission checks. It should use current policy, current route availability and the original task requirements. A successful admission decision from several minutes earlier does not establish that the next destination remains authorized. Nor does a recently healthy provider prove that it can accept this particular task now.

In the synthetic example, A becomes unavailable after dispatch and its result is unknown. B remains incapable of reading the required diagrams. C remains outside the allowed destination set. The correct admission outcome is still no candidate. A fallback policy that chooses C because it is healthy would change a hard requirement into a preference at precisely the moment the system is under pressure.

A route can also become eligible only after the owner approves a different task. For example, the owner might authorize a text-only summary with the diagrams excluded and the resulting limitation disclosed. That is a new contract, not a successful execution of the original one. The resulting artifact should carry the changed scope so that another reader does not mistake it for a complete incident analysis.

The public Ethen routing article discusses eligibility and separates different routing paths. I use it as attributed context, rather than evidence that my proposed recovery design is implemented there. Its discussion also overlaps with the general subject of admission. This article's distinct focus is the state of the failed attempt and the permissions that must survive a route change.

Tool effects make retries harder

Generating a summary and creating its draft artifact are different operations. The model may be called again without necessarily repeating the artifact creation. Combining both into one opaque job makes that distinction harder to enforce. A clearer design records completion of each step and places effectful operations behind explicit boundaries.

Suppose the application asked a storage service to create draft D, then lost the response. Repeating the creation with a fresh identifier may create another draft. Reusing a documented idempotency mechanism may help, but only within that mechanism's actual scope and retention period. The application still needs to know what request identity means and what happens when the original request is no longer recognized.

Reconciliation might query a stable artifact identifier or consult an operation receipt. If neither is available, the system can pause and ask the owner to inspect the destination. That costs time, but it keeps the uncertainty visible. Declaring the operation absent because the local database lacks a receipt would confuse absence of evidence with evidence of absence.

A recovery budget should include those investigation costs. Counting only the second model call makes fallback look cheaper than it is. A workflow may spend less on inference and more on duplicate cleanup, owner review or support. The useful accounting unit is the completed, verified task, with recovery effort included consistently.

Preserve the owner's authority

An owner might approve automatic fallback within a specific destination set. Another might approve only the explicitly selected provider. Both are reasonable product choices if the application represents them accurately. The mistake is to convert a general desire for reliability into permission for every recovery path the system knows how to execute.

For the incident summary, permission to create a private draft does not become permission to email it when the draft service fails. Permission to use one provider does not necessarily cover another provider with different data handling. These are task-specific boundaries. They should be evaluated against the current authorization record, rather than inferred from a generic checkbox whose meaning has grown over time.

The interface should explain material recovery changes before asking for a decision. “Try again” is insufficient when the new attempt changes destination or scope. A useful review shows what remains the same, what changes and what is unresolved about the previous attempt. This lets the owner evaluate a concrete proposal without reading raw traces or assuming that every failure happened before dispatch.

The application must also support refusal. If the owner declines a broader destination set, the task should remain paused or fail with an understandable reason. A system that keeps asking until permission expands is not preserving the original boundary. Its recovery policy is pressuring the user to weaken it.

Verify the output after recovery

Passing admission only establishes that a route may attempt the task. It does not establish that the resulting summary is correct. The same verification requirements should apply after fallback as after the preferred route. Otherwise the system becomes least careful about quality when it is already operating outside its normal path.

For this example, a verifier would check the required fields, distinguish source facts from interpretation and identify unsupported incident claims. A valid output structure is necessary but insufficient. A well-formed summary that invents a cause still fails the factual requirement. A second provider's fluent response should not inherit success from the first provider's intended behavior.

Verification records should identify the actual route and attempt that produced the checked artifact. If a reviewer edits the summary, the record should identify the reviewed revision. The final result must not accidentally point to an earlier unchecked response. Recovery often creates several plausible artifacts, which makes precise linkage more important than a single green status badge.

A partial output can be useful if its limits are explicit. The owner may accept a draft with an unresolved field for further investigation. That acceptance should not be counted as a fully verified outcome unless the task definition allows it. Reporting partial completion honestly makes reliability measurements more informative.

Test the refusal paths

A synthetic admission fixture can test the three invented routes without calling a model. It should admit A when its evidence is current, reject B for a required input capability and reject C for destination policy. Making A unavailable should produce an empty eligible set. The local companion check exercises these decisions as illustrations, not as certification of a deployed gateway.

The state fixtures should cover a failure before dispatch, an acknowledged completion and an unknown result after dispatch. Only the first supports a simple assertion that this attempt caused no remote effect. An unknown result must remain unknown until a reconciliation record resolves it. The fixture should not erase the previous attempt when a new one is proposed.

Other useful tests revoke the owner's fallback permission between attempts, change the input revision or expire the capability evidence. Each should prevent automatic continuation under the old assumptions. A test that always reaches a successful alternative demonstrates availability in a narrow scenario; it does not demonstrate that the policy protects boundaries when no safe alternative exists.

Live integration testing would add provider-specific behavior, cancellation handling and realistic failure injection in a suitable test environment. Those checks are not completed by this article. The distinction matters: a small local fixture can validate the reasoning in a worked example while leaving actual network behavior and production reliability unverified.

Measure the consequences of recovery

A routing report should distinguish successful first attempts, successful recoveries, unresolved outcomes and tasks refused by policy. Those categories describe different operational burdens. Combining them into a single success percentage can hide a growing dependence on manual reconciliation or increasingly permissive fallback settings.

Record the costs and latency of the whole task, including abandoned attempts and verification. Also record where a fallback changed an approved contract after an explicit owner decision. That change may be useful, but it should not be reported as identical service under the original requirements. Comparison requires a stable definition of what counts as a satisfactory outcome.

Small examples should not become performance claims. The synthetic routes in this article establish no provider ranking, savings rate or production availability. They expose decision points that a measured evaluation would need to include. Real claims would require controlled task cohorts, documented policies and evidence that each counted success met the applicable contract.

Keep a recovery receipt readable

A concise recovery receipt would identify the original task, the failed attempt, the proposed route and the admission result. It would also state whether the previous outcome is known. That last field changes the interpretation of every subsequent action: a new request may be a straightforward retry or an additional attempt alongside an unresolved earlier one.

For the synthetic incident summary, a refused receipt would say that no available route meets both the input and destination requirements. It would not describe the result as a generic provider outage. The owner could then choose between waiting and authorizing a genuinely different task.

The receipt should reference evidence rather than imply that its own existence certifies correctness. A record can faithfully describe a bad decision. Reviewers still need to inspect the contract, route evidence and verification result. Readability makes that investigation easier; it does not replace it.

Keep a safe ending available

A recovery system needs a legitimate ending in which no further request is sent. Without that ending, every failure eventually becomes a reason to loosen a constraint. The resulting system may look resilient while gradually moving beyond the owner's permissions and the task's intended meaning.

The useful question is not simply whether another route can answer. It is whether that route can perform this task under current authorization, and whether the previous attempt leaves uncertainty that must be resolved first. That question connects routing, job state, approval and verification into one recovery decision.

Write the failure record before selecting another route. It should identify the current contract, the previous dispatch state and the authority needed for the next effect. If those cannot be established, stopping is an accurate recovery result. Related methods are discussed in the writing and research collections.

Sources

Back to blog