Faros and the evidence gates between model access and model development

A model-development roadmap can describe a sensible future without establishing that the future has arrived. Integration, routing, evaluation, adaptation and training may appear in one diagram, but they represent different changes to a system. Each needs a different kind of evidence before it can support a capability claim.

Faros is useful for examining this distinction because its public material describes both current integration and gated development work. The engineering question is how a gate should turn uncertainty into a decision: continue with the existing approach, run a better experiment, or promote a specific candidate under a bounded contract.

This article analyzes company-authored public material and proposes an original gate-review worksheet. It does not establish my personal contribution to Faros, independently test its runtime, or report proprietary model training. All numerical examples are synthetic. No invented example below is a Faros benchmark or evidence that a development gate has been passed.

State the present capability before describing the roadmap

The Faros page describes external model integration and rules-based routing, and says Ethen does not currently train its own models. It distinguishes workstreams that are being designed or evaluated from future adaptation and foundation-model training, and reports no validated workstream. These are attributed company statements at the time of checking, not independently reproduced implementation facts.

The company’s learned-routing protocol is explicitly a planned experiment. It proposes a cost improvement threshold and a quality non-inferiority boundary, with additional constraints. The publication does not report an achieved improvement. A reader should carry that evidence status into every summary of the program.

The important distinction is between a gate’s existence and its attainment. A published criterion can make a research direction more inspectable. It cannot establish that the necessary data were collected, the comparison was executed correctly, or the criterion was satisfied. Those facts require a run record and a justified interpretation.

This also makes status wording technically meaningful. Designing a harness differs from evaluating a candidate with that harness. Evaluating differs from validating a claim within a defined scope. A future capability should not inherit the status of the most mature component in the surrounding diagram.

The company’s model-admission essay addresses whether a model belongs in a supported product path. The question here is a research promotion decision: whether a changed policy or capability has sufficient comparative evidence to replace an existing approach. These decisions can inform each other, but qualifying access does not establish an experimental improvement, and a favorable experiment does not automatically qualify every deployment context.

Identify exactly what the proposed change changes

Before evaluating a candidate, describe the intervention. Does it change which external model receives a request? Does it change context preparation, tool selection or verification? Does it update learned parameters? Does it introduce a new dataset or a different acceptance rule?

These questions matter because a measured improvement can have several causes. A route may look cheaper because it receives easier tasks. A verifier may accept more outputs because its rule became less strict. A cache may reduce billed requests without improving selection. A provider price change may reduce cost while the system’s decisions stay identical.

For a gate review, I would record the intervention as a small versioned object: candidate identity, modified components, unchanged comparators, task population, acceptance rule and cost boundary. This is a proposed method. It is not a description of an internal Faros implementation.

Table: original distinction between development activities and the evidence they would need.
Activity Change being evaluated Evidence needed for a bounded claim
Model integration A supported external access path Compatible requests, responses and failure handling
Rules routing A deterministic selection policy Versioned rules and task-level decisions
Learned routing A policy fitted from data Training provenance, comparison design and out-of-sample outcomes
Specialist adaptation Learned parameters change for a task Rights-cleared data, model identity and sealed evaluation
Foundation-model training A model is trained at foundational scale Data rights, training records, compute accounting and broad evaluation

A product label can cover several activities, but an evidence claim should name the one actually tested. Otherwise readers cannot tell whether an improvement belongs to model capability, system integration or a revised accounting rule.

Use a comparator worth keeping

A weak baseline makes an ambitious candidate easier to promote. The useful question is whether the candidate improves on an approach a competent team would actually use. For routing, that might mean rules that already account for task class, compatibility, cost and permitted data handling.

The comparator should receive reasonable tuning on development data. The candidate should receive its own declared development budget. Both should then face evaluation data that were not repeatedly inspected while selecting them. The point is to compare approaches under an honest process, rather than compare a carefully tuned candidate with an intentionally neglected alternative.

There is no universal strongest baseline. A support workflow with deterministic state checks may need a different comparator from open-ended writing. The review should state why the chosen baseline represents the relevant decision, and how sensitive the conclusion might be to another plausible comparator.

It is also useful to retain simple controls. A fixed model path can reveal whether a complicated router adds value. A cache-controlled comparison can reveal whether accounting gains depend on cache placement. Those controls answer narrower questions; they should not be allowed to replace the primary comparison after results become inconvenient.

A synthetic worksheet that separates cost from quality

Imagine two hypothetical configurations evaluated on 1,000 tasks each. Rules produce 800 accepted outcomes at a total measured-boundary cost of USD 400. A candidate produces 790 at USD 320. These are invented inputs for arithmetic, not observations from Faros or any deployed product.

Define cost per verified outcome as all included cost divided by accepted outcomes. Under that definition, rules cost USD 0.50 per accepted outcome. The candidate costs approximately USD 0.405. Its point-estimate cost reduction is about 19.0%, while its accepted-outcome rate is one percentage point lower.

Table: synthetic gate arithmetic; acceptance counts do not establish verifier accuracy.
Quantity Rules Candidate
Evaluated tasks 1,000 1,000
Accepted outcomes 800 790
Included cost, USD 400 320
Cost per accepted outcome, USD 0.5000 0.4051
Accepted-outcome rate 80.0% 79.0%

This looks attractive if the only criterion is the point estimate of cost. It does not settle a quality gate that uses uncertainty. A one-percentage-point observed difference can coexist with a confidence bound that permits a larger loss. The data design and analysis rule determine whether the uncertainty is tolerable.

For illustration only, treat the two samples as independent Bernoulli trials and use a simple normal approximation. The standard error of the difference is approximately 1.81 percentage points. Subtracting 1.645 standard errors from the observed minus-one-point difference gives a one-sided 95% lower bound near minus 3.97 percentage points.

That illustrative bound would not establish non-inferiority against a minus-two-point margin. This is not the analysis appropriate to every routing study: paired tasks, clustering, repeated attempts and adaptive selection can require different estimators. The worked example demonstrates why a favorable point estimate cannot stand in for the predeclared uncertainty calculation.

The company protocol proposes at least 15% lower cost per verified outcome and a one-sided 95% lower quality-difference bound above minus two percentage points. In the synthetic independent-sample example, the cost point estimate clears the first comparison, but the illustrative quality bound fails the second. No company experiment is being graded by this invented table.

Decide what the verifier is allowed to accept

A verified outcome is only as well defined as its acceptance process. A deterministic check may be strong for a precisely specified output but weak for an incomplete specification. A human rubric may capture useful qualities while introducing variation. A model judge may be convenient while sharing biases with the producer.

The gate should therefore include a verifier review. Which failure classes are known? How are ambiguous cases handled? Is a sample independently adjudicated? Can the producer influence the acceptance rule or hide the evidence that would cause rejection? Those questions apply before cost ratios become meaningful.

Suppose a candidate produces a patch that passes tests but omits a requested migration note. If the original task requires the note, the omission belongs in the outcome definition. Redefining success as tests pass would make the candidate look better by changing the task, not necessarily by improving the system.

For the proposed worksheet, acceptance disagreements would remain visible in a separate field. Unresolved cases would not be silently converted into successes. Their treatment would be fixed before comparing candidates, and a sensitivity analysis would explain whether plausible alternative adjudications change the decision.

Logs should support the claim, not merely exist

A routing study needs to know what choices were available when a decision was made. The selected model alone cannot establish that another option was eligible, cheaper or allowed to receive the data. Candidate availability, compatibility constraints and policy exclusions affect the comparison.

For a learned policy, the record may also need the probability with which a configuration was selected, depending on the evaluation method. Without appropriate logging and assumptions, historical outcomes from selectively chosen routes cannot be treated as if every option had been assigned randomly. A retrospective spreadsheet does not remove selection bias.

The proposed record should include task class, eligible choices, selected configuration, policy version, context transformation, verifier version, outcome, cost and missingness reason. Sensitive content can remain protected; the public claim can describe the schema and method without exposing private prompts or credentials.

Missing records deserve their own analysis. If the most expensive failures disappear from the ledger, cost per accepted outcome becomes artificially favorable. If timeouts are counted differently between arms, reliability comparisons may be misleading. The review should reconcile task counts from enrollment through final disposition.

Evaluate a model change as a new uncertainty

An external model may change while the surrounding product name remains familiar. Even a pinned local artifact may be replaced by a new version. A routing policy fitted to the previous behavior may select differently or lose the advantage that justified its use.

For this proposed process, a model change would trigger a bounded assurance review. The team would rerun representative tasks, inspect relevant failure categories, compare verifier behavior and assess whether the original eligibility assumptions still hold. A change might be safe for one task class and unsuitable for another.

The model-card documentation helps identify intended use and limitations, but a card cannot establish performance inside an application’s current tools and context pipeline. The application still needs evidence for its own configuration. Documentation and application testing should complement each other.

Rollback also needs a meaningful contract. Returning to a previous model is only possible if access, compatibility and authorization still exist. A release process should retain the old policy’s identity and its dependencies, then verify that the alternative is usable rather than treating rollback as a label in a dashboard.

Data rights are a gate, not a footnote

A promising dataset cannot be assumed available for every downstream use. Evaluation access, redistribution, model training and publication may have different permissions. A dataset’s presence in a repository does not establish that all of those activities are authorized.

The proposed review should trace each data source to its allowed uses and restrictions. Derived labels need provenance too, especially when generated from confidential material or licensed outputs. Removing obvious identifiers may be useful, but it does not automatically resolve every contractual or disclosure constraint.

This has a direct effect on research design. A sealed test set must be usable by the evaluation process, while its contents may need to remain undisclosed. A public methods description can explain selection, exclusions and adjudication without publishing the underlying private examples. The claim should state what outside reviewers can and cannot reproduce.

Similarly, a personal engineering article needs specific contribution evidence and item-level disclosure permission. Private plans cannot become proof of completed implementation merely because they are detailed. A responsible account distinguishes the intended work, the performed work and the material cleared for readers.

Write the decision record before opening the results

For the proposed gate worksheet, the decision record should explain what happens in three possible outcomes: clear success, clear failure and an inconclusive result. A study designed only to recognize success leaves too much freedom to reinterpret an inconvenient observation after the fact.

Clear failure might mean a safety constraint is violated even when cost improves. An inconclusive result might mean the quality interval remains too wide to support promotion. These outcomes are different. In the synthetic worksheet, the uncertain quality bound does not establish that the candidate is worse by four points; it establishes that the illustrative analysis has not ruled out an unacceptable loss.

That distinction changes the next action. A safety failure may require redesign. An inconclusive quality comparison may justify a better-powered study, provided the additional data collection follows a declared plan. Quietly repeating the experiment until a favorable interval appears would create a different evidence process from the one readers were promised.

The record should also distinguish a development pilot from a confirmatory evaluation. Pilot data can help find defects, estimate variability and refine the task definition. Once those data guide revisions, they should not be described as a sealed independent test of the final candidate. A separate evaluation boundary is needed for that claim.

For a practical review meeting, I would display the original gate, complete outcome counts, missing-record reconciliation, cost boundary and relevant uncertainty calculation together. Then reviewers can see whether a proposed promotion follows the declared rule or requests an explicit amendment. An amendment may be reasonable, but it should not retroactively turn the old criterion into a passed gate.

The discussion should retain negative evidence. If one task class improves while another regresses, a narrower rollout may be justified, but a population-wide improvement claim would be inappropriate. If latency tails worsen while average latency improves, the user’s actual deadline may decide which result matters.

Finally, define what new observation would reverse the promotion. Monitoring should preserve the relevant outcome and accounting definitions so the post-release evidence can be compared meaningfully. A gate is useful when it creates both an entry condition and a reasoned path back to the comparator. Otherwise it risks becoming a ceremonial checkpoint that the system never revisits.

Promotion should authorize a bounded use

Even a passed experiment would not establish unlimited deployment readiness. Its conclusion would apply to a task population, a configuration and an outcome definition. A release decision would still need operational controls, incident handling, monitoring and a plan for conditions outside the evaluated scope.

A useful promotion statement might authorize one candidate for one task class under specified limits, while retaining the comparator for unsupported or uncertain cases. That is more precise than declaring the whole program validated. It also makes future regressions easier to investigate because the approved boundary is explicit.

For Faros, the public distinction between current integration and gated development provides a starting point for those questions. It does not fill the absent independent run record. The next meaningful evidence would connect a particular candidate to a rights-cleared evaluation, a credible comparator, a reviewed verifier and complete outcomes.

Readers can find related methodology through Research and the essay collection in Writing. The decision worth preserving is the gate itself: what observation would justify changing the system, and what observation would require keeping the existing approach? A roadmap becomes technically useful when both answers remain possible.

Back to blog