What would disprove this AI research claim?

An AI research report can look convincing before it contains a testable claim. It may describe an elegant architecture, cite respected papers and show a plausible example. None of those elements establishes that the proposed system improves an outcome. The missing question is often the most useful one: what observation would make the author withdraw the claim?

That question turns a narrative into a research plan. It forces the author to specify a comparison, define success and expose conditions under which the proposed improvement disappears. It also gives readers a way to distinguish a design proposal from measured evidence without dismissing the proposal's engineering value.

This article develops an original review exercise around a synthetic routing study. The proposed claim is that a new selection policy improves verified task completion relative to a fixed baseline without violating task constraints. The numbers are invented to test reasoning and arithmetic. No models were called, no production traces were inspected and no Ethen benchmark was reproduced. The exercise demonstrates what an accountable claim would require.

Turn the headline into a decision rule

“The router is better” leaves too much undefined. Better might mean cheaper, faster, more accurate or more likely to return any answer. Those objectives can disagree. A policy that produces more responses by ignoring a required input format is not necessarily more successful at the original task.

A testable claim identifies the task population, comparison policy, outcome definition and important constraints. For this exercise, each task requires an output that satisfies a stated structure and a factual check against its supplied evidence. A task also has a permitted destination set. Completion outside that set does not qualify as a success, regardless of answer quality.

The baseline policy and candidate policy must be fixed before their results are compared. Otherwise the author can repeatedly adjust the candidate using the evaluation tasks and report the final comparison as though those tasks were untouched. That would measure adaptation to the evaluation set as well as any general improvement in routing.

A decision rule should also describe acceptable tradeoffs. If a study permits a small decrease in completion in exchange for lower cost, that tolerance should be explicit before inspecting outcomes. The author should explain why the tolerance is meaningful for the intended users. An arbitrary threshold chosen after seeing the results makes the claim harder to evaluate.

Label the evidence stage

A proposal states what might be worth building. A protocol states how a future comparison would be conducted. An experiment reports what actually happened under that protocol. A synthesis interprets existing research. These forms can all be valuable, but they answer different questions and should not borrow one another's authority.

The Ethen routing protocol is relevant public context because it presents a planned research comparison. Its existence is not a result demonstrating that a learned router improves completion or cost. I cite it as a protocol, and this article does not claim to have implemented or run it.

The Center for Open Science's preregistration guidance explains the role of specifying a plan before observing results. A local protocol file saved before a toy calculation is a useful demonstration of sequence, but it is not an independently registered empirical study. That distinction matters when an author uses the language of preregistration to describe an internal note.

A reader should be able to identify the evidence stage from the introduction and methods, rather than discovering it at the end. “No measured results” is useful information. It prevents a conceptual diagram, a local arithmetic check or a synthetic trace from being interpreted as an operational performance claim.

Define the unit being compared

The unit in the exercise is a logical task, not a successful provider response. A task may include several attempts, verification work and a final refusal. Counting responses as the unit can make a retry-heavy policy appear more productive even if it completes no additional tasks.

Each task receives a stable identifier and a recorded input revision. Both policies are evaluated against the same synthetic task definitions. This paired structure makes the example easy to inspect: a difference in completion can be traced to a specific task instead of being hidden inside two totals from unrelated populations.

A real paired study would still need to address order effects and environmental change. If one policy runs when a provider is healthy and the other runs during an outage, the comparison includes availability differences. If the first run changes shared state, the second may inherit a different task. Pairing task labels alone does not remove those effects.

The experiment would therefore document scheduling, isolation and any shared caches or external tools. It would also record which attempted tasks were excluded and why. Excluding difficult tasks from only one policy can change the question being answered. The population described in the headline should match the population that actually contributes to the result.

Use a verifier that can reject plausible prose

The success definition in this exercise includes structure and factual support. A response with all required fields can still contain an unsupported cause or an incorrect comparison. Conversely, a cautious response that identifies a missing source may be more useful than a complete-looking answer that fills the gap with speculation.

A verifier needs a documented rubric and access to the relevant evidence. It should distinguish unsupported statements, contradictions and interpretive judgments. Where human review is used, reviewers need enough context to apply the rubric consistently. Where an automated checker is used, the author should describe its known limits rather than treating a machine-generated score as self-validating.

The Ethen article on research-report evidence explicitly discusses a local lexical check and its limitations. That is an attributed description of a particular checking approach. This essay's review exercise goes further in a different direction: it asks whether the research claim itself has a predefined condition for failure.

For a real comparison, the verifier should be versioned and protected against learning which policy produced the artifact where feasible. Otherwise expectations about the candidate can influence judgment. Disagreements should be retained, with an escalation rule for required claims. The goal is an inspectable outcome definition, not a ritual in which every plausible answer eventually receives approval.

Inspect the paired outcomes

The synthetic cohort contains twenty tasks. The baseline completes sixteen. The candidate completes fifteen. Those totals immediately contradict a simple claim that the candidate improves the observed completion count in this cohort. They do not establish a general population ranking, because the cohort is invented and no empirical sampling occurred.

Looking at task pairs gives a more informative account. Both policies complete twelve tasks. The baseline alone completes four. The candidate alone completes three. Both fail one. These categories sum to twenty tasks, and the completion totals follow directly: twelve plus four for the baseline, twelve plus three for the candidate.

The candidate wins on three tasks but loses on four. Its total is not rescued by emphasizing only the wins. A report that presents three success stories without the corresponding losses would give the reader a distorted account of the same fixture. The arithmetic in the table checks the category totals and the difference of one completed task.

This fixture is intentionally negative for the improvement claim. It demonstrates a review habit, not a statistical finding: include evidence that would disappoint the hypothesis. If a real study produced an unfavorable result, the author should report it under the planned analysis rather than redefining success until the original headline becomes defensible.

All twenty synthetic task pairs

These task identifiers and outcomes are invented to demonstrate a negative case. They were constructed for this essay, not sampled from model executions. “Pass” means the fictional task meets the fixed verification rule; “fail” means it does not. Neither label certifies a real model.

Table 1: Synthetic worked example
Synthetic task Baseline Candidate
T01 Pass Pass
T02 Pass Pass
T03 Pass Pass
T04 Pass Pass
T05 Pass Pass
T06 Pass Pass
T07 Pass Pass
T08 Pass Pass
T09 Pass Pass
T10 Pass Pass
T11 Pass Pass
T12 Pass Pass
T13 Pass Fail
T14 Pass Fail
T15 Pass Fail
T16 Pass Fail
T17 Fail Pass
T18 Fail Pass
T19 Fail Pass
T20 Fail Fail
Table 2: Synthetic worked example
Paired category Task count
Both pass 12
Baseline only passes 4
Candidate only passes 3
Both fail 1
Total pairs 20

The baseline passes 12 + 4 = 16 tasks; the candidate passes 12 + 3 = 15. Completion is 80% versus 75%, a five-percentage-point difference in this invented cohort. The candidate's three wins and four losses must both be considered. These counts support no population inference, confidence interval or empirical claim of superiority.

Keep cost and quality connected

Suppose the invented baseline costs forty accounting units across the cohort and the candidate costs thirty-six. The candidate has lower total cost but also one fewer verified completion. Cost per verified completion is forty divided by sixteen, or 2.50 units, for the baseline. It is thirty-six divided by fifteen, or 2.40 units, for the candidate.

Those numbers answer different questions. The candidate spends less overall and less per counted success in the synthetic fixture. It does not complete more tasks. Whether that tradeoff is acceptable depends on the task's requirements and the predeclared decision rule. The author should not turn a favorable cost ratio into evidence of improved completion.

The costs here are arbitrary accounting units, not provider prices. A real study would define which inference, retry, tool, verification and human costs are included. It would also explain whether infrastructure and development costs are part of the comparison. Changing accounting boundaries between policies can manufacture apparent efficiency.

A ratio also needs a defined treatment of zero successes. Division by zero should not become a reassuring zero-cost result. Reporting total cost and no verified completion is clearer than inventing a finite cost-per-success number. The companion arithmetic check covers that boundary so the worked example remains honest about what its metric can express.

Cost does not erase a failed primary claim

Table 3: Synthetic worked example
Synthetic measure Baseline Candidate
Total accounting units 40 36
Verified completions 16 15
Units per verified completion 2.50 2.40

The candidate's cost per completion is 4% lower, while its completion rate is five percentage points lower. These arbitrary units are not dollars, tokens or provider prices. Whether that tradeoff is acceptable requires a decision rule established independently of this convenient example. If the primary claim is more completions on these tasks, lower cost answers a different question.

Separate exploration from confirmation

After inspecting the twenty invented outcomes, an author might notice that the candidate does better on short structured tasks and worse on longer document tasks. That could motivate a new hypothesis. It should not automatically become the main confirmed finding of the original comparison.

Subgroup analysis is useful when its status is clear. A planned subgroup tests an explicitly stated question. An exploratory subgroup generates a question for another study. Both should include unfavorable cases and explain how the subgroup was defined. A category chosen because it gives the candidate a strong result is especially vulnerable to overinterpretation.

The next comparison should use tasks not already used to tune the candidate or choose the subgroup. Otherwise the apparent confirmation includes information learned from the first analysis. Maintaining separate development and evaluation records makes that boundary more visible. It also helps explain why a result on a familiar task collection may fail to transfer.

An article can present this process as original research commentary without claiming an original empirical discovery. The contribution is the reasoning framework and the worked example. Actual discovery would require new observations, documented methods and evidence that supports the stronger claim. Keeping those levels distinct protects both the reader and the author's credibility.

Record the baseline's strength

An improvement against a weak baseline may establish less than the headline suggests. A baseline that routes every task identically, ignores known capabilities or never retries can be easy to beat. The author should explain why the comparison represents a meaningful alternative for the intended setting.

A useful baseline might apply explicit eligibility requirements and a fixed ordering among admitted routes. Its policy should be documented well enough for another reviewer to understand the decisions. Calling it “rules-based” is insufficient when the actual rules are unspecified or deliberately omit information available to the candidate.

The candidate must also receive comparable inputs unless the study explicitly evaluates additional information. If it sees richer capability evidence, the observed difference may arise from the evidence rather than the selection method. A factorial design could separate those contributions in a real experiment, but this synthetic exercise does not perform one.

Baseline maintenance matters over time. Provider availability and model versions can change, so a stale comparison may answer a historical question rather than today's deployment decision. A published report should identify the versions and evaluation window. It should avoid implying that one completed study settles every future routing choice.

Make missing data part of the result

Some tasks may remain pending at the reporting cutoff. Others may be canceled, fail verification or lack a complete cost record. Those states should be visible. Treating all missing outcomes as failures or silently dropping them can each distort the interpretation, depending on why the data are missing.

The report should state the cutoff and provide separate counts for completed, failed, pending and canceled tasks. It should explain how each category enters the primary analysis. If unresolved provider effects exist, they belong in the evidence record. A task with unknown completion should not be assigned a confident outcome merely to finish the table.

Missing cost data require similar care. An incomplete retry log can make a policy appear inexpensive. A report can provide a sensitivity analysis showing how plausible missing costs affect the conclusion, but it must label the assumptions. Estimated values should not be mixed with observed values without a way for the reader to distinguish them.

The synthetic cohort is complete by construction, so it has no missing records. That convenience is a limitation. The example checks whether the stated arithmetic follows from its fixture; it does not demonstrate that a real logging system would collect every necessary record or that a deployed verifier would classify them correctly.

A claim ledger makes review concrete

A claim ledger should identify the statement, its evidence, its status and the action needed before publication. For this article, the paired counts are synthetic and mechanically checked. The general research recommendations are proposed methods. The statements about public Ethen materials are attributed to those pages, rather than independently verified implementation facts.

An unresolved claim should remain explicitly unresolved in a research report. A source that repeats the same assertion does not necessarily add independent support. A company publication can be authoritative about what the company publicly says while remaining insufficient evidence of production behavior. Personal contribution claims require additional evidence connecting a named person to the relevant work.

The ledger also helps prevent attribution drift during editing. A sentence that begins as “the publisher describes” can become “the system guarantees” when compressed. That change increases the evidentiary burden. Comparing the final prose with its ledger catches such changes before a polished article accidentally promises more than its sources support.

A report should identify what its reviewers actually checked. Arithmetic and consistency checks establish whether a calculation follows its inputs. They do not constitute independent scientific peer review or verify that the inputs are observations. Accurate review labels make this distinction visible.

Give reviewers the unfavorable table

A review package should make the full paired outcome table available, including losses and joint failures. A reader who sees only selected examples cannot reconstruct the comparison. For the twenty-task fixture, publishing the four categories is enough to reproduce the completion totals, while task-level records would allow more detailed inspection.

The package should also state which questions remain unanswered. Here, there is no measured latency, no real provider cost and no empirical estimate of future performance. Adding a polished chart would not supply those missing observations. The absence of those measurements limits the claim even when the arithmetic is correct.

A reviewer can still judge whether the proposed method is coherent and whether the evidence labels are accurate. That is a worthwhile review of a design essay. It should be recorded as such, with empirical conclusions reserved for a study that actually collects the necessary observations.

Publish the question with the answer

A strong research article tells readers what would count against its conclusion. It records a meaningful baseline, a fixed outcome definition and the limitations of its evidence. When the result is negative, it reports the negative result. When the work is a proposal, it explains the proposed test without borrowing the language of completed experiments.

Start with the claim, name the observation that would defeat it, and show that observation alongside the favorable cases. In this example, fifteen completions cannot establish an improvement over sixteen. The next empirical study would need real tasks, a defensible sample and a fixed analysis plan. Related methodological commentary appears in research and writing.

Sources

Back to blog