TuringCorp

DeciderExamples › Research

Recorded run · Research · 2026-09-15

We submit next week. The effect showed up in the main run and in two of the three repeats. Which abstract sentence do we send?

Two candidate answers were submitted. Decider picked one and reported how close the call was. The reasoning below is the run's own output, published verbatim.

Option A

The effect appeared in the main run and in two of the three repeats; we note that it may be sensitive to seed choice and needs more repetitions.

Option B

The effect appeared in the main run and in two of the three repeats; the third repeat and the sensitivity analysis are reported in Section 4.

SUGGESTION: A
Confidence: 86.0% · recorded 2026-09-15

Option A is the stronger abstract sentence because abstracts must stand alone and transparently characterize the key result. Stating that the effect appeared in the main run and in two of three repeats, then explicitly noting possible seed sensitivity and the need for more repetitions, gives a calibrated, honest summary of a partial replication. Option B shares the first clause but replaces the missing replication detail with a cross-reference to Section 4; that conveys no information in the abstract, leaves abstract-only readers unable to tell whether the third repeat and sensitivity analysis were supportive, and risks misleading by omission—especially because the third repeat is the one that did not show the effect. A section reference is also generally discouraged in abstracts. The counterargument that Option B is more concise and confident is not sufficient: the conciseness gain is small given the identical first clause, and the confident tone is achieved by withholding relevant limitation information rather than by stronger evidence. A single hedged caveat is standard reproducibility practice and remains defensible even if further analysis is favorable, while Option B is potentially misleading under the actual unfavorable replication pattern and may appear evasive to reviewers who will see Section 4 anyway. Therefore Option A is preferred for scientific integrity, self-containedness, and accurate reporting.

How to read this

The confidence value is calibrated against outcomes, not a vibe: on JudgeBench, judgements Decider rates 90%+ were correct 99.6% of the time, and the 80–90% band 94.0%. The lowest band flags near-ties where either choice is defensible — an answer, not a failure. Full tables: the product page.

This page is a recorded run published as a worked example. It is not a live call, and nothing here is charged. Raw judgements behind our numbers are published at api.turingcorp.net.

Try your own two options All Research examples

Prefer the other side? Option B is right there above — the point of a calibrated reading is that you can tell a decisive call from a coin flip.