TuringCorp

DeciderExamples › Research

Recorded run · Research · 2026-09-15

A metric moved four percent this quarter and the team reads it two ways. The report needs one headline — which reading do we lead with?

Two candidate answers were submitted. Decider picked one and reported how close the call was. The reasoning below is the run's own output, published verbatim.

Option A

Lead with composition: most of the four percent comes from the mix of new accounts, which score lower in every period on record. Recommend waiting a quarter before we call it a behaviour shift.

Option B

Lead with behaviour: the move is smaller but still visible inside each cohort separately. Recommend investigating now, with the account mix as the first caveat.

SUGGESTION: A
Confidence: 80.0% · recorded 2026-09-15

The single headline should lead with composition: most of the four percent move is driven by account mix, specifically new accounts that score lower in every period on record. That pattern is stable and structural/mechanical, so it is the dominant explanation of the overall change and the safest basis for a report headline. Leading with behaviour would elevate a smaller within-cohort residual above the main driver. The within-cohort movement is visible and worth monitoring or investigating as a follow-up caveat, but it is not strong enough to headline as a behaviour shift. Making behaviour the headline risks the classic Simpson's-paradox misread: treating a mix artifact as a genuine performance change, over-alarming readers, and misdirecting attention or resources toward the wrong intervention. The prudent conclusion is therefore to lead with composition, wait a quarter before declaring a behaviour shift, and keep the smaller cohort-level signal under active review rather than discarding it. This preserves report integrity by ensuring the one headline explains what actually moved while still acknowledging the secondary behavioural evidence.

How to read this

The confidence value is calibrated against outcomes, not a vibe: on JudgeBench, judgements Decider rates 90%+ were correct 99.6% of the time, and the 80–90% band 94.0%. The lowest band flags near-ties where either choice is defensible — an answer, not a failure. Full tables: the product page.

This page is a recorded run published as a worked example. It is not a live call, and nothing here is charged. Raw judgements behind our numbers are published at api.turingcorp.net.

Try your own two options All Research examples

Prefer the other side? Option B is right there above — the point of a calibrated reading is that you can tell a decisive call from a coin flip.