Available on Poe
Pick the better answer.
Get a confidence score you can act on.
Decider is an arbitration layer for AI workflows. Give it a task and two candidate answers; it returns the stronger one, a calibrated confidence level and a short reason — so you can auto-approve, escalate, or ask again.
Open Decider on Poe See pricing
List price $0.50 per decision · Launch offer $0.25 for the first month
Not ready to pay? Open the app and try a free recorded run first — pick a domain, read a full decision end to end, nothing is charged.
Or read them here first: 27 recorded decisions — real questions, both candidate answers, the pick, the confidence and the full reasoning.
Confidently wrong is the real bug
Single-shot generation rarely signals when it is unsure. In pipelines, a wrong-but-confident answer costs far more than a slow one.
Choosing is a different skill
Pick the best of two drafts sounds easy and is not. It rewards judgement, not fluency — which is exactly what we built for.
A verdict you can automate on
A single decisive field plus a confidence tier lets you route work: auto-accept, review, or re-run. No parsing of prose.
How it works
- Describe the task. Paste the question or context (a few sentences is enough).
- Paste both candidates. Option A and Option B — drafts, model outputs, vendor proposals.
- Read the verdict.
betterOption, a confidence percentage, and a concise reason.
{"task": "…the question and background…",
"option_A": "…candidate A…",
"option_B": "…candidate B…"}
{"betterOption": "option_A",
"confidence": "91.7%",
"reason": "…short justification…"}
Every call is one decision over one A/B pair. Most decisions come back in well under a minute.
Decider — arbitration & quality gate
A second opinion your agents can act on. One structured verdict per call.
Best for
- Choosing between two drafts, plans or vendor proposals
- Final QA before a deliverable ships
- A confidence gate inside an agent pipeline
- Triaging factual-risk claims before publication
What you get
- A forced verdict — never "it depends"
- A calibrated confidence tier (monotonic)
- A concise reason suitable for logs and audit trails
- Consistent structured output for automation
Evaluated results
| Benchmark | Decider | Published references |
|---|---|---|
| JudgeBench 620 judgements, 6 failures disclosed | ≈92.5% | — |
| ContextualJudgeBench 2,000 pairs, all 8 splits | 67.1% vs 65.4% same-set direct baseline | o1 55.3 · o3-mini 52.6 · R1 51.9 |
The part that is actually different: the confidence signal
On JudgeBench the raw pick rate ties the direct baseline — roughly 92.5% against 92.7%. What you cannot get from a single pass is how close the call was. Every judgment below carries a calibrated value, measured against outcomes:
| Confidence reported | Share of judgments | Observed accuracy | What it means |
|---|---|---|---|
| ≥ 90% | 45.6% | 99.6% | one candidate clearly stronger |
| 80–90% | 29.7% | 94.0% | one candidate stronger |
| 70–80% | 13.2% | 84.1% | closer call |
| < 70% | 10.5% | 67.7% | near-tie — either choice is defensible |
A reading, not a verdict you have to obey
You usually already lean one way. What you cannot do on your own is price your own uncertainty — that is the part Decider adds, and both outcomes are useful:
- Decider is confident too. Commit to that pick and stop re-reading both options.
- Decider comes back low. That is not a failure — it is the answer you actually needed: the two are genuinely close, so take the one you already preferred and spend your time somewhere else.
A high band means the comparison was decisive. A low band means the decision is yours to make, on grounds outside the two answers.
We publish what we measure, including the runs that fail, the tasks that did not complete, and the splits where the baseline wins. We do not claim perfection or make absolute guarantees. Every number here, with the raw judgments behind it, is published at api.turingcorp.net.
Team — long-horizon deliverables
For expert work that deserves more than one pass — research, technical analysis, research writing and business decisions — judged on substance, criterion by criterion, rather than on how well it reads.
Coming soon. Team is in final validation on Poe — we are not opening it until the delivery experience matches the work it produces. Pricing at launch: list $2.50 per task · launch offer $0.99.
Where it is strongest
- Scientific and technical analysis (Physics PhD, Chemistry PhD)
- Business decisions (Consulting MBA, Finance MBA)
- Research and multi-section analysis that a single pass leaves thin
What it costs you
- Minutes, not seconds — and longer for the heaviest tasks
- More per task than a single direct call
- What that buys: in a side-by-side comparison it is the preferred draft
Measured on expert-level tasks
| Same judging pipeline, same 38 tasks | Score / 100 |
|---|---|
| Team | 63.3 |
| Direct model baseline | 57.4 |
| Official o3 draft | 55.6 |
| Domain | Team | Direct baseline | Official o3 draft |
|---|---|---|---|
| Chemistry PhD | 75.5 | 69.4 | 60.8 |
| Consulting MBA | 70.8 | 70.3 | 64.9 |
| Finance MBA | 54.6 | 48.2 | 57.5 |
| Physics PhD 8 tasks | 49.6 | 37.8 | 35.2 |
The result that matters: it is the draft that gets picked
Shown the same question and two drafts — ours and the official o3 draft — and nothing else (no rubric, no scores, no knowledge of the measurements below), our own decision kernel picked ours 36 times out of 38, at a median confidence of 85.9%. Swapping the two drafts between option A and option B on a 10-task sample changed nothing: 10/10 kept the same pick. This is supporting evidence, not the primary measurement.
A two-point gap in a headline score means little to most readers. Being the preferred draft in an almost one-sided comparison is the result we would actually act on — and it is the one we would ask you to hold us to.
Why this benchmark
ProfBench scores expert work criterion by criterion against an official rubric, on tasks drawn from Physics and Chemistry PhD work and from Finance and Consulting MBA practice. That makes it a deliberately differentiated benchmark: a system that only writes convincingly cannot produce a differentiated score on it. We chose it because it is the kind of test our approach should either pass or fail visibly — research, technical analysis, research writing and business decisions.
Judge calibration on the official o3 draft: 74.0% agreement with the official labels, F1 0.761, mean difference +2.4 points between our scale and the official one. Full tables, exclusions and raw judgments: api.turingcorp.net.
Pricing
Simple per-call pricing. Charged in Poe compute points; shown here in USD.
| Product | List price | Launch offer | Unit |
|---|---|---|---|
| Decider | $0.50 | $0.25 | per decision |
| Team coming soon | $2.50 | $0.99 | per task |
Because each bot is priced separately, there is no tiered pricing to decode: you pay per call, and that is the whole model.
FAQ
- Is Decider a chatbot?
- No. It does not write essays or brainstorm with you. It judges: given a task and two candidates, it returns the better one, a confidence level and a short reason.
- Why not just ask a frontier model to pick?
- You can — and you will get an opinion. Decider is built specifically for the judgement step and returns a structured, decision-ready verdict with a confidence tier, so the result can drive automation rather than a conversation.
- When should I not use it?
- When the two options differ in cost or strategy rather than quality — that is a business decision. Decider is for questions with a defensible better answer.
- What happens if the input is too thin to judge?
- Expect a lower confidence tier. A low tier is a signal: gather more context, or route the decision to a human.
- Where do the numbers come from?
- From our own runs on the official benchmark protocols, with failures and exclusions disclosed. Reference scores are the published numbers under the same protocol, equal-weighted. Full tables and raw judgments live at api.turingcorp.net.
- Decider ties the baseline on accuracy — why would I pay for it?
- Because you are not buying a better answer, you are buying a reading of your own uncertainty — and both outcomes save you something. Judgments Decider rates ≥90% were correct 99.6% of the time on JudgeBench, so a high band is permission to commit and move on. A low band is equally useful: it tells you the two options really are close, so take the one you already preferred instead of re-reading both again.
- How is Team different from asking a bigger model?
- Measured criterion by criterion against ProfBench's official rubric on expert-level tasks, Team scores 63.3 against 57.4 for a direct model baseline and 55.6 for the official o3 draft — same 38 tasks, same judging pipeline. The result we would rather be judged on: in a side-by-side comparison where our draft and the official o3 draft were shown together, the pick went to ours 36 times out of 38, and swapping the two positions changed nothing (10/10).
- Is the API the same thing?
- Same capability, two channels. On Poe these run as apps on the Poe platform: public, self-serve, pay per call, no separate approval. The bare API at api.turingcorp.net is a small-scale preview by invitation, for people calling the models directly (iAsk@turingcorp.net).
- How is this billed?
- Through Poe, in compute points, at the per-call price above. No subscription, no minimum.