TuringCorp

Available on Poe

Pick the better answer.
Get a confidence score you can act on.

Decider is an arbitration layer for AI workflows. Give it a task and two candidate answers; it returns the stronger one, a calibrated confidence level and a short reason — so you can auto-approve, escalate, or ask again.

Open Decider on Poe See pricing

List price $0.50 per decision · Launch offer $0.25 for the first month

Not ready to pay? Open the app and try a free recorded run first — pick a domain, read a full decision end to end, nothing is charged.

Or read them here first: 27 recorded decisions — real questions, both candidate answers, the pick, the confidence and the full reasoning.

Confidently wrong is the real bug

Single-shot generation rarely signals when it is unsure. In pipelines, a wrong-but-confident answer costs far more than a slow one.

Choosing is a different skill

Pick the best of two drafts sounds easy and is not. It rewards judgement, not fluency — which is exactly what we built for.

A verdict you can automate on

A single decisive field plus a confidence tier lets you route work: auto-accept, review, or re-run. No parsing of prose.

How it works

  1. Describe the task. Paste the question or context (a few sentences is enough).
  2. Paste both candidates. Option A and Option B — drafts, model outputs, vendor proposals.
  3. Read the verdict. betterOption, a confidence percentage, and a concise reason.
{"task": "…the question and background…",
 "option_A": "…candidate A…",
 "option_B": "…candidate B…"}

{"betterOption": "option_A",
 "confidence": "91.7%",
 "reason": "…short justification…"}

Every call is one decision over one A/B pair. Most decisions come back in well under a minute.

Decider — arbitration & quality gate

A second opinion your agents can act on. One structured verdict per call.

Best for

  • Choosing between two drafts, plans or vendor proposals
  • Final QA before a deliverable ships
  • A confidence gate inside an agent pipeline
  • Triaging factual-risk claims before publication

What you get

  • A forced verdict — never "it depends"
  • A calibrated confidence tier (monotonic)
  • A concise reason suitable for logs and audit trails
  • Consistent structured output for automation

Evaluated results

Self-run on the official protocols; failures disclosed. References are published scores under the same protocol, equal-weighted; our numbers are our own runs and are reported here without vendor or architecture detail.
BenchmarkDeciderPublished references
JudgeBench 620 judgements, 6 failures disclosed ≈92.5%
ContextualJudgeBench 2,000 pairs, all 8 splits 67.1% vs 65.4% same-set direct baseline o1 55.3 · o3-mini 52.6 · R1 51.9

The part that is actually different: the confidence signal

On JudgeBench the raw pick rate ties the direct baseline — roughly 92.5% against 92.7%. What you cannot get from a single pass is how close the call was. Every judgment below carries a calibrated value, measured against outcomes:

Confidence reportedShare of judgmentsObserved accuracyWhat it means
≥ 90%45.6%99.6%one candidate clearly stronger
80–90%29.7%94.0%one candidate stronger
70–80%13.2%84.1%closer call
< 70%10.5%67.7%near-tie — either choice is defensible
JudgeBench, 614 judgements. Use the band you are comfortable with as your own threshold; for high-stakes or irreversible decisions, apply your own review policy.

A reading, not a verdict you have to obey

You usually already lean one way. What you cannot do on your own is price your own uncertainty — that is the part Decider adds, and both outcomes are useful:

  • Decider is confident too. Commit to that pick and stop re-reading both options.
  • Decider comes back low. That is not a failure — it is the answer you actually needed: the two are genuinely close, so take the one you already preferred and spend your time somewhere else.

A high band means the comparison was decisive. A low band means the decision is yours to make, on grounds outside the two answers.

We publish what we measure, including the runs that fail, the tasks that did not complete, and the splits where the baseline wins. We do not claim perfection or make absolute guarantees. Every number here, with the raw judgments behind it, is published at api.turingcorp.net.

Pricing

Simple per-call pricing. Charged in Poe compute points; shown here in USD.

ProductList priceLaunch offerUnit
Decider$0.50$0.25per decision
Team coming soon$2.50$0.99per task
Launch offer applies to the first month. One decision = one A/B pair; one task = one complete deliverable.

Because each bot is priced separately, there is no tiered pricing to decode: you pay per call, and that is the whole model.

FAQ

Is Decider a chatbot?
No. It does not write essays or brainstorm with you. It judges: given a task and two candidates, it returns the better one, a confidence level and a short reason.
Why not just ask a frontier model to pick?
You can — and you will get an opinion. Decider is built specifically for the judgement step and returns a structured, decision-ready verdict with a confidence tier, so the result can drive automation rather than a conversation.
When should I not use it?
When the two options differ in cost or strategy rather than quality — that is a business decision. Decider is for questions with a defensible better answer.
What happens if the input is too thin to judge?
Expect a lower confidence tier. A low tier is a signal: gather more context, or route the decision to a human.
Where do the numbers come from?
From our own runs on the official benchmark protocols, with failures and exclusions disclosed. Reference scores are the published numbers under the same protocol, equal-weighted. Full tables and raw judgments live at api.turingcorp.net.
Decider ties the baseline on accuracy — why would I pay for it?
Because you are not buying a better answer, you are buying a reading of your own uncertainty — and both outcomes save you something. Judgments Decider rates ≥90% were correct 99.6% of the time on JudgeBench, so a high band is permission to commit and move on. A low band is equally useful: it tells you the two options really are close, so take the one you already preferred instead of re-reading both again.
How is Team different from asking a bigger model?
Measured criterion by criterion against ProfBench's official rubric on expert-level tasks, Team scores 63.3 against 57.4 for a direct model baseline and 55.6 for the official o3 draft — same 38 tasks, same judging pipeline. The result we would rather be judged on: in a side-by-side comparison where our draft and the official o3 draft were shown together, the pick went to ours 36 times out of 38, and swapping the two positions changed nothing (10/10).
Is the API the same thing?
Same capability, two channels. On Poe these run as apps on the Poe platform: public, self-serve, pay per call, no separate approval. The bare API at api.turingcorp.net is a small-scale preview by invitation, for people calling the models directly (iAsk@turingcorp.net).
How is this billed?
Through Poe, in compute points, at the per-call price above. No subscription, no minimum.