Not a single model. Not a router.
A collaborative inference engine where multiple reasoning
paths cross-examine each other — producing answers of a depth that no
individual model can reach alone.
Our Models
Access is self-serve. Create an account and add credit at
Agent Pass, then send the
Agent Pass it issues as Authorization: Bearer <pass> to
https://api.turingcorp.net/v1. A pass lasts 7 days and can be
re-rolled from that page at any time.
Agents can do the same over the API — see
agent-pass llms.txt or
our agent brief.
Prefer not to write code? Decider is also available as a ready-to-use product on Poe — sign in there and use it directly, no TuringCorp key needed. Decider on Poe — pricing, FAQ, and recorded decisions you can read before you pay.
Benchmark Note: Independent objective evaluation across open benchmarks. Not a third-party leaderboard. Self-run with official protocols; details per product below.
ProfBench — 40 expert-level tasks, ten each in Physics PhD, Chemistry PhD, Finance MBA and Consulting MBA, scored criterion by criterion with the official rubric.
Self-run with the official rubric protocol: per-criterion Yes/No with the official prompt verbatim, temperature 0 / top_p 1, no output cap. Team delivery is content + reason. 38 of 40 tasks scored: two tasks were excluded because no answer was produced for them during the run (disclosed rather than imputed). The DeepSeek V4 Flash baseline and the official o3 draft were read by the same judging pipeline, and all three columns are reported on the 38 tasks every model completed.
| Model | Team | DeepSeek V4 Flash (direct) | official o3 draft |
|---|---|---|---|
| Score / 100 | 63.3 | 57.4 | 55.6 |
| Domain | tasks | Team | DeepSeek V4 Flash (direct) | official o3 draft |
|---|---|---|---|---|
| Chemistry PhD | 10 | 75.5 | 69.4 | 60.8 |
| Consulting MBA | 10 | 70.8 | 70.3 | 64.9 |
| Finance MBA | 10 | 54.6 | 48.2 | 57.5 |
| Physics PhD | 8 | 49.6 | 37.8 | 35.2 |
Team, the DeepSeek V4 Flash baseline and the official o3 draft were all read by the same judging pipeline (per-criterion Yes/No, official prompt, temperature 0, no output cap), and are reported on the 38 tasks that all three completed — so the three self-run columns are directly comparable.
All three self-run columns come from one judging pipeline on the same 38 tasks: Team 63.3 · DeepSeek V4 Flash (direct) 57.4 · official o3 draft 55.6.
| Agreement with official labels | F1 | Mean score difference |
|---|---|---|
| 74.0% | 0.761 | +2.4 pts |
Our judging pipeline reproduces the official per-criterion labels on the o3 draft with 74.0% agreement (F1 0.761) and a mean score difference of +2.4 points — the numbers above sit on our pipeline's scale, calibrated against the official one.
36 : 1 — team draft : official o3 draft, 38 tasks, 1 unresolved.
On the same 38 tasks we also asked our own decision kernel to choose between the team draft and the official o3 draft. It was shown the question and the two drafts and nothing else — no rubric, no scores — and it selected the team draft 36 times.
This head-to-head run is supporting evidence, not the primary measurement: the per-criterion rubric columns above are the measurement, and the kernel reached its conclusion without ever seeing them.
The kernel reports a confidence value with every judgment, so a preference is never a bare claim. Median confidence on this benchmark is 85.9%, and it is not uniform across domains: highest on Physics PhD (95.3%) and Chemistry PhD (87.0%), lowest on Finance MBA (83.8%). It describes how far apart the two drafts are in the kernel's judgment: the lower it is, the closer the two are in quality — where either draft is a reasonable choice.
| Domain | tasks | Median confidence | Kernel picks the team draft |
|---|---|---|---|
| Chemistry PhD | 10 | 87.0% | 9/10 |
| Consulting MBA | 10 | 84.5% | 10/10 |
| Finance MBA | 10 | 83.8% | 9/10 |
| Physics PhD | 8 | 95.3% | 8/8 |
The 38-task run used a single fixed presentation order, with the team draft always offered as option A; the position-swap audit above, not this run, is the control for that.
| Domain | o3 | Grok 4 | DeepSeek-R1 |
|---|---|---|---|
| Chemistry PhD | 51.6 | 67.9 | 48.2 |
| Consulting MBA | 71.2 | 67.4 | 59.0 |
| Finance MBA | 44.5 | 41.3 | 39.1 |
| Physics PhD | 45.4 | 30.9 | 39.1 |
| Overall (40 tasks) | 53.2 | 51.9 | 46.3 |
Official evaluation of the official reference drafts, produced by the official judging pipeline. A different judge from the self-run columns above, so these rows are context only — not a head-to-head comparison.
Self-run with the official per-criterion rubric protocol (official prompt verbatim, temperature 0, top_p 1, no output cap). Our judging pipeline reproduces the official criterion labels on the o3 draft with 74.0% agreement and a mean difference of +2.4 points. Confidence describes how close the two drafts were in that judgment.
JudgeBench, 620 pairs (sources: MMLU-Pro / LiveBench Reasoning / Math / LiveCodeBench)
Self-run with the official judging protocol. Both columns use first successful verdict per pair; the six pairs whose first verdict failed are disclosed rather than imputed. Reference model measured on the same judged set.
| Model | Decider | DeepSeek V4 Flash (direct) |
|---|---|---|
| Accuracy % | 92.5 | 92.2 |
Agreement between the two: 96.1%.
| Segment | pairs | Decider | DeepSeek V4 Flash (direct) |
|---|---|---|---|
| Knowledge (MMLU-Pro) | 303 | 88.8 | 87.8 |
| Reasoning (LiveBench) | 149 | 98.0 | 96.6 |
| Math (LiveBench) | 90 | 92.2 | 95.6 |
| Code (LiveCodeBench) | 72 | 97.2 | 97.2 |
Confidence is emitted with every judgment and calibrated against outcomes on this benchmark (JudgeBench, 2026-09-08). It describes how far apart the two candidates are — and the accuracy column shows what that delivered here. Set your own threshold from the accuracy column; for high-stakes or irreversible decisions, apply your own review policy.
| Confidence band | judgments | share | observed accuracy | what the value means | suggested use |
|---|---|---|---|---|---|
| >= 90% | 283 | 45.6% | 99.6 | one candidate clearly stronger | Act on it |
| 80-90% | 184 | 29.7% | 94.0 | one candidate stronger | Go with it |
| 70-80% | 82 | 13.2% | 84.1 | closer call | Act on it after a quick look |
| < 70% | 65 | 10.5% | 67.7 | near-tie — evenly matched | Either choice is fine |
Every judgment ships with a calibrated confidence value. A high value means the comparison was decisive and the pick can be acted on directly; a low value means the two are evenly matched, where either choice is defensible and the decision belongs to criteria outside the answers.
ContextualJudgeBench (Salesforce), full official set of 2,000 pairs across all 8 splits; official reference values from the paper's Table 2.
Self-run with the official vanilla pairwise protocol; consistent accuracy (both response orders judged correctly); random floor 25%; failures disclosed (12 orders rerun-excluded). Reference model measured on the same judged set.
| Model | Overall consistent accuracy % |
|---|---|
| Decider | 67.1 |
Completed 1991 / 2000 official pairs.
Full official 8-split run completed (2026-09-09). Consistent accuracy 67.1% vs reference 65.4%: the advantage is strongest on faithfulness and refusal tasks, while the deliberately near-tie splits sit in the 46-60% range, reflecting the benchmark's difficulty design. Confidence tiers below are calibrated on this benchmark.
| Split | pairs | Decider | DeepSeek V4 Flash (direct) | o1 | o3-mini | r1 |
|---|---|---|---|---|---|---|
| Refusal (Answerable) | 250 | 97.2 | 94.4 | 96.0 | 95.2 | 92.0 |
| Faithfulness (QA) | 250 | 92.3 | 90.8 | 84.4 | 76.4 | 72.0 |
| Faithfulness (Summarization) | 250 | 68.7 | 67.2 | 59.2 | 58.0 | 50.4 |
| Completeness (Summarization) | 251 | 66.8 | 64.5 | 63.7 | 59.8 | 60.6 |
| Refusal (Unanswerable) | 250 | 59.8 | 62.8 | 48.4 | 34.4 | 52.0 |
| Completeness (QA) | 250 | 52.2 | 51.6 | 48.4 | 40.4 | 41.2 |
| Conciseness (QA) | 255 | 53.7 | 50.2 | 15.3 | 20.8 | 20.4 |
| Conciseness (Summarization) | 244 | 46.1 | 41.4 | 27.0 | 35.7 | 26.2 |
Confidence is emitted with every judgment and calibrated against outcomes on this benchmark (ContextualJudgeBench, 2026-09-09). It describes how far apart the two candidates are, and the accuracy column shows what that delivered here. This benchmark deliberately contains near-tie splits, so the accuracies run lower by design — the bands describe the comparison, not a promise about either answer.
| Confidence band | judgments | share | observed accuracy | what the value means | suggested use |
|---|---|---|---|---|---|
| >= 90% | 789 | 19.8% | 83.3 | one candidate clearly stronger | Act on it |
| 80-90% | 1486 | 37.3% | 76.4 | one candidate stronger | Go with it |
| 70-80% | 1184 | 29.7% | 63.6 | closer call | Act on it after a quick look |
| < 70% | 529 | 13.3% | 55.4 | near-tie — evenly matched | Either choice is fine |
Every judgment ships with a calibrated confidence value. A high value means the comparison was decisive and the pick can be acted on directly; a low value means the two are evenly matched, where either choice is defensible.
Self-run with the official protocol over the full 2,000 official pairs. Consistent accuracy requires the same correct pick in both presentation orders (random floor 25%). 12 orders (0.3%) were rerun-excluded after repeated platform failures; the reference model was measured on the same judged set. Previously published results measured on a subset are archived in the repository changelog.