TuringCorp — Published Benchmark Results

Static, script-free version of our published results for readers that do not execute JavaScript. Last updated: 2026-09-13T11:45:00+08:00. Single source of truth: /benchmarks/latest.json.

All numbers are self-run with each benchmark's official protocol (not a third-party leaderboard). Coverage and exclusions are disclosed per benchmark; excluded items were rerun-excluded, never imputed. Confidence is a reading of the comparison: a high value means one candidate was clearly stronger and the pick can be acted on directly; a low value means the two were close in quality, where either choice is defensible.

Team Junior

ProfBench — 2026-09-10

ProfBench — 40 expert-level tasks, ten each in Physics PhD, Chemistry PhD, Finance MBA and Consulting MBA, scored criterion by criterion with the official rubric.

Self-run with the official rubric protocol: per-criterion Yes/No with the official prompt verbatim, temperature 0 / top_p 1, no output cap. Team delivery is content + reason. 38 of 40 tasks scored: two tasks were excluded because no answer was produced for them during the run (disclosed rather than imputed). The DeepSeek V4 Flash baseline and the official o3 draft were read by the same judging pipeline, and all three columns are reported on the 38 tasks every model completed.

ModelTeamDeepSeek V4 Flash (direct)official o3 draft
Score / 10063.357.455.6

Score by domain

DomaintasksTeamDeepSeek V4 Flash (direct)official o3 draft
Chemistry PhD1075.569.460.8
Consulting MBA1070.870.364.9
Finance MBA1054.648.257.5
Physics PhD849.637.835.2

Team, the DeepSeek V4 Flash baseline and the official o3 draft were all read by the same judging pipeline (per-criterion Yes/No, official prompt, temperature 0, no output cap), and are reported on the 38 tasks that all three completed — so the three self-run columns are directly comparable.

All three self-run columns come from one judging pipeline on the same 38 tasks: Team 63.3 · DeepSeek V4 Flash (direct) 57.4 · official o3 draft 55.6.

Judge calibration

Agreement with official labelsF1Mean score difference
74.0%0.761+2.4 pts

Our judging pipeline reproduces the official per-criterion labels on the o3 draft with 74.0% agreement (F1 0.761) and a mean score difference of +2.4 points — the numbers above sit on our pipeline's scale, calibrated against the official one.

Head-to-head arbitration by our decision kernel

36 : 1 — team draft : official o3 draft, 38 tasks, 1 unresolved.

On the same 38 tasks we also asked our own decision kernel to choose between the team draft and the official o3 draft. It was shown the question and the two drafts and nothing else — no rubric, no scores — and it selected the team draft 36 times.

This head-to-head run is supporting evidence, not the primary measurement: the per-criterion rubric columns above are the measurement, and the kernel reached its conclusion without ever seeing them.

Confidence reported with every judgment

The kernel reports a confidence value with every judgment, so a preference is never a bare claim. Median confidence on this benchmark is 85.9%, and it is not uniform across domains: highest on Physics PhD (95.3%) and Chemistry PhD (87.0%), lowest on Finance MBA (83.8%). It describes how far apart the two drafts are in the kernel's judgment: the lower it is, the closer the two are in quality — where either draft is a reasonable choice.

DomaintasksMedian confidenceKernel picks the team draft
Chemistry PhD1087.0%9/10
Consulting MBA1084.5%10/10
Finance MBA1083.8%9/10
Physics PhD895.3%8/8

The 38-task run used a single fixed presentation order, with the team draft always offered as option A; the position-swap audit above, not this run, is the control for that.

Official reference (official judging pipeline)

Domaino3Grok 4DeepSeek-R1
Chemistry PhD51.667.948.2
Consulting MBA71.267.459.0
Finance MBA44.541.339.1
Physics PhD45.430.939.1
Overall (40 tasks)53.251.946.3

Official evaluation of the official reference drafts, produced by the official judging pipeline. A different judge from the self-run columns above, so these rows are context only — not a head-to-head comparison.

Self-run with the official per-criterion rubric protocol (official prompt verbatim, temperature 0, top_p 1, no output cap). Our judging pipeline reproduces the official criterion labels on the o3 draft with 74.0% agreement and a mean difference of +2.4 points. Confidence describes how close the two drafts were in that judgment.

Decider Junior

JudgeBench — 2026-09-08

JudgeBench, 620 pairs (sources: MMLU-Pro / LiveBench Reasoning / Math / LiveCodeBench)

Self-run with the official judging protocol. Both columns use first successful verdict per pair; the six pairs whose first verdict failed are disclosed rather than imputed. Reference model measured on the same judged set.

ModelDeciderDeepSeek V4 Flash (direct)
Accuracy %92.592.2

Agreement between the two: 96.1%.

SegmentpairsDeciderDeepSeek V4 Flash (direct)
Knowledge (MMLU-Pro)30388.887.8
Reasoning (LiveBench)14998.096.6
Math (LiveBench)9092.295.6
Code (LiveCodeBench)7297.297.2

Confidence calibration

Confidence is emitted with every judgment and calibrated against outcomes on this benchmark (JudgeBench, 2026-09-08). It describes how far apart the two candidates are — and the accuracy column shows what that delivered here. Set your own threshold from the accuracy column; for high-stakes or irreversible decisions, apply your own review policy.

Confidence bandjudgmentsshareobserved accuracywhat the value meanssuggested use
>= 90%28345.6%99.6one candidate clearly strongerAct on it
80-90%18429.7%94.0one candidate strongerGo with it
70-80%8213.2%84.1closer callAct on it after a quick look
< 70%6510.5%67.7near-tie — evenly matchedEither choice is fine

Every judgment ships with a calibrated confidence value. A high value means the comparison was decisive and the pick can be acted on directly; a low value means the two are evenly matched, where either choice is defensible and the decision belongs to criteria outside the answers.

ContextualJudgeBench — 2026-09-09

ContextualJudgeBench (Salesforce), full official set of 2,000 pairs across all 8 splits; official reference values from the paper's Table 2.

Self-run with the official vanilla pairwise protocol; consistent accuracy (both response orders judged correctly); random floor 25%; failures disclosed (12 orders rerun-excluded). Reference model measured on the same judged set.

ModelOverall consistent accuracy %
Decider67.1

Completed 1991 / 2000 official pairs.

Full official 8-split run completed (2026-09-09). Consistent accuracy 67.1% vs reference 65.4%: the advantage is strongest on faithfulness and refusal tasks, while the deliberately near-tie splits sit in the 46-60% range, reflecting the benchmark's difficulty design. Confidence tiers below are calibrated on this benchmark.

SplitpairsDeciderDeepSeek V4 Flash (direct)o1o3-minir1
Refusal (Answerable)25097.294.496.095.292.0
Faithfulness (QA)25092.390.884.476.472.0
Faithfulness (Summarization)25068.767.259.258.050.4
Completeness (Summarization)25166.864.563.759.860.6
Refusal (Unanswerable)25059.862.848.434.452.0
Completeness (QA)25052.251.648.440.441.2
Conciseness (QA)25553.750.215.320.820.4
Conciseness (Summarization)24446.141.427.035.726.2

Confidence calibration

Confidence is emitted with every judgment and calibrated against outcomes on this benchmark (ContextualJudgeBench, 2026-09-09). It describes how far apart the two candidates are, and the accuracy column shows what that delivered here. This benchmark deliberately contains near-tie splits, so the accuracies run lower by design — the bands describe the comparison, not a promise about either answer.

Confidence bandjudgmentsshareobserved accuracywhat the value meanssuggested use
>= 90%78919.8%83.3one candidate clearly strongerAct on it
80-90%148637.3%76.4one candidate strongerGo with it
70-80%118429.7%63.6closer callAct on it after a quick look
< 70%52913.3%55.4near-tie — evenly matchedEither choice is fine

Every judgment ships with a calibrated confidence value. A high value means the comparison was decisive and the pick can be acted on directly; a low value means the two are evenly matched, where either choice is defensible.

Self-run with the official protocol over the full 2,000 official pairs. Consistent accuracy requires the same correct pick in both presentation orders (random floor 25%). 12 orders (0.3%) were rerun-excluded after repeated platform failures; the reference model was measured on the same judged set. Previously published results measured on a subset are archived in the repository changelog.