# TuringCorp — published benchmark results

Last updated: **2026-09-13T11:45:00+08:00**. Plain Markdown, no JavaScript. Machine-readable source: https://api.turingcorp.net/benchmarks/latest.json

All numbers are self-run with each benchmark's official protocol. Coverage and exclusions are disclosed per benchmark.

Method paper (preprint, DOI): https://doi.org/10.6084/m9.figshare.33684823

## Team Junior v1.1
## ProfBench — 2026-09-10

ProfBench — 40 expert-level tasks, ten each in Physics PhD, Chemistry PhD, Finance MBA and Consulting MBA, scored criterion by criterion with the official rubric.

| model | score / 100 |
|---|---|
| Team | 63.3 |
| DeepSeek V4 Flash (direct) | 57.4 |
| official o3 draft | 55.6 |


| domain | tasks | Team | DeepSeek V4 Flash (direct) | official o3 draft |
|---|---|---|---|---|
| Chemistry PhD | 10 | 75.5 | 69.4 | 60.8 |
| Consulting MBA | 10 | 70.8 | 70.3 | 64.9 |
| Finance MBA | 10 | 54.6 | 48.2 | 57.5 |
| Physics PhD | 8 | 49.6 | 37.8 | 35.2 |

Judge calibration on the o3 draft: 74.0% agreement with official labels, F1 0.761, mean difference +2.4 points (our scale vs the official one).

### Head-to-head arbitration by our own decision kernel

- Result: **36 : 1** — team draft : official o3 draft, 38 tasks, 1 unresolved.
- The kernel was shown the question and the two drafts only — no rubric, no scores.

- Position swap. On all 37 tasks where the kernel had made a pick, the comparison was re-run with the two drafts exchanged between option A and option B. 36 kept the same draft. The single change came on a task the kernel had already flagged as a near-tie (23% confidence); every pick it made at 80% confidence or above — 23 of the 37 — came back identical.
- No self-grading. The judgment comes from a panel of independent models rather than one model scoring its own output, which removes the self-preference bias of a single-judge setup.

This head-to-head run is supporting evidence, not the primary measurement: the per-criterion rubric columns above are the measurement, and the kernel reached its conclusion without ever seeing them.

### Confidence reported with every judgment

| domain | tasks | median confidence | kernel picks the team draft |
|---|---|---|---|
| Chemistry PhD | 10 | 87.0% | 9/10 |
| Consulting MBA | 10 | 84.5% | 10/10 |
| Finance MBA | 10 | 83.8% | 9/10 |
| Physics PhD | 8 | 95.3% | 8/8 |

The kernel reports a confidence value with every judgment, so a preference is never a bare claim. Median confidence on this benchmark is 85.9%, and it is not uniform across domains: highest on Physics PhD (95.3%) and Chemistry PhD (87.0%), lowest on Finance MBA (83.8%). It describes how far apart the two drafts are in the kernel's judgment: the lower it is, the closer the two are in quality — where either draft is a reasonable choice.

The 38-task run used a single fixed presentation order, with the team draft always offered as option A; the position-swap audit above, not this run, is the control for that.

### Official reference (official judging pipeline, context only)

| domain | o3 | Grok 4 | DeepSeek-R1 |
|---|---|---|---|
| Chemistry PhD | 51.6 | 67.9 | 48.2 |
| Consulting MBA | 71.2 | 67.4 | 59.0 |
| Finance MBA | 44.5 | 41.3 | 39.1 |
| Physics PhD | 45.4 | 30.9 | 39.1 |
| Overall | 53.2 | 51.9 | 46.3 |

Official evaluation of the official reference drafts, produced by the official judging pipeline. A different judge from the self-run columns above, so these rows are context only — not a head-to-head comparison.

Protocol and coverage:
- Self-run with the official rubric protocol: per-criterion Yes/No with the official prompt verbatim, temperature 0 / top_p 1, no output cap. Team delivery is content + reason. 38 of 40 tasks scored: two tasks were excluded because no answer was produced for them during the run (disclosed rather than imputed). The DeepSeek V4 Flash baseline and the official o3 draft were read by the same judging pipeline, and all three columns are reported on the 38 tasks every model completed.

## Decider Junior
## JudgeBench — 2026-09-08

JudgeBench, 620 pairs (sources: MMLU-Pro / LiveBench Reasoning / Math / LiveCodeBench)

| model | accuracy % |
|---|---|
| Decider | 92.5 |
| DeepSeek V4 Flash (direct) | 92.2 |

| segment | pairs | Decider | DeepSeek V4 Flash (direct) |
|---|---|---|---|
| Knowledge (MMLU-Pro) | 303 | 88.8 | 87.8 |
| Reasoning (LiveBench) | 149 | 98.0 | 96.6 |
| Math (LiveBench) | 90 | 92.2 | 95.6 |
| Code (LiveCodeBench) | 72 | 97.2 | 97.2 |

### Confidence calibration

| confidence band | judgments | share | observed accuracy | what the value means | suggested use |
|---|---|---|---|---|---|
| >= 90% | 283 | 45.6% | 99.6 | one candidate clearly stronger | Act on it |
| 80-90% | 184 | 29.7% | 94.0 | one candidate stronger | Go with it |
| 70-80% | 82 | 13.2% | 84.1 | closer call | Act on it after a quick look |
| < 70% | 65 | 10.5% | 67.7 | near-tie — evenly matched | Either choice is fine |

Confidence is emitted with every judgment and calibrated against outcomes on this benchmark (JudgeBench, 2026-09-08). It describes how far apart the two candidates are — and the accuracy column shows what that delivered here. Set your own threshold from the accuracy column; for high-stakes or irreversible decisions, apply your own review policy.

Protocol and coverage:
- Self-run with the official judging protocol. Both columns use first successful verdict per pair; the six pairs whose first verdict failed are disclosed rather than imputed. Reference model measured on the same judged set.

## ContextualJudgeBench (CJB) — 2026-09-09

ContextualJudgeBench (Salesforce), full official set of 2,000 pairs across all 8 splits; official reference values from the paper's Table 2.

| model | overall consistent accuracy % |
|---|---|
| Decider | 67.1 |

Completed 1991 / 2000 official pairs. Consistent accuracy = the same pick correct in both presentation orders (random floor 25%).

| split | pairs | Decider | DeepSeek V4 Flash (direct) | o1 | o3-mini | r1 |
|---|---|---|---|---|---|---|
| Refusal (Answerable) | 250 | 97.2 | 94.4 | 96.0 | 95.2 | 92.0 |
| Faithfulness (QA) | 250 | 92.3 | 90.8 | 84.4 | 76.4 | 72.0 |
| Faithfulness (Summarization) | 250 | 68.7 | 67.2 | 59.2 | 58.0 | 50.4 |
| Completeness (Summarization) | 251 | 66.8 | 64.5 | 63.7 | 59.8 | 60.6 |
| Refusal (Unanswerable) | 250 | 59.8 | 62.8 | 48.4 | 34.4 | 52.0 |
| Completeness (QA) | 250 | 52.2 | 51.6 | 48.4 | 40.4 | 41.2 |
| Conciseness (QA) | 255 | 53.7 | 50.2 | 15.3 | 20.8 | 20.4 |
| Conciseness (Summarization) | 244 | 46.1 | 41.4 | 27.0 | 35.7 | 26.2 |

### Confidence calibration

| confidence band | judgments | share | observed accuracy | what the value means | suggested use |
|---|---|---|---|---|---|
| >= 90% | 789 | 19.8% | 83.3 | one candidate clearly stronger | Act on it |
| 80-90% | 1486 | 37.3% | 76.4 | one candidate stronger | Go with it |
| 70-80% | 1184 | 29.7% | 63.6 | closer call | Act on it after a quick look |
| < 70% | 529 | 13.3% | 55.4 | near-tie — evenly matched | Either choice is fine |

Confidence is emitted with every judgment and calibrated against outcomes on this benchmark (ContextualJudgeBench, 2026-09-09). It describes how far apart the two candidates are, and the accuracy column shows what that delivered here. This benchmark deliberately contains near-tie splits, so the accuracies run lower by design — the bands describe the comparison, not a promise about either answer.

Protocol and coverage:
- Self-run with the official vanilla pairwise protocol; consistent accuracy (both response orders judged correctly); random floor 25%; failures disclosed (12 orders rerun-excluded). Reference model measured on the same judged set.

