Kodiak as a Judge: Accept When Confident, Escalate When Unsure
I turned Kodiak, my open 1B decision model, into a local judge API: it picks the better of two answers in about 50 ms, accepts only the verdicts it's sure of, and catches its own position bias by judging every pair both ways.
Every team that builds with LLMs ends up needing a judge. Which of these two answers is better? Did the new prompt beat the old one? Is this summary faithful to the source? The usual answer is to ask a big LLM, which works, but it's slow and it isn't free when you're running thousands of comparisons.
A recent paper from Carnegie Mellon, JEV-as-a-Judge: Accept When Confident, Escalate When Unsure, makes a simple argument: use a small, fast model that returns probabilities instead of essays, accept the verdicts it's confident about, and send the rest to the expensive judge. That idea is exactly what I've been building Kodiak to do. So this week I built Kodiak Judge: a local API and web page around Kodiak-v0.4-1B, my open 1B decision model.
What it does
You give it a question and two answers. It returns one of three verdicts (answer 1 is better, answer 2 is better, or they're about equal), a confidence, and an action: accept or escalate. It runs on my DGX Spark at about 50 milliseconds per pair, and that already includes judging the pair twice (more on that below).
When the verdict can be read off the text, it's sure, and it's right:
return x * x beats return x + x at 0.98, with the same verdict in both orders. A summary that sticks to the facts beats one that invents a cancellation:
Position bias, and why every pair gets judged twice
Here's the problem with judges, big or small: they often prefer whichever answer sits in a particular slot. The MT-Bench paper documented it for GPT-4 and Claude, and Kodiak has it too. When I swapped the two answers on 150 expert-judged MT-Bench pairs, 18% of Kodiak's verdicts changed.
So Kodiak Judge never trusts one order. Every pair runs twice, once as given and once swapped:
- The two runs' probabilities are averaged, so a mild position preference cancels out in the score.
- Both verdicts are shown, so you can see when the order mattered.
- If the two orders disagree, the pair is escalated, no matter how confident one order looked.
This example is the one that convinced me rule 3 matters:
Is ibuprofen safe for a dog? Answer 1 is right: it's toxic. With the long, wrong answer in the second slot, Kodiak picked it at 0.99 confidence. Swap the answers and it leans toward whatever is in the second slot again, now the right one, but only at 0.49. Averaging can't fix that: the confident wrong 0.99 outweighs the weak right 0.49, so the average still leans the wrong way (0.73). What saves it is rule 3. The two orders disagree, so there's no verdict at all; the pair goes to a person or an LLM. A judge that ran once would have accepted a dangerous answer at 99%.
(An earlier version of the page still printed the averaged lean in big letters on escalated pairs, which read like a verdict. Now an escalated pair says "No verdict: sent for review" and shows the lean small, marked "not trusted".)
I checked every example by hand
Here's every pair I tried, with the right answer, what Kodiak said in each order (both translated to the original order), and what the judge did, at the default 0.7 threshold:
| Example | Right answer | As given | Swapped | Judge |
|---|---|---|---|---|
| Faithful vs. made-up summary | 1 | 1 (0.88) | 1 (0.87) | accepted, correct |
| Working vs. broken code | 1 | 1 (0.98) | 1 (0.99) | accepted, correct |
| Capital of Australia | 1 | 2 (0.96) | 1 (0.72) | escalated |
| Ibuprofen for a dog | 1 | 2 (0.99) | 1 (0.48) | escalated |
| Boiling point of water | 1 | 1 (0.44) | 1 (0.94) | escalated (low confidence) |
| Same answer, different style | tie | 2 (0.79) | 2 (0.40) | escalated (low confidence) |
| Pay off a card or invest? (both partly right) | arguable | tie (0.55) | 1 (0.54) | escalated |
| "Good morning" in Spanish | 1 | 1 (0.48) | tie (0.55) | escalated |
| How many legs does a spider have? | 2 | 2 (0.99) | 1 (0.66) | escalated |
| What does HTTP 404 mean? | 1 | tie (0.64) | 1 (0.91) | escalated |
Two accepted, both right. Eight escalated, none accepted wrongly. But look at the "as given" and "swapped" columns: on these short factual pairs Kodiak mostly prefers whatever is in the second slot, sometimes at 0.99. On the eight it escalated, its averaged lean was right only 3 times. So the honest summary is: Kodiak's own judgment on world-knowledge questions is close to a coin flip with a strong slot preference, and the both-orders check is what keeps that from becoming confident wrong answers.
Where it's weak, and how it knows
Kodiak is a 1B encoder. It's good at reading and comparing text, and it knows very little about the world. Ask it which answer correctly names the capital of Australia and the verdict flips with the order, so it escalates:
That's the behaviour I want from a small judge: a fast first pass that hands off what it can't read off the page. It's also where the JEV paper draws its line: the small judge is strong "wherever a verdict can be read off the text."
The numbers
I measured it on 150 pairs from MT-Bench's expert human judgments (CC BY 4.0): 50 where the first answer won, 50 where the second won and 50 ties, so random guessing gets 33%.
| Accept at | Kodiak decides | Right when it decides | One order only |
|---|---|---|---|
| 0.6 | 68% | 64% | 60% of 71% |
| 0.7 | 57% | 67% | 63% of 59% |
| 0.8 | 43% | 71% | 68% of 49% |
| 0.9 | 35% | 69% | 67% of 35% |
Judging everything with no threshold, it's right 54% of the time. At a 0.8 threshold it decides 43% of pairs and gets 71% of those right; everything else goes to the expensive judge.
For context: in the MT-Bench paper's setup that counts ties (random = 33%), human experts agree with each other 63–67% of the time, and GPT-4 agrees with them 66% of the time. Our 150 pairs are a balanced sample, not the paper's full set, so the numbers aren't directly comparable. But it tells you how hard this task is: even experts disagree about a third of the time. Kodiak isn't a GPT-class judge, and I'm not claiming it is. What it can be is the cheap first pass in front of one.
Judging both orders helps exactly where you'd set a threshold: at 0.8, one order alone gets 68% of 49%; both orders get 71% of 43%. And escalating on disagreement removes the 18% of verdicts that depended on position. With no threshold at all, that alone raises accuracy on what's accepted from 54% to 60%.
The API
It's a small Python service with no web framework, just the standard library:
curl http://127.0.0.1:8790/v1/judge -H 'Content-Type: application/json' \
-d '{"question": "...", "answer_1": "...", "answer_2": "...", "threshold": 0.7}'
{"verdict": "first", "confidence": 0.98, "action": "accept",
"reason": "confident, and the same verdict in both orders",
"orders": {"as_given": {"verdict": "first"}, "swapped": {"verdict": "first"}, "disagree": false}}
When the action is "escalate", treat verdict as a lean, not an answer. There's also a batch endpoint (up to 64 pairs per call; six pairs took 82 ms) and a Python API. The model is open under Apache-2.0 on Hugging Face, and Kodiak's code and full build log are on GitHub.
What I learned
- Ask the judge twice. Swapping the order costs one more forward pass and catches a class of confident mistakes you'd never see otherwise.
- Averaging isn't enough; disagreement is the signal. A 0.99-confident wrong answer in one order can outvote a weak right one. The fact that the orders disagree is what tells you not to trust either.
- A small judge's best feature is knowing when to stop. Kodiak isn't good at everything, but it's honest enough to route on.
- Measure the trade-off, not one number. "54% accurate" sounds bad. "71% accurate on the 43% of cases it takes, and it hands off the rest" is a useful system.
Next I want to measure the full cascade: Kodiak first, a big model on whatever it escalates, and the total accuracy and cost against the big model alone. That's the experiment the JEV paper ran, and it's the one that would tell you whether this belongs in your pipeline. I'd also like to train the second-slot preference out of Kodiak itself, rather than only catching it.


