The short version
Imagine two teachers marking the same exam. One is strict and rarely gives high marks; the other is generous. If your child's paper goes to the strict teacher and your neighbour's goes to the generous one, the neighbour's child looks better on paper — even if both did equally well. That isn't fair, and it isn't the child's fault.
Jurika fixes exactly this. It notices when a judge marks everyone a little lower (or higher) than the others, and gently adjusts for it — so every team is compared on the same footing, no matter which judges it happened to get.
In short: your team is judged on its work, not on the luck of who judged it.
The problem with a simple average
Judges are human. Some are strict and mark everyone low; others are generous. So a team's plain average depends partly on who judged it. A strong team seen by strict judges can end up below a weaker team seen by lenient ones. This is the single biggest, best-documented flaw in panel judging — and most competitions simply ignore it.
What Jurika does about it
It can't tell from a judge's own average whether they're strict or just happened to judge weaker teams. So it compares each judge to their fellow judges on the teams they shared. If a judge consistently marks a few points below everyone else who saw the same teams, Jurika learns that and lifts their teams back up. If another marks high, it eases their teams down. What's left is each team's fair score — the level it would have reached under a perfectly even panel.
A quick example
Team X is marked 80 by a strict judge. Team Y is marked 90 by a generous one.
On the teams those two judges both saw, the first consistently marks about 10 points lower.
A plain average says Y wins. Jurika learns the first judge is strict and the second generous, lifts X's 80 and lowers Y's 90, and finds they are really about equal. Team X is no longer punished just for drawing the tougher judge.
The safeguards
- Trust earned by evidence. A judge who saw only a couple of teams isn't trusted as much as one who saw many — their correction is gently pulled toward neutral, so we never over-adjust on thin data.
- The scale never drifts. All the adjustments are balanced to cancel out, so fair scores stay on the same 0–100 scale as the original marks.
- It knows when it can't help. The correction only works if judges are linked through shared teams. If they aren't, Jurika says so and falls back to the plain average rather than guessing — it never pretends to a fairness it can't deliver.
- Strictness vs. randomness. Being consistently strict is fine and correctable. A judge whose marks scatter randomly is a different problem — Jurika measures that separately and shows it to the organiser instead of quietly “correcting” it away.
Want the actual maths — the model, the algorithm, and the references?