<!-- llms-explorer concept facts · https://llms-explorer.com/tree/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/ · pack 2026-09-25 · ~13539 tokens -->

# LLM-as-judge bias & calibration (kappa, binary-vs-Likert, position/verbosity/self-preference)

> Depth-first rabbithole dossier for LLM-as-judge bias & calibration (kappa, binary-vs-Likert, position/verbosity/self-preference); source-anchored research pack.

Parent: [Eval-Driven Development for LLM Apps](https://llms-explorer.com/tree/eval-driven-development-for-llm-apps/) · 7 facets · 72 facts · page: https://llms-explorer.com/tree/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/

## Definitions

- An LLM judge is a model that scores or ranks model outputs in place of a human rater. *Bias* means the judge's verdict changes with a factor that should not matter, such as answer order, answer length, or whether the judge wrote the text itself. *Calibration* covers two jobs: measuring how far the judge's labels sit from human labels using a chance-corrected statistic, and correcting the judge's output. Correction can happen through protocol (order swapping, choice of scale), through statistics (length regression, probability-weighted scores, human-anchored estimators), or through the choice o — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/rabbithole-synthesis.md#definition`

## Structure and components

- 1. In MT-Bench pairwise judging, GPT-4 gave the same verdict after swapping answer order in 65.0% of cases; 30.0% of cases favored the first position. https://arxiv.org/html/2306.05685v4 2. In the same test, Claude-v1 was position-consistent in only 23.8% of cases and favored the first position in 75.0%; GPT-3.5 was consistent in 46.2% and favored the first position in 50.0%. https://arxiv.org/html/2306.05685v4 3. Answer order alone can flip outcomes: with order manipulation, Vicuna-13B beat ChatGPT on 66 of 80 Vicuna-Benchmark queries under a ChatGPT-family judge. https://arxiv.org/abs/2305.1 — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/practice.md#position-bias`
- - **C1.** Position bias means the verdict changes when the order of the candidates is swapped. [M,E] https://arxiv.org/abs/2305.17926 - **C2.** On MT-Bench, the verdict stayed the same after a swap 23.8% of the time for Claude-v1, 46.2% for GPT-3.5 and 65.0% for GPT-4. [M,H,E,P] https://arxiv.org/html/2306.05685v4 - **C3.** In the same test, the first position was favored 75.0% of the time by Claude-v1, 50.0% by GPT-3.5 and 30.0% by GPT-4. [M,P] https://arxiv.org/html/2306.05685v4 - **C4.** The direction of the bias depends on the model: some judges favor the first slot (primacy) and others th — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/rabbithole-synthesis.md#c-position-bias`

## How it works

- **D1. What causes self-preference?** — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/rabbithole-synthesis.md#disagreements-side-by-side-not-averaged`
- 37. Chen et al., "Do LLM Evaluators Prefer Themselves for a Reason?" (arXiv 2504.03846, 2025-04-04), separated legitimate from harmful self-preference. They found that "stronger models prefer themselves mostly legitimately," but "harmful self-preference persists when evaluator models err as generators." https://arxiv.org/abs/2504.03846 38. Li et al. (arXiv 2601.03444, 2026-01-06) compared 0–5, 0–10, and 0–100 scales using ICC. Human panels were stable across scales. LLM agreement moved with the scale, and 0–5 gave the highest human–LLM alignment. https://arxiv.org/abs/2601.03444 — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/history.md#d-revisions-2025-2026`
- - Default to binary, criterion-specific judges with a written critique before the verdict (claims 24–26, 29). Reserve Likert scales for cases where you have validated that the judge uses the scale consistently. - Calibrate every judge on a held-out human-labeled set. Report kappa together with TPR and TNR, never raw agreement alone (claims 17–21). - For pairwise judging, always run both orders and record a tie when the two orders disagree (claim 6). Expect the most flips on near-tie pairs (claim 4). - Length-control win rates or penalize padding explicitly (claims 10–11). - Do not let a model — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/practice.md#operational-implications-derived-from-the-claims-above`

## Measurements and reference values

- - **B1.** G-Eval (Liu et al., arXiv 2303.16634, submitted 2023-03-29) scored generated text with GPT-4 and chain-of-thought. It reported Spearman 0.514 with humans on summarization. [H] https://arxiv.org/abs/2303.16634 - **B2.** On a 1–5 scale, direct numeric scores clump on one value (often 3), which lowers variance. LLMs also emit whole numbers even when asked for decimals, which produces ties. [M,H] https://arxiv.org/html/2303.16634 - **B3.** G-Eval's fix reads the probabilities of the score tokens and computes score = Σ p(sᵢ)·sᵢ. The result is a continuous, finer-grained score. This was an — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/rabbithole-synthesis.md#b-llm-as-judge-emerges-with-bias-caveats-march-june-2023`
- - https://arxiv.org/abs/2306.05685 — Zheng et al. 2023, MT-Bench / Chatbot Arena (LLM-as-a-judge) - https://arxiv.org/html/2306.05685 — same, full text (Tables 2–3) - https://arxiv.org/abs/2305.17926 — Wang et al. 2023, Large Language Models are not Fair Evaluators - https://arxiv.org/html/2406.12624 — Thakur et al. 2024, Judging the Judges: alignment and vulnerabilities - https://arxiv.org/abs/2606.19544 — Norman, Rivera, Hughes 2026, Reliability without Validity - https://pubmed.ncbi.nlm.nih.gov/2348207/ — Feinstein & Cicchetti 1990, High agreement but low kappa I - https://www.sciencedirect — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/edge-cases.md#sources`
- - **D1.** Verbosity bias means the judge prefers a longer answer when the extra length adds no quality. [M,E] https://arxiv.org/html/2306.05685 - **D2.** In Zheng et al.'s "repetitive list" attack, which pads answers without adding content, Claude-v1 and GPT-3.5 failed 91.3% of the time and GPT-4 failed 8.7% of the time. The attack used 23 cases. [M,H,E,P] https://arxiv.org/html/2306.05685v4 - **D3.** Saito et al. (arXiv 2310.10076, 2023-10-16) define verbosity bias relative to human labels. Their metric is the accuracy-parity gap P(Y′=Y|S=Y) − P(Y′=Y|S=1−Y), grouped by which answer is longer. — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/rabbithole-synthesis.md#d-verbosity-length-bias`
- **D3. Does reasoning before the verdict reduce bias?** - *Yes:* few-shot prompting and chain-of-thought (C11, H5); Multiple Evidence Calibration (B8); long CoT reduces harmful self-preference (E9); explaining the rating raises human correlation (G12); critique iterations reached over 90% expert agreement (G4). - *No or worse:* reasoning models are not reliably less self-biased (E7); position bias grows with reasoning-trace length (C9); G-Eval's auto-CoT did not reliably help (G13). - The "yes" evidence comes from explicit rationales that the prompt asks for. The "no" evidence comes from native — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/rabbithole-synthesis.md#disagreements-side-by-side-not-averaged`
- 16. The order in which answers appear can decide the outcome. With ChatGPT as judge, Vicuna-13B beat ChatGPT on 66 of 80 queries just by swapping the order. https://arxiv.org/abs/2305.17926 17. Default-prompt consistency after swapping positions was 65.0% for GPT-4, 46.2% for GPT-3.5 and 23.8% for Claude-v1. https://arxiv.org/html/2306.05685 18. Standard mitigation: run each comparison in both orders and count a win only if the same answer wins both times (MT-Bench's "conservative" rule). A related method, Balanced Position Calibration, averages the result over both orders. https://arxiv.org/h — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/edge-cases.md#position-bias`
- 6. Position bias means the verdict changes when the order of the candidates is swapped. In one test, ChatGPT acted as the judge, and Vicuna-13B "beat" ChatGPT on 66 of 80 queries purely because of answer order. — https://arxiv.org/abs/2305.17926 7. When two similar answers were swapped on MT-Bench, the verdict stayed the same 23.8% of the time for Claude-v1, 46.2% for GPT-3.5, and 65.0% for GPT-4. Claude-v1 favored the first position 75.0% of the time, and GPT-4 favored it 30.0% of the time. — https://arxiv.org/html/2306.05685 8. The direction of the bias depends on the model. Some judges favo — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/mechanism.md#b-position-bias`
- 46. With a small human-labeled subset, the judge's bias can be corrected without introducing new bias. The AutoEval estimator averages the judge's scores over all items, then adds a rectifier: the mean of (human − λ·judge) on the labeled items. The estimate stays unbiased for any fixed λ, even if the judge itself is biased. — https://arxiv.org/html/2403.07008 47. The tuning parameter λ goes toward 0 when the judge is uninformative, which falls back to the classical human-only estimate. With GPT-4 as the judge, the method raised the effective human sample size by up to 50%. — https://arxiv.org/ — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/mechanism.md#g-calibration-against-humans-statistical-correction`
- 8. Under a "repetitive list" attack that padded answers without adding content, Claude-v1 and GPT-3.5 judges failed on 91.3% of 23 cases; GPT-4 failed on 8.7%. https://arxiv.org/html/2306.05685v4 9. GPT-4 prefers longer answers more than human raters do when the answers are of similar quality. https://arxiv.org/abs/2310.10076 10. Length bias can be removed statistically. Length-controlled AlpacaEval fits a GLM on length difference and predicts the preference at zero length difference. https://arxiv.org/abs/2404.04475 11. Length control raised AlpacaEval's Spearman correlation with LMSYS Chatbo — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/practice.md#verbosity-length-bias`
- - **H1.** AutoEval (Boyeau et al.) averages the judge's scores over all items, then adds a rectifier: the mean of (human − λ·judge) on a small human-labelled subset. The estimate stays unbiased for any fixed λ, even if the judge itself is biased. [M] https://arxiv.org/html/2403.07008 - **H2.** λ goes toward 0 when the judge is uninformative, which falls back to the classical human-only estimate. With GPT-4 as the judge, the method raised the effective human sample size by up to 50%. [M] https://arxiv.org/abs/2403.07008 - **H3.** AutoCalibrate (arXiv 2309.13308) calibrates an off-the-shelf judg — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/rabbithole-synthesis.md#h-calibration-against-humans-and-abstention`
- | Position | Evidence | Source | |---|---|---| | The judge recognizes its own text (identity) | Self-recognition correlates linearly with self-preference, and fine-tuning evidence supports a causal link (E2) | https://arxiv.org/abs/2404.13076 | | The judge prefers familiar text (low perplexity) | Higher scores for low-perplexity text regardless of who wrote it (E3, E4) | https://arxiv.org/abs/2410.21819 ; https://arxiv.org/abs/2405.01724 | | The judge is uncertain on items it fails | Only 51% of significant examples survive a quality baseline, but the survivors hold 89.6% of the probability ma — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/rabbithole-synthesis.md#disagreements-side-by-side-not-averaged`
- **D6. Pairwise or pointwise?** - *Pointwise is safer:* under distractor features, pairwise verdicts flip about 35% of the time against about 9% for absolute scores (D9), and pairwise exposes position bias (C12). https://arxiv.org/abs/2504.14716 - *Pairwise is more stable on subjective criteria* (G14). https://eugeneyan.com/writing/llm-evaluators/ - *Unresolved context:* the canonical MT-Bench agreement figures come mostly from pairwise setups (B10–B11). https://arxiv.org/html/2306.05685 — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/rabbithole-synthesis.md#disagreements-side-by-side-not-averaged`
- 14. Verbosity bias means the judge prefers a longer answer even when the extra length adds no quality. In the "repetitive list" attack, Claude-v1 and GPT-3.5 fell for the padded answer 91.3% of the time, and GPT-4 fell for it 8.7% of the time. — https://arxiv.org/html/2306.05685 15. Saito et al. define verbosity bias against human labels, not in absolute terms. Their metric is the accuracy-parity gap P(Y′=Y|S=Y) − P(Y′=Y|S=1−Y), grouped by which answer is longer. On HH-RLHF, GPT-4 scores 0.328 and GPT-3.5 scores 0.428. GPT-4 prefers long answers more than humans do. — https://arxiv.org/html/23 — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/mechanism.md#c-verbosity-length-bias`
- 19. Self-preference (also called self-enhancement) means the judge rates its own outputs higher. Zheng et al. saw GPT-4 give itself a 10% higher win rate and Claude-v1 give itself a 25% higher one. The authors said the data were too limited to confirm the bias. — https://arxiv.org/html/2306.05685 20. A judge's ability to recognize its own outputs correlates linearly with the strength of its self-preference. Panickssery et al. changed self-recognition by fine-tuning and found the evidence consistent with a causal link. — https://arxiv.org/abs/2404.13076 21. Wataoka et al. propose a different me — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/mechanism.md#d-self-preference-and-familiarity-bias`
- | Pass | Focus | New claims | Cumulative | Rate | |---|---|---|---|---| | 0 | Canonical papers (Zheng, Wang, Panickssery, Dubois, Shi, Thakur) | 17 | 17 | 100% | | 1 | Primary-text numbers, self-preference mechanism, scales, juries | 17 | 34 | 50% | | 2 | Kappa paradox, G-Eval, AutoEval, CALM, Arize, Saito, pairwise/pointwise | 10 | 44 | 23% | | 3 | Kappa formula and weights, interpretation bands, JUDGE-BENCH | 3 | 47 | 6% | — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/mechanism.md#pass-log-new-information-rate`
- **D10. What is the Feinstein–Cicchetti remedy?** - H (paper II, read via PubMed): no single omnibus index fixes the paradox, so report ppos and pneg separately. https://pubmed.ncbi.nlm.nih.gov/2189948/ - M (paper I, secondary account after a 403): report prevalence and bias indices alongside κ. https://www.sciencedirect.com/science/article/abs/pii/089543569090158L - These are probably the recommendations of the two companion papers, not a true conflict. Neither report confirmed this. — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/rabbithole-synthesis.md#disagreements-side-by-side-not-averaged`
- Statistics and methodology: - https://journals.sagepub.com/doi/10.1177/001316446002000104 — Cohen 1960, kappa - https://pubmed.ncbi.nlm.nih.gov/843571/ — Landis & Koch 1977 - https://pubmed.ncbi.nlm.nih.gov/2348207/ — Feinstein & Cicchetti 1990, High agreement but low kappa I - https://www.sciencedirect.com/science/article/abs/pii/089543569090158L — same paper, publisher page - https://pubmed.ncbi.nlm.nih.gov/2189948/ — Cicchetti & Feinstein 1990, paper II (ppos/pneg) - https://pmc.ncbi.nlm.nih.gov/articles/PMC3900052/ — McHugh 2012, the kappa statistic - https://en.wikipedia.org/wiki/Cohen%27 — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/rabbithole-synthesis.md#sources`
- 1. Jacob Cohen introduced the kappa coefficient in "A Coefficient of Agreement for Nominal Scales", *Educational and Psychological Measurement* 20(1):37–46, 1960. https://journals.sagepub.com/doi/10.1177/001316446002000104 2. Landis & Koch (1977, *Biometrics* 33:159–174) proposed the verbal bands still quoted today. For example, a kappa of 0.61–0.80 counts as "substantial" agreement. https://pubmed.ncbi.nlm.nih.gov/843571/ 3. Feinstein & Cicchetti (1990) showed the "high agreement but low kappa" paradox. In a binary 2×2 table, imbalanced marginal totals can make kappa low even when observed ag — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/history.md#a-agreement-statistics-predate-llms-1960-1990`
- The new-information rate is falling but did not reach <5% on two consecutive passes. A pass 3 would most likely add in-paper numbers from Saito, CALM, and Shi, plus human-annotator length bias. It would probably not change the story told here. — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/history.md#pass-log`
- 12. In MT-Bench, GPT-4 gave itself a win rate about 10% above the human baseline, and Claude-v1 gave itself about 25% above it; GPT-3.5 showed no clear self-preference. https://arxiv.org/html/2306.05685v4 13. GPT-4 and Llama 2 can recognize their own outputs with non-trivial accuracy. After fine-tuning, self-recognition ability correlated linearly with the strength of self-preference. https://arxiv.org/abs/2404.13076 14. An alternative mechanism: LLM judges score lower-perplexity (more familiar) text higher than humans do. On this view, self-preference is a familiarity effect, not identity rec — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/practice.md#self-preference-self-enhancement`
- - Pass 0: broad draft from primary bias papers, 22 claims. - Pass 1: disconfirming sources and mechanisms, with 9 new claims (rate 9/31 ≈ 29%). - Verdict: BUDGET_EXHAUSTED (soft stop, not saturation). One or two more passes would probably still pay off, most likely on: probability-weighted scoring, judge panels/juries, and per-criterion rubric anchoring. — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/practice.md#depth-passes`

## Problems, failure modes and limitations

- 19. Saito et al., "Verbosity Bias in Preference Labeling by LLMs" (arXiv 2310.10076, 2023-10-16), found that "GPT-4 prefers longer answers more than humans." They proposed a metric for this bias. https://arxiv.org/abs/2310.10076 20. Prometheus (Kim et al., arXiv 2310.08491, 2023-10-12) trained a 13B judge on 1,000 score rubrics. It reached Pearson 0.897 with humans, against 0.882 for GPT-4. This made rubric-conditioned Likert scoring the open-model standard. https://arxiv.org/abs/2310.08491 21. AutoCalibrate (Liu et al., arXiv 2309.13308, 2023-09-23) calibrated an off-the-shelf judge without g — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/history.md#c-biases-isolated-and-measured-one-at-a-time-late-2023-2024`
- 5. G-Eval (Liu et al., arXiv 2303.16634, submitted 2023-03-29) used GPT-4 with chain-of-thought to score generated text. It reported Spearman 0.514 with humans on summarization. https://arxiv.org/abs/2303.16634 6. G-Eval computes its score as a probability-weighted sum over the score tokens, Σ p(sᵢ)·sᵢ. The reason is that on a 1–5 scale "one digit usually dominates the distribution," and LLMs output integers even when asked for decimals. This was an early calibration step for Likert-style direct scoring. https://arxiv.org/html/2303.16634 7. G-Eval reported that G-Eval-4 "always gives higher sc — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/history.md#b-llm-as-judge-emerges-with-bias-caveats-built-in-march-june-2023`
- LLM-as-judge primary papers: - https://arxiv.org/abs/2303.16634 — Liu et al. 2023, G-Eval - https://arxiv.org/html/2303.16634 — G-Eval full text - https://arxiv.org/abs/2305.01937 — Chiang & Lee 2023, LLMs as an alternative to human evaluation - https://arxiv.org/abs/2305.17926 — Wang et al. 2023, Large Language Models are not Fair Evaluators - https://arxiv.org/html/2305.17926 — Wang et al. full text - https://arxiv.org/abs/2306.05685 — Zheng et al. 2023, MT-Bench and Chatbot Arena - https://arxiv.org/html/2306.05685 — Zheng et al. full text - https://arxiv.org/html/2306.05685v4 — Zheng et al — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/rabbithole-synthesis.md#sources`
- 21. Verbosity bias means preferring the longer answer when the answers are of similar quality. GPT-4 prefers longer answers more than human raters do. https://arxiv.org/abs/2310.10076 22. Verbosity bias depends heavily on the judge. On the repetitive-list attack, Claude-v1 and GPT-3.5 failed 91.3% of the time, while GPT-4 failed 8.7%. https://arxiv.org/html/2306.05685 23. Length can be gamed. Changing only the verbosity instruction moved the AlpacaEval baseline's win rate from 22.9% to 64.3%. With length control, the range narrowed to 41.9–51.6%. https://arxiv.org/html/2404.04475v1 24. Control — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/edge-cases.md#verbosity-length-bias`
- 1. The canonical LLM-as-judge paper names three biases: position, verbosity, and self-enhancement. It also names limited reasoning ability as a separate limitation. — https://arxiv.org/abs/2306.05685 2. In the same paper, GPT-4 judges agree with human preferences over 80% of the time. That matches the human-to-human agreement level. — https://arxiv.org/abs/2306.05685 3. Agreement depends on how ties are counted. Without ties, GPT-4 agrees with experts 85% of the time and humans agree with each other 81% of the time. With ties included, the figures fall to 66–70% and 60–63%. — https://arxiv.org — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/mechanism.md#a-parts-of-the-mechanism-what-a-judge-bias-is`
- 1. Strong judges such as GPT-4 agree with human preferences more than 80% of the time, which matches human–human agreement. This is the baseline claim that later work qualifies. https://arxiv.org/abs/2306.05685 2. Percent agreement is high for almost every judge, so it cannot tell judges apart. Judges with more than 90% agreement still differed by more than 10 points in the scores they assigned. https://arxiv.org/html/2406.12624 3. Scott's π tells judges apart better than percent agreement does. Thakur et al. use Scott's π, not Cohen's kappa. Humans reached 96.2±1.07 π, and the best LLM judge — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/edge-cases.md#agreement-metrics-kappa-and-its-relatives`
- - **A1.** Jacob Cohen introduced the kappa coefficient in 1960 ("A Coefficient of Agreement for Nominal Scales", *Educational and Psychological Measurement* 20(1):37–46). [H] https://journals.sagepub.com/doi/10.1177/001316446002000104 - **A2.** κ = (Pr(a) − Pr(e)) / (1 − Pr(e)). Pr(a) is the observed agreement. Pr(e) is the agreement expected by chance, computed from the marginal totals. [M] https://pmc.ncbi.nlm.nih.gov/articles/PMC3900052/ - **A3.** Landis & Koch (1977, *Biometrics* 33:159–174) proposed the verbal bands still quoted today. For example, κ 0.61–0.80 counts as "substantial". [M, — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/rabbithole-synthesis.md#a-agreement-statistics-before-llms-1960-1990`
- - **E1.** Zheng et al. saw GPT-4 give itself about a 10% higher win rate and Claude-v1 about 25% higher. GPT-3.5 showed no clear self-preference. The authors wrote that "our study cannot determine whether the models exhibit a self-enhancement bias." [M,H,E,P] https://arxiv.org/html/2306.05685v4 - **E2.** Panickssery, Bowman & Feng (arXiv 2404.13076) found "a linear correlation between self-recognition capability and the strength of self-preference bias" in GPT-4 and Llama 2. They changed self-recognition by fine-tuning, and the result is consistent with a causal link. [M,H,E,P] https://arxiv.o — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/rabbithole-synthesis.md#e-self-preference-self-enhancement-and-familiarity`
- **D2. How large is self-preference, and does it shrink with scale?** - *It shrinks sharply:* after the gold correction, DBG falls from 21.6% (Llama-3.1-8B) to 0.4% (Llama-3.1-70B) (E6). https://arxiv.org/html/2506.02592v1 - *Its harmful form grows:* stronger models show more harmful self-preference when they are wrong (E9). https://arxiv.org/abs/2504.03846 - *Inconclusive or absent:* the MT-Bench authors could not confirm it (E1), and one RAG study found none (E11). https://arxiv.org/html/2306.05685v4 ; https://arxiv.org/abs/2410.20833 - These studies measure different things: an unconditional — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/rabbithole-synthesis.md#disagreements-side-by-side-not-averaged`
- **D4. How large is verbosity bias?** - *Severe:* 91.3% failure for Claude-v1 and GPT-3.5 on padding (D2); GPT-4 prefers length more than humans (D4); the win rate was gameable from 22.9% to 64.3% (D7); AlpacaEval needed length control even with a GPT-4-class annotator (D5–D6). - *Small:* GPT-4 failed the padding attack only 8.7% of the time (D2); under a single rubric, bias was below 0.011 for all 21 judges (D10). - *Over-correction risk:* humans also value length to some degree (D4, D8), so fully removing length may over-correct. - The size of the bias depends on the judge's generation and on — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/rabbithole-synthesis.md#disagreements-side-by-side-not-averaged`
- 27. Self-recognition and self-preference are linked. When models were fine-tuned, self-recognition ability correlated linearly with the strength of self-preference. The authors read this as evidence that self-recognition causes self-preference. https://arxiv.org/abs/2404.13076 28. Alternative explanation: judges rate low-perplexity text higher than humans do, whether or not they wrote it. On this view, self-preference is a special case of preferring familiar text. https://arxiv.org/abs/2410.21819 ; https://arxiv.org/abs/2405.01724 29. The original MT-Bench study was inconclusive. GPT-4 favored — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/edge-cases.md#self-preference`

## Comparisons and alternatives

- - **Concept:** LLM-as-judge bias & calibration (kappa, binary-vs-Likert, position/verbosity/self-preference) - **Parent context:** Eval-Driven Development for LLM Apps - **Run:** frontier-2026-09-25 · /rabbithole (report-only mode) · written 2026-09-25 - **Verdict:** BUDGET_EXHAUSTED (soft stop after 4 passes; the new-information rate was falling but had not reached two consecutive passes under 5%) — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/mechanism.md`
- **In scope:** how the field came to know that LLM judges have position, verbosity (length), and self-preference biases. Also how agreement with humans is measured (Cohen's kappa and its alternatives), and the binary-versus-Likert scoring debate. Also the calibration methods proposed to correct these biases. Every claim is tied to a primary paper, a statistics source, or a dated practitioner post. — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/history.md#scope`
- **In scope:** where LLM judges disagree with humans, and how that gets measured. This covers agreement metrics (kappa and its relatives), score format (binary vs ordinal/Likert), and three named biases: position, verbosity/length and self-preference. For each one, the report asks where it holds, where it breaks, and what evidence goes against it. **Out of scope:** judge-prompt design in general, reward models, eval-harness tooling, and other bias families (authority, gender, beauty, etc.). Other bias families appear here only to show that humans share them. — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/edge-cases.md#scope`
- - **Is self-preference a bias at all?** Zheng et al. measured a 10–25% uplift but could not confirm a bias (claim 17). Panickssery et al. and Wataoka et al. give mechanisms for a real bias (self-recognition, claim 24; perplexity, claim 33). Chen et al. 2025 argue that much of it is legitimate: stronger models are often actually better (claim 37). These positions are not resolved. - **Are LLM judges as reliable as humans?** Zheng et al. say GPT-4 matches human–human agreement (85% vs 81%, claims 13–14). Thakur et al. say chance-corrected agreement shows a clear gap (π 88 vs 96.2, claim 29), and — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/history.md#unresolved-disagreements`
- - https://journals.sagepub.com/doi/10.1177/001316446002000104 — Cohen 1960, kappa - https://pubmed.ncbi.nlm.nih.gov/843571/ — Landis & Koch 1977 - https://pubmed.ncbi.nlm.nih.gov/2348207/ — Feinstein & Cicchetti 1990 (I) - https://pubmed.ncbi.nlm.nih.gov/2189948/ — Cicchetti & Feinstein 1990 (II) - https://arxiv.org/abs/2303.16634 — G-Eval - https://arxiv.org/html/2303.16634 — G-Eval full text - https://arxiv.org/abs/2305.01937 — Chiang & Lee 2023 - https://arxiv.org/abs/2305.17926 — Wang et al., not fair evaluators - https://arxiv.org/html/2305.17926 — Wang et al. full text - https://arxiv.or — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/history.md#sources`
- - **Is self-preference about identity or about familiarity?** Panickssery et al. tie it causally to self-recognition (https://arxiv.org/abs/2404.13076). Wataoka et al. tie it to low perplexity, which applies to any familiar text (https://arxiv.org/abs/2410.21819). Chen et al. show that part of the naive signal is real quality, and that the remainder falls sharply with scale (https://arxiv.org/html/2506.02592v1). All three positions stand. None has been ruled out. - **How large is verbosity bias?** Zheng et al. found 91.3% failure for weaker judges (https://arxiv.org/html/2306.05685), and Saito — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/mechanism.md#unresolved-disagreements`
- - https://arxiv.org/abs/2306.05685 — Zheng et al. 2023, Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena - https://arxiv.org/html/2306.05685 — same, full text (tables 2, 4, 5) - https://arxiv.org/abs/2305.17926 — Wang et al. 2023, Large Language Models are not Fair Evaluators - https://arxiv.org/html/2406.07791 — Shi et al. 2024, Judging the Judges: position bias - https://arxiv.org/html/2406.12624 — Thakur et al. 2024, Judging the Judges: alignment and vulnerabilities - https://arxiv.org/abs/2606.19544 — Reliability without Validity (2026 preprint) - https://arxiv.org/html/2310.10076v1 — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/mechanism.md#sources`
- IN: how an LLM judge's own biases distort scores (position, verbosity/length, self-preference); how to calibrate a judge against human labels (Cohen's kappa vs raw agreement vs correlation); the choice of output scale (binary pass/fail vs Likert/direct score vs pairwise); and the specific mitigations for each bias. OUT: the parent domain (eval-driven development in general), sibling concepts (reward models, benchmark design, RAG metrics, agent eval harnesses), and judge fine-tuning as a training topic. — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/practice.md#scope`
- - **Why self-preference happens.** Panickssery et al. attribute it to self-recognition, citing a linear correlation after fine-tuning (https://arxiv.org/abs/2404.13076). Wataoka et al. attribute it to low perplexity, i.e. familiarity (https://arxiv.org/abs/2410.21819). Chen et al. argue that much of it is legitimate, because stronger models often do produce better answers. They also find that stronger models show *more harmful* self-preference when they are wrong, and that CoT reduces it (https://arxiv.org/abs/2504.03846). None of these accounts has displaced the others. - **Binary vs graded s — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/practice.md#unresolved-disagreements`
- 24. Practitioner case for binary pass/fail: a 1–5 score is not actionable (the difference between a 3 and a 2 is unclear), it rarely tracks domain-expert judgment, and it causes "metric sprawl." https://hamel.dev/blog/posts/llm-judge/ 25. Binary outputs also let you apply standard classification metrics (precision, recall, kappa) to the judge directly. https://eugeneyan.com/writing/llm-evaluators/ 26. Decomposing a holistic score into many binary questions (BINEVAL) matched human score distributions better and avoided the ceiling effects of holistic LLM judges. https://arxiv.org/abs/2606.27226 — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/practice.md#output-format-binary-vs-likert-vs-pairwise`
- - **Binary vs Likert.** Practitioners argue that binary is more actionable and that 1–5 scores do not track expert judgment (claim 11). Empirical work finds that 0–5 gives the best human alignment (claim 12) and that fine-grained ordinal scales match ranking (claim 14). No study found here compares binary directly with Likert on actionability, so the two sides measure different things. - **Cause of self-preference.** Three explanations compete: self-recognition (claim 27), familiarity or low perplexity (claim 28), and evaluator uncertainty on items the judge fails (claim 31). Separately, some — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/edge-cases.md#unresolved-disagreements-listed-side-by-side-not-averaged`
- | Pass | Focus | New claims | Total | New-info rate | |---|---|---|---|---| | 0 | Canonical papers: abstracts, dates, named biases | 20 | 20 | 100% | | 1 | Kappa lineage, Scott's π switch, binary-vs-Likert, disconfirming self-preference paper | 14 | 34 | 41% | | 2 | In-paper numbers: Wang conflict rates, G-Eval weighting mechanism, kappa paradox abstracts | 4 (plus numeric refinements of existing claims) | 38 | ~11–15% | — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/history.md#pass-log`
- 17. High percent agreement can hide large score disagreements. Even the best judges stayed well below inter-human agreement and could differ from human scores by up to 5 points. https://arxiv.org/abs/2406.12624 18. Example: llama-3-8b showed 80% raw agreement with humans, but its Cohen's kappa was only 0.62. https://eugeneyan.com/writing/llm-evaluators/ 19. Correlation metrics (Spearman, Kendall) do not correct for chance agreement. They usually run higher than kappa and can overstate judge quality. https://eugeneyan.com/writing/llm-evaluators/ 20. Kappa has its own failure mode. If class prev — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/practice.md#calibration-metrics-kappa-vs-agreement-vs-correlation`
- **D12. What drives position bias?** - The judge model: consistency ranges from 23.8% to 65.0% (C2). - The quality gap between candidates, with prompt-component length mattering little (C7). - Reasoning-trace length, for reasoning models (C9). - C7 and C9 measure different lengths (prompt components vs generated trace), so they do not contradict each other. The three drivers have not been tested together. — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/rabbithole-synthesis.md#disagreements-side-by-side-not-averaged`
- **Saturation:** soft stop after 3 deepening passes, not true saturation. New claims per pass were 15 → 11 → 7 → 3. New-information rate was roughly 42% → 20% → 8%, so it never dropped below 5% twice in a row. One more pass on weighted κ / Krippendorff's α for ordinal judges, and on direct binary-vs-Likert comparisons, would likely still add claims. **Handoffs (out of scope, for concept-family-explorer):** authority, beauty and gender bias in judges; how reward models inherit judge bias; panel/jury aggregation design; perplexity-based familiarity bias as a concept of its own. — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/edge-cases.md#quality-gate`
- 27. Cohen's kappa is κ = (Pr(a) − Pr(e)) / (1 − Pr(e)), where Pr(a) is observed agreement and Pr(e) is agreement expected by chance from the marginals. — https://pmc.ncbi.nlm.nih.gov/articles/PMC3900052/ 28. McHugh argues that the model behind "chance agreement" (raters guessing) is not warranted. She proposes that any kappa below 0.60 means inadequate agreement in health contexts. — https://pmc.ncbi.nlm.nih.gov/articles/PMC3900052/ 29. The common interpretation bands (Landis & Koch: 0.21–0.40 fair, 0.41–0.60 moderate, and so on) rest on opinion and are "by no means universally accepted". — ht — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/mechanism.md#e-agreement-statistics-kappa-and-its-relatives`
- - **G1.** In Arize's spelling-corruption test on a 0–10 scale, GPT-4 scores collapsed toward the extremes (1 or 10). Binary labels separated clean passages from corrupted ones with low variance. [M] https://arize.com/blog-course/numeric-evals-for-llm-as-a-judge/ - **G2.** Arize reran the test in September 2025 on GPT-5-nano, Claude Opus 4, Qwen3-235B and o3. Numeric scores still "drift, flatten, or reverse" across models and scales, while categorical labels stayed stable across runs. [M] https://arize.com/blog/testing-binary-vs-score-llm-evals-on-the-latest-models/ - **G3.** Hamel Husain's cas — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/rabbithole-synthesis.md#g-output-format-binary-vs-likert-vs-pairwise`
- 11. Practitioner argument for binary: nobody knows how to act on a "3 vs 4", and in the author's data, expert judgments did not correlate with 1–5 metrics. Pass/fail forces the expert to state what matters. https://hamel.dev/blog/posts/llm-judge/ 12. Evidence against binary-only: across six benchmarks (MT-Bench, SummEval, STS-B, etc.), a 0–5 scale gave the highest human–LLM alignment when results were pooled across tasks. Switching scales substantially changed human–LLM agreement even when reliability within each rater group was high. https://arxiv.org/abs/2601.03444 13. A pooled reliability f — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/edge-cases.md#binary-vs-likert`
- 38. Direct numeric scores clump. On a 1–5 scale, one value (often 3) dominates, which lowers variance. LLMs also emit whole numbers even when asked for decimals, which produces ties. — https://arxiv.org/html/2303.16634 39. G-Eval's fix reads the probabilities of the score tokens and computes the expected score: score = Σ p(sᵢ)·sᵢ. The result is a continuous, finer-grained value. — https://arxiv.org/html/2303.16634 40. LLM judges show skewed rating distributions and low "inter-sample" agreement. They are also sensitive to prompt changes that humans would consider insignificant. — https://arxiv. — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/mechanism.md#f-output-format-binary-vs-likert-vs-continuous`
- - **F1.** JUDGE-BENCH covers 20 NLP datasets and 11 LLMs. Judge reliability varies with the property being judged, the expertise of the human raters, and whether the text is human-written or model-generated. The authors conclude that every judge must be validated against human labels before use. [M] https://arxiv.org/abs/2406.18403 - **F2.** Percent agreement is high for almost every judge, so under percent agreement judges "are rarely discriminable." [M,H,E] https://arxiv.org/html/2406.12624 - **F3.** Thakur et al. report that "judges with high percent agreement can still assign vastly differ — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/rabbithole-synthesis.md#f-measuring-judge-human-agreement`
- **D7. Are LLM judges as reliable as humans?** - *Parity:* GPT-4–expert agreement is 85% against 81% human–human with ties excluded, and 66–70% against 60–63% with ties included (B11). - *A clear gap:* the best judge scores π 88 against 96.2 for humans (F4). Agreement falls 33–41 points when chance-corrected (F8), and the "contains" matcher shows ranking and scoring diverging (F7). - *The reference is itself biased:* humans are vulnerable to bias too (F17). - These figures use different metrics (percent agreement vs Scott's π vs κ) and cannot be compared directly. — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/rabbithole-synthesis.md#disagreements-side-by-side-not-averaged`
- **The next pass would most likely pay off on these items:** 1. A direct binary-vs-Likert comparison using a chance-corrected statistic (D5 open gap). A lead: the practice report saw a search snippet attributing "binary beats {0, 0.5, 1}" results to AgentJudgeBench (arXiv 2608.26623). The snippet could not be confirmed from the abstract, so the claim is excluded here. 2. Reading the Thakur et al. full text to settle D8 and D9. 3. Reconciling the self-preference-vs-scale tension (D2): DBG vs Chen et al. 2025 under a common metric. 4. Weighted κ and Krippendorff's α for ordinal judges. The edge-c — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/rabbithole-synthesis.md#saturation-verdict`
- The curve is falling. Saturation, defined as two consecutive passes under 5%, was not reached. At least one more pass would likely still pay off, on the open binary-vs-Likert gap and on newer position-bias mechanism work. — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/mechanism.md#pass-log-new-information-rate`

## Facts and statements

- **In scope:** how an LLM judge's verdicts depart systematically from human or ground-truth labels. That covers position, verbosity (length) and self-preference bias, plus familiarity, leniency, and score-distribution skew. Also in scope: how each bias is measured, how judge–human agreement is quantified (percent agreement, Cohen's κ, Scott's π, weighted κ, ICC, TPR/TNR, McDonald's ω), how the output format (binary, Likert/pointwise, pairwise) changes reliability, and the calibration and debiasing methods. — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/rabbithole-synthesis.md#scope`
- **In scope:** how an LLM judge's verdict departs from human or ground-truth labels in a systematic way (position, verbosity, self-preference, familiarity, leniency, score-distribution skew). How each bias is measured. How agreement with humans is quantified (percent agreement, Cohen's kappa, Scott's pi, weighted kappa, ICC). How the output format (binary, Likert, pairwise, pointwise) changes reliability. How judges are calibrated or debiased (order swapping, length regression, probability-weighted scores, juries, human-anchored estimators). — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/mechanism.md#scope`
- Met. The report uses 20+ independent sources across arxiv.org, aclanthology-indexed papers, pmc.ncbi.nlm.nih.gov, sciencedirect.com, en.wikipedia.org, hamel.dev and arize.com. The arxiv.org sources come from distinct author groups. Disconfirming sources were found for verbosity bias (2606.19544), for pure binary advocacy (2601.03444), and for naive self-preference measurement (2506.02592). — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/mechanism.md#quality-gate`
- Met. The report uses 12 independent primary or near-primary sources across 4 hosts (arxiv.org, pubmed.ncbi.nlm.nih.gov, hamel.dev, eugeneyan.com). Disconfirming sources were sought and included: arXiv 2504.03846 against self-preference as pure bias, PubMed 2348207 against kappa as sufficient, and arXiv 2310.05657 against a binary-only default. — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/practice.md#quality-gate`
- Caveats: - The MT-Bench bias figures come from 2023-era judges (GPT-4, Claude-v1, GPT-3.5) and small samples (for example, 23 cases in the verbosity attack). Re-measure them on current judges. - Claims 23 and 26 come from 2026 preprints and were checked at abstract level only. - A search snippet attributed "binary beats {0, 0.5, 1}" results to AgentJudgeBench (arXiv 2608.26623). I could not confirm that from its abstract, so it is excluded. — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/practice.md#quality-gate`
- **Out of scope:** eval-driven development as a practice (the parent), reward modelling and RLHF length exploitation, benchmark design and contamination, agent/trajectory evaluation, eval tooling products, and bias families other than the three named ones (authority, gender, beauty). Those bias families appear here only as evidence that humans share them. — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/rabbithole-synthesis.md#scope`
- **Handoffs for concept-family-explorer:** these siblings surfaced in the reports and were not absorbed here: - reward-model length bias and RLHF length exploitation - agent and trajectory evaluation with LLM judges - benchmark contamination and self-evaluation leakage - authority, beauty and gender bias in judges - jury/panel aggregation design - perplexity-based familiarity bias as a concept of its own - prediction-powered inference for judge-labelled metrics - fine-tuned judge models (Prometheus-style) — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/rabbithole-synthesis.md#saturation-verdict`
- - These arXiv sources are unreviewed 2026 preprints: 2601.03444, 2601.22548, 2604.11581, 2605.06672 (single author), 2606.19544 and 2606.27226. - The full-text numbers were read from arXiv HTML for Zheng, Wang, G-Eval, Thakur, Shi, Saito, Dubois and DBG. Most other papers were checked at abstract level only. - The Feinstein–Cicchetti publisher pages returned 403. Their claims rest on PubMed abstracts and one secondary account (D10). - The MT-Bench bias figures come from 2023-era judges (GPT-4, Claude-v1, GPT-3.5) and small samples, such as 23 cases in the padding attack. They should be re-meas — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/rabbithole-synthesis.md#caveats`
- 35. Humans are biased judges too. Both human and LLM judges are vulnerable to perturbations such as misinformation oversight and authority bias, so a human baseline is not bias-free. https://arxiv.org/abs/2402.10669 36. Frontier judges still show significant biases on specific tasks across 12 bias types (CALM framework). Good overall performance does not mean a judge is unbiased on every task. https://arxiv.org/abs/2410.02736 — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/edge-cases.md#cross-cutting`
- - Reward-model length bias and RLHF length exploitation - Agent and trajectory evaluation with LLM judges - Benchmark contamination and self-evaluation leakage — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/history.md#handoffs-not-researched-siblings-for-concept-family-explorer`
- 31. The CALM framework catalogs 12 distinct LLM-judge biases and finds that significant biases persist in specific applications even for strong models. https://arxiv.org/abs/2410.02736 — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/practice.md#taxonomy-breadth`
- Handoffs for concept-family-explorer (siblings surfaced, not researched): reward-model bias, judge ensembles/juries, prediction-powered inference for judge-labeled metrics, fine-tuned judge models (Prometheus-style). — source: `~/.global-ai-hub/research-runs/frontier-2026-09-25/llm-as-judge-bias-calibration-kappa-binary-vs-likert-position-verbosity-self-preference/reports/practice.md#depth-passes`

## Related concepts

- bias — is a part of LLM-as-judge bias & calibration (kappa, binary-vs-Likert, position/verbosity/self-preference)
- self-preference — is a part of LLM-as-judge bias & calibration (kappa, binary-vs-Likert, position/verbosity/self-preference)
- position — is a part of LLM-as-judge bias & calibration (kappa, binary-vs-Likert, position/verbosity/self-preference)
- verbosity — is a part of LLM-as-judge bias & calibration (kappa, binary-vs-Likert, position/verbosity/self-preference)
- kappa — is a part of LLM-as-judge bias & calibration (kappa, binary-vs-Likert, position/verbosity/self-preference)
- calibration — is a part of LLM-as-judge bias & calibration (kappa, binary-vs-Likert, position/verbosity/self-preference)
- binary-vs-Likert — is a part of LLM-as-judge bias & calibration (kappa, binary-vs-Likert, position/verbosity/self-preference)
- LLM-as-judge — is a part of LLM-as-judge bias & calibration (kappa, binary-vs-Likert, position/verbosity/self-preference)
