Eval-Driven Development for LLM Applications (2024-2026)

Parent: Global AI Hub Research Corpus · researched 2026-05-31· 1 source · 0 concepts

Research date: 2026-05-31

Eval-Driven Development for LLM Applications (2024-2026)

1. Eval-Driven Development as a discipline

The "Three Gulfs" framing (Shankar & Husain)

Building evals FROM error analysis (qualitative coding)

2. Eval levels and the offline/online split

Known biases

Judge calibration / validation against humans

4. Tooling landscape

5. Metric design

6. CI/CD, regression testing, eval-gated deploys

7. Agent / trajectory evaluation

8. Anti-patterns

9. Sources

10. Contested / low-confidence areas

Scope-overlap notes (for concept-tree placement)

Children

← the whole tree · 3D view· how to read this page