Eval-Driven Development for LLM Apps

Parent: AI Agent Ecosystems · researched 2026-06-01T02:26:42.058Z· 18 sources · 10 concepts · skill ai-agent-engineering

AI & agent-engineering family ROUTER. Split into: ai-agents-orchestration (agent frameworks, multi-agent, memory, planning, guardrails, coding/GUI agents, autonomous loops, eval); ai-rag-retrieval (RA

ai-agent-engineering

Children

Frontier under this node: Agent/trajectory evaluation, CI/CD eval gating & regression suites, Criteria drift & Who-Validates-the-Validators (EvalGen, SPADE), Error analysis & qualitative coding (open/axial coding of traces), Eval levels (assertion / LLM-as-judge / human; offline vs online flywheel), Golden datasets & synthetic eval-data (silver-to-gold), LLM-as-judge bias & calibration (kappa, binary-vs-Likert, position/verbosity/self-preference), Metric design (rubric, pairwise, pass@k vs pass^k), The Three Gulfs (Specification/Generalization/Comprehension), Tooling (promptfoo, DeepEval, Ragas, Braintrust, LangSmith)

← the whole tree · 3D view· how to read this page