Eval-Driven Development for LLM Apps
Parent: AI Agent Ecosystems · researched 2026-06-01T02:26:42.058Z· 18 sources · 10 concepts · skill ai-agent-engineering
AI & agent-engineering family ROUTER. Split into: ai-agents-orchestration (agent frameworks, multi-agent, memory, planning, guardrails, coding/GUI agents, autonomous loops, eval); ai-rag-retrieval (RA
ai-agent-engineering
- AI & agent-engineering family ROUTER. Split into: ai-agents-orchestration (agent frameworks, multi-agent, memory, planning, guardrails, coding/GUI agents, autonomous loops, eval); ai-rag-retrieval (RAG, iterative retrieval, vector/graph datastores); ai-llm-model-layer (training, fine-tuning, alignment/RLHF, compression, inference serving, transformer/multimodal architecture, model selection, observability); ai-mcp-sdk-prompting (MCP servers/builder, Anthropic SDK, prompt engineering, context engineering, LLM frameworks, tool-search, prompt lookup). Route to the matching sub-hub. [source]
- This hub routes to on-demand reference files under references/. See each spoke for depth. [source]
Children
- Error analysis & qualitative coding (open/axial coding of traces) (frontier)
- The Three Gulfs (Specification/Generalization/Comprehension) (frontier)
- Criteria drift & Who-Validates-the-Validators (EvalGen, SPADE) (frontier)
- Eval levels (assertion / LLM-as-judge / human; offline vs online flywheel) (frontier)
- LLM-as-judge bias & calibration (kappa, binary-vs-Likert, position/verbosity/self-preference) (frontier)
- Golden datasets & synthetic eval-data (silver-to-gold) (frontier)
- Metric design (rubric, pairwise, pass@k vs pass^k) (frontier)
- CI/CD eval gating & regression suites (frontier)
- Agent/trajectory evaluation (frontier)
- Tooling (promptfoo, DeepEval, Ragas, Braintrust, LangSmith) (frontier)
Frontier under this node: Agent/trajectory evaluation, CI/CD eval gating & regression suites, Criteria drift & Who-Validates-the-Validators (EvalGen, SPADE), Error analysis & qualitative coding (open/axial coding of traces), Eval levels (assertion / LLM-as-judge / human; offline vs online flywheel), Golden datasets & synthetic eval-data (silver-to-gold), LLM-as-judge bias & calibration (kappa, binary-vs-Likert, position/verbosity/self-preference), Metric design (rubric, pairwise, pass@k vs pass^k), The Three Gulfs (Specification/Generalization/Comprehension), Tooling (promptfoo, DeepEval, Ragas, Braintrust, LangSmith)