<!-- llms-explorer concept facts · https://llms-explorer.com/tree/learning-measurement-training-evaluation/ · pack 2026-09-08 · ~3429 tokens -->

# Learning Measurement & Training Evaluation

> Reference for measuring training and enablement program effectiveness across the full evaluation stack — from post-workshop smile sheets to executive ROI reports and cross-system xAPI analytics.

Parent: [Technical Instruction & Engineering Education](https://llms-explorer.com/tree/technical-instruction-engineering-education/) · 19 facets · 40 facts · page: https://llms-explorer.com/tree/learning-measurement-training-evaluation/

## Learning Measurement & Training Evaluation

- Reference for measuring training and enablement program effectiveness across the full evaluation stack - from post-workshop smile sheets to executive ROI reports and cross-system xAPI analytics. — [source](https://llms-explorer.com/sources/mdb-context-hub/learning-measurement-evaluation/#learning-measurement-training-evaluation)
- See references/learning-measurement-context.md for the full deep-reference with worked examples, decision tables, and vendor comparisons. — [source](https://llms-explorer.com/sources/mdb-context-hub/learning-measurement-evaluation/#learning-measurement-training-evaluation)

## Evaluation Framework Overview

- Three frameworks dominate training program evaluation. They are complementary, not competing: — [source](https://llms-explorer.com/sources/mdb-context-hub/learning-measurement-evaluation/#evaluation-framework-overview)
- When to use each: Kirkpatrick levels 1–4 apply to any formal training; New World Kirkpatrick adds the pre-design planning discipline; Phillips Level 5 is appropriate for 5–10% of programs - those that are high-cost, strategically critical, and directly tied to measurable business KPIs. — [source](https://llms-explorer.com/sources/mdb-context-hub/learning-measurement-evaluation/#evaluation-framework-overview)

## Level 1: Reaction

- Measures whether participants found training favorable, engaging, and relevant. Relevance is the most predictive sub-component. — [source](https://llms-explorer.com/sources/mdb-context-hub/learning-measurement-evaluation/#level-1-reaction)
- Limitation: A 2009 study (n=335) found no statistically significant correlation between Level 1 scores and Level 3 behavior change. Alliger & Janak (1989) found the Level 1→Level 2 causal link produced only r=.23. — [source](https://llms-explorer.com/sources/mdb-context-hub/learning-measurement-evaluation/#level-1-reaction)

## Level 2: Learning

- Measures acquisition of knowledge, skills, attitude, confidence, and commitment (5 components per the New World Model). — [source](https://llms-explorer.com/sources/mdb-context-hub/learning-measurement-evaluation/#level-2-learning)

## Level 3: Behavior

- Measures whether participants apply what they learned on the job. Most organizations skip formal Level 3 evaluation - Kennedy et al. (2013) found 40%+ report it is not required by management. — [source](https://llms-explorer.com/sources/mdb-context-hub/learning-measurement-evaluation/#level-3-behavior)

## Level 4: Results

- Measures whether targeted organizational outcomes occurred. Requires pre-established KPI baselines and isolation methodology. — [source](https://llms-explorer.com/sources/mdb-context-hub/learning-measurement-evaluation/#level-4-results)

## Shared Accountability Model

- L1/L2 = L&D accountability; L3/L4 = shared accountability between L&D, managers, and the business. — [source](https://llms-explorer.com/sources/mdb-context-hub/learning-measurement-evaluation/#shared-accountability-model)

## Required Drivers (Level 3 Addition)

- Post-training reinforcement systems management must provide: coaching, job aids, work review, recognition. [TENTATIVE: ~85% vs ~15% application figures are from Kirkpatrick-affiliated sources - treat as directional.] — [source](https://llms-explorer.com/sources/mdb-context-hub/learning-measurement-evaluation/#required-drivers-level-3-addition)

## Leading Indicators (Level 4 Bridge)

- Short-term observations signaling whether critical behaviors are on track to produce results. — [source](https://llms-explorer.com/sources/mdb-context-hub/learning-measurement-evaluation/#leading-indicators-level-4-bridge)

## Backward Design Mandate

- Plan evaluation in reverse: L4 KPI → L3 behaviors → L2 objectives → L1 design. — [source](https://llms-explorer.com/sources/mdb-context-hub/learning-measurement-evaluation/#backward-design-mandate)

## Phillips ROI Methodology (Level 5)

- Isolation methods (descending reliability): control group → trend-line → forecasting model → manager estimates → participant estimates. — [source](https://llms-explorer.com/sources/mdb-context-hub/learning-measurement-evaluation/#phillips-roi-methodology-level-5)
- Apply to ~5–10% of programs: high-cost, strategically critical, tied to measurable KPI, with a viable isolation method. — [source](https://llms-explorer.com/sources/mdb-context-hub/learning-measurement-evaluation/#phillips-roi-methodology-level-5)

## xAPI, SCORM, and cmi5

- xAPI adoption ~17% (verified-as-of: 2026-06-16) despite 10+ years. SCORM still dominant at 81.7%. — [source](https://llms-explorer.com/sources/mdb-context-hub/learning-measurement-evaluation/#xapi-scorm-and-cmi5)

## Training Transfer

- Baldwin & Ford (1988): transfer depends on trainee characteristics, training design, and work environment. Supervisor support is the single strongest driver (ρ=.51, Human Factors 2019). — [source](https://llms-explorer.com/sources/mdb-context-hub/learning-measurement-evaluation/#training-transfer)
- LTSI (Holton, Bates & Ruona, 2000): 16-factor validated instrument for diagnosing transfer barriers. — [source](https://llms-explorer.com/sources/mdb-context-hub/learning-measurement-evaluation/#training-transfer)

## L&D Dashboard Design

- Three audiences: Executive (ROI, time-to-proficiency), Manager (behavior change, team completion), L&D ops (engagement analytics). — [source](https://llms-explorer.com/sources/mdb-context-hub/learning-measurement-evaluation/#ld-dashboard-design)
- Top metrics: time-to-productivity, 30/60/90-day retention, manager reinforcement score, application rate, completion rate (process only - not performance). — [source](https://llms-explorer.com/sources/mdb-context-hub/learning-measurement-evaluation/#ld-dashboard-design)

## Evaluation Edge Cases

- Skip evaluation for cohorts under ~20 or one-time deliveries where overhead exceeds value. — [source](https://llms-explorer.com/sources/mdb-context-hub/learning-measurement-evaluation/#evaluation-edge-cases)
- Retrospective evaluation: use trend-line data, untrained comparison cohort, or SME estimates with documented confidence reduction. — [source](https://llms-explorer.com/sources/mdb-context-hub/learning-measurement-evaluation/#evaluation-edge-cases)
- ROI report structure: Executive summary → Methodology → Financial detail → Intangibles → Sensitivity analysis. — [source](https://llms-explorer.com/sources/mdb-context-hub/learning-measurement-evaluation/#evaluation-edge-cases)
- xAPI/LRS privacy: anonymize actor IDs for PII; GDPR applies in EU; FERPA applies to educational institutions. — [source](https://llms-explorer.com/sources/mdb-context-hub/learning-measurement-evaluation/#evaluation-edge-cases)

## Where this helps

- Deciding how much evaluation rigor a given training program deserves — a one-off onboarding session needs different measurement than a multi-week technical certification tied to revenue. — [source](https://llms-explorer.com/tree/learning-measurement-training-evaluation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Building an L&D dashboard that actually answers what each audience needs (executive ROI/time-to-proficiency, manager behavior change/completion, L&D ops engagement analytics) instead of one generic completion-rate report. — [source](https://llms-explorer.com/tree/learning-measurement-training-evaluation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Defending a training budget or enablement investment to leadership with a Phillips ROI case, when the program is high-cost and tied to a measurable business KPI. — [source](https://llms-explorer.com/tree/learning-measurement-training-evaluation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Diagnosing why training "didn't stick" on the job by tracing the gap through Kirkpatrick's levels — distinguishing a Level 1 satisfaction problem from a Level 3 transfer problem. — [source](https://llms-explorer.com/tree/learning-measurement-training-evaluation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*

## How to apply this

- Use backward design: start from the Level 4 business KPI you actually want to move, work back to the Level 3 behaviors that would produce it, then the Level 2 objectives and Level 1 design that build those behaviors. — [source](https://llms-explorer.com/tree/learning-measurement-training-evaluation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Reserve Phillips ROI (Level 5) isolation analysis for the roughly 5-10% of programs that are high-cost, strategically critical, and tied to a measurable KPI — applying it to every program wastes the overhead the framework itself warns against. — [source](https://llms-explorer.com/tree/learning-measurement-training-evaluation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Pick an isolation method by its reliability ranking — control group first, then trend-line, then forecasting model, then manager estimates, then participant estimates — and document which one was used and why. — [source](https://llms-explorer.com/tree/learning-measurement-training-evaluation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Build the Required Drivers (coaching, job aids, work review, recognition) into the post-training plan up front, since Level 3 behavior change depends on management reinforcement as much as on the training itself. — [source](https://llms-explorer.com/tree/learning-measurement-training-evaluation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*

## Antipatterns

- Treating a high Level 1 "smile sheet" score as proof the training worked — Alliger & Janak (1989) found only a weak r=.23 correlation between Level 1 and Level 2, and a 2009 study found no significant Level 1-to-Level 3 correlation at all. — [source](https://llms-explorer.com/tree/learning-measurement-training-evaluation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Claiming a Level 4 business result without first establishing Level 3 behavior change — results without a plausible behavior-change mechanism are not attributable to the training. — [source](https://llms-explorer.com/tree/learning-measurement-training-evaluation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Running full Phillips ROI methodology on every program regardless of cost or strategic weight, instead of reserving it for the roughly 5-10% of programs where the isolation overhead is justified. — [source](https://llms-explorer.com/tree/learning-measurement-training-evaluation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Reporting completion rate as if it were a performance metric — it is a process metric, and conflating the two hides whether the training actually changed behavior. — [source](https://llms-explorer.com/tree/learning-measurement-training-evaluation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*

## Known issues

- Level 1 reaction scores are a weak predictor of downstream learning and behavior change, so relying on them alone as an evaluation signal is misleading despite being the easiest data to collect. — [source](https://llms-explorer.com/tree/learning-measurement-training-evaluation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- Most organizations never formally evaluate Level 3 — Kennedy et al. (2013) found 40%+ report management does not require it — meaning the levels most tied to business impact are also the least measured in practice. — [source](https://llms-explorer.com/tree/learning-measurement-training-evaluation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- The roughly 85% vs 15% application-rate figures commonly cited for the Required Drivers model come from Kirkpatrick-affiliated sources and should be treated as directional, not as a validated statistic. — [source](https://llms-explorer.com/tree/learning-measurement-training-evaluation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*
- xAPI adoption remains around 17% after more than a decade, with SCORM still dominant at roughly 82%, so cross-system learning-record analytics built on xAPI often cannot assume broad tooling support yet. — [source](https://llms-explorer.com/tree/learning-measurement-training-evaluation/) *(AI-suggested, synthesized from this pack's existing facts — not extracted from a source document.)*

## Context files

- [Learning Measurement & Training Evaluation](https://llms-explorer.com/downloads/sources/mdb-context-hub/learning-measurement-evaluation.md)
