Every L&D function reports on training and most of what gets reported does not answer the question the executive actually asked, which is whether the workforce can now do something it could not do before.
The gap is not laziness. Measuring that honestly is harder than the industry admits, and a lot of what is sold as measurement is theatre.
The four things usually reported, and what each one is worth
| Metric | What it measures | What it does not |
|---|---|---|
| Enrollment | A procurement decision | Anything about the workforce |
| Completion rate | Whether the program was finishable | Whether anyone learned |
| Satisfaction score | How the experience felt | Learning. Sometimes inversely |
| Assessment score | Recall under test conditions | Performance on the job |
Each row is a step closer to the real question and none of them arrives.
Completion is worth reporting because an unfinished program cannot have worked, so it is a necessary condition. Treating it as the outcome is the error.
Completion is a floor, not a result. It tells you the program was possible to finish. It tells you nothing about what the finishing produced.
Why satisfaction scores are worse than useless
The end-of-course rating is the most collected metric in corporate learning and among the least informative.
Research on student evaluations and on training reaction measures has repeatedly found weak correlation between how highly learners rate an experience and how much they actually learned. In some studies the relationship runs the wrong way: instruction that produces better long-term retention is rated lower, because it is harder.
The mechanism is straightforward once stated. Effortful retrieval, spacing and difficulty all improve retention and all feel worse than fluent, well-delivered restudy that produces a comfortable sense of understanding. Learners rate the feeling, and the feeling is generated by fluency rather than by learning.
So a program optimised on satisfaction scores drifts toward being pleasant. That is not a neutral outcome. It actively selects against the design features that work.
Satisfaction data is worth collecting for one thing only: detecting operational problems. If a cohort reports the platform crashes or the timing is impossible, that is real information. As a measure of effectiveness it should not appear in a report to an executive.
The Kirkpatrick problem
Most L&D functions use some version of the four-level model: reaction, learning, behaviour, results.
The framework is useful for organising thinking and it has a structural problem in practice, which is that effort and rigour fall off sharply as you climb it. Level one is collected universally because it is trivial. Level four is claimed frequently and measured properly almost never.
The result is an evaluation culture that reports heavily on the level that matters least, gestures at the level that matters most, and skips the middle where the honest work is.
Level three, behaviour change on the job, is the most useful and most neglected. It is harder than a test and far easier than isolating business results, and it is where an evaluation effort should concentrate.
What can be measured honestly
Four things, in ascending order of value and difficulty.
Assessed competence against a pre-defined standard. This requires the standard to exist before the program starts, which is why scoping and measurement are the same conversation. If the standard says a graduate can commission a PLC on this equipment class and resolve the common fault modes unsupervised, then the assessment is watching someone do that, not a multiple-choice paper about it.
Retention at a delay. Assess again at three months. This is the single most underused measurement in corporate learning and it is cheap. It also distinguishes a program that produced durable capability from one that produced a good week. Expect the second assessment to be worse and treat the size of the drop as the finding.
Observed behaviour change. Structured observation on the job against specific behaviours, by supervisors who know what they are looking for. Subjective, and considerably better than a test score because it measures the thing you actually bought.
Operational indicators tied closely to the trained capability. Not company revenue. Something proximate: first-time-fix rate for a maintenance program, changeover time for a lean program, defect escape rate for a quality program, mean time to resolve for a support program. The closer the indicator sits to the trained behaviour, the more attributable the movement.
Assessment inside the engine
Astra Trainer's format produces assessment data as a by-product rather than as a separate exercise. Practice questions sit inside every lesson, quizzes follow each topic, and each course ends with a ten-question final exam, so mistakes surface during training rather than on the job.
Passing the final exam produces a verified certificate in the learner's name, and certified learners join an expert network answering other people's questions, which is itself a usable signal about who has genuinely retained the material.
That gives you levels one and two of the model in the platform. Levels three and four still require observation and operational data from your side of the arrangement, and any provider claiming to deliver those from a learning platform alone is overstating what a platform can see. You can see how programs are structured here.
The attribution problem, stated plainly
The reason honest measurement stops short of a clean ROI figure is that training is never the only thing that changed.
Over the six months of a program, the organisation also hired, lost people, changed a process, bought equipment, shifted a product mix, and experienced whatever the market did. If throughput improved, several of those are candidate explanations and the training is one.
Establishing that training caused a business outcome requires either a controlled comparison or a set of assumptions substantial enough that the answer is mostly determined by the assumptions.
Which is what most published training ROI figures are. Take a business improvement, assume a share of it was caused by training, divide by cost, report a multiple. The assumption is doing the work and it is rarely stated.
What to ask when someone presents a training ROI number. What was the comparison group? What else changed in that period? What share of the improvement was assumed rather than measured, and where does that assumption come from? If the answer to the first is "none", the number is an estimate wearing the clothes of a measurement. That does not make it useless, but it should be labelled.
There is a respectable position here: for many programs, precise financial attribution is not achievable and the right response is to measure competence and behaviour rigorously, report operational indicators with appropriate caveats, and decline to invent the rest.
An L&D function that says "we can demonstrate capability gained and behaviour changed, and we can show these operational indicators moved, and we cannot isolate the financial contribution" is being more credible than one presenting a multiple to two decimal places. Executives who work with data recognise the difference.
Designing a program you can evaluate
Five decisions, all made before the program starts, that determine whether evaluation is possible at all.
Write the competence standard first, in verifiable terms. Everything else depends on this.
Baseline before you train. Assess the cohort against the standard at the start. Without it you cannot distinguish what the program produced from what people already knew, and the temptation afterwards is to attribute all of it.
Stagger the rollout. This is the highest-value trick available. If a hundred people are being trained, train fifty now and fifty in four months. You have a comparison group for free, the second group is not disadvantaged because everyone gets trained, and the operational comparison between the two is the closest thing to real evidence you will get without running an experiment.
Choose the operational indicator in advance and write down what movement would count. Picking the indicator afterwards, from whatever moved, is how honest teams accidentally produce misleading reports.
Book the follow-up assessment at three months into the plan, because it will not happen otherwise.
What to report upward
A short structure that survives scrutiny.
What the program was for. The roles, the headcount, the standard.
What people can now do, assessed against that standard, with the baseline alongside it.
What held at three months, including the drop.
What changed on the job, from observation, with its subjectivity acknowledged.
What moved operationally, with the staggered comparison if you have one and a clear statement of what else changed if you do not.
What it cost, including the productive time, which is usually the largest line and usually missing.
What you still cannot say. Include this deliberately. It is the item that makes the rest believable.
What to take from this
Completion is a floor. Satisfaction is close to noise and sometimes points the wrong way, because effective learning feels harder than ineffective learning.
Assessed competence against a standard written beforehand is the honest core of the measurement, and a follow-up at three months is the cheapest high-value addition available.
Stagger the rollout. It gives you a comparison group at no cost and it is the difference between evidence and assertion.
Business-outcome attribution is genuinely hard and most published ROI multiples are assumptions presented as measurements.
And reporting what you cannot conclude makes everything you do conclude more credible, not less.
Are satisfaction scores worth collecting?
Only for detecting operational problems such as platform failures or impossible scheduling. As a measure of effectiveness they are weak and sometimes inverted, because difficulty improves retention and lowers ratings.
What is the single best measure of training effectiveness?
Assessed competence against a standard defined before the program started, re-assessed at three months, ideally supported by structured observation on the job.
Can training ROI be calculated?
Rarely with real precision. Everything else in the organisation changes during a program, so attribution requires either a comparison group or large assumptions. Staggering the rollout gives you the comparison group cheaply.
What should we baseline?
The cohort's competence against the standard, before training starts. Without a baseline, you cannot separate what the program produced from what people already knew.
What can a learning platform actually measure?
Completion and assessed knowledge, which Astra Trainer produces through in-lesson practice, topic quizzes and a final exam per course. Behaviour change and operational outcomes need observation and data from your side. You can see how programs are structured here.
