Measuring whether it actually worked
The success criteria were written at the start precisely so this question could be answered. Almost nobody goes back and answers it.
Software purchases are evaluated exhaustively before the decision and almost never after it. Once the thing is in, attention moves on, and whether it delivered what was claimed becomes a matter of impression.
This is why organisations accumulate subscriptions nobody can justify and repeat the same purchasing mistakes. For a concrete example of time-management reading after purchase, this page can be used to identify the behaviours and measures that rollout should actually change.
You need a before figure
The measurement that matters is a comparison, which means capturing the baseline before rollout. Once the old process is gone, it cannot be reconstructed.
Whatever the success criteria were — hours spent, errors per month, days to close, response time — measure them in the weeks before go-live. This is the step that makes the whole exercise possible and it takes a fortnight of ordinary record-keeping.
Without a before figure, the after figure is a number with nothing to compare it to, and the review becomes a discussion of how people feel about the change.
Measure the same thing, the same way
The comparison only holds if the measurement is identical. Changing the definition between before and after — a different sample, a different period, a different way of counting — produces a difference that is an artefact.
Write the method down when you take the baseline, and follow it exactly when you repeat it.
Wait long enough
Measuring a month after go-live measures the disruption, not the outcome. Everything is worse during a transition, and a review at that point will conclude the purchase failed.
Six months is a reasonable point for most changes: long enough for the process to settle, soon enough to act. A twelve-month review is also worth having, because some effects only appear across a full cycle.
Measure use as well as outcome
Two figures, and they answer different questions. Is it being used — how many of the intended users, how often, for what proportion of the intended work? And did the outcome move?
Low use with an unchanged outcome is an adoption problem. High use with an unchanged outcome is a more uncomfortable finding: the product works and the benefit was not there, which is a lesson about the original business case.
Look for the shadow processes
The most informative question in a post-implementation review is what people still do outside the system. The private spreadsheet, the message thread where decisions actually happen, the manual step nobody mentioned.
Each one is a gap between the design and the work, and each is degrading whatever the system reports. Asking directly and without blame produces honest answers; most people are happy to explain the workaround they built.
Write the answer down
Record what was measured, what changed, and what you would do differently. It informs the next purchase, it supports or challenges the next renewal, and it is the only mechanism by which an organisation gets better at buying software.
For operational controls that often change during adoption, NCSC 10 Steps to Cyber Security provides a public baseline.