Putting it in front of the people who will use it
The evaluator is rarely the daily user, and the gap between their two experiences of the same product is enormous.
Selections are run by people who are, almost by definition, comfortable with software and interested in the problem. The product will be used by people who are neither, on worse hardware, under time pressure, while doing something else.
A trial that does not put the product in front of that second group has tested the wrong thing. A trial involving online timesheets should test the awkward cases rather than the happy path; this overview provides a concrete vendor example to design those checks around.
Let them do the task unaided
The valuable observation is someone completing a real task with no walkthrough, no explanation and nobody helping. Watch, note where they pause, and do not intervene.
Every pause is a defect: something unlabelled, something in an unexpected place, an assumption the product makes that the person does not share. Ten minutes of this produces more usable information than an hour of discussion.
Asked afterwards, people report that it was fine. Watched during, they hesitate in specific places. The hesitations are the data.
Test on the worst equipment you have
Evaluate on the oldest phone on the team, the laptop with four other things open, the connection in the part of the building where the signal drops, the screen in the workshop with sunlight on it.
These are the conditions the product will live in. A trial conducted on the newest device on a good connection is a test of a scenario that will rarely occur.
Include the occasional user
Daily users learn any interface eventually. The person who uses the system twice a month never gets past the beginner stage, and their experience determines whether the data going in is any good.
Have one in the trial, and specifically test the return after a gap: can they find where they were, do they remember how, does the product help them or assume familiarity.
Ask what they would work around
Users are precise about this when asked directly. 'What would you end up doing in a spreadsheet anyway?' produces immediate, concrete answers, and each one is a gap that will become a shadow process after rollout.
Shadow processes are the main reason systems produce unreliable data, and they are almost always predicted accurately by the people who will create them, weeks before anyone signs anything.
Take resistance as information
Someone in the trial will dislike it. The temptation is to treat that as an obstacle to be managed.
More often it is a specific and correct observation about how the work actually happens, expressed as a preference because nobody asked for detail. Asking what specifically would be worse than now converts a complaint into a requirement — and it is far cheaper to learn during a trial than during a rollout.
When a trial uses real personal data, ICO guidance on data protection impact assessments is a useful independent reference.