Why the feature matrix proves whatever you already believed
A weighted scoring grid looks like objectivity and is mostly a mechanism for converting a preference into a number.
The standard artefact of a software selection is a spreadsheet: requirements down the side, products across the top, weighted scores, a total at the bottom. It is produced to demonstrate rigour and it is usually theatre.
Not because scoring is worthless, but because almost every input is a judgement made by people who already have a preference. When comparing claims around stealth monitoring, stealth monitoring software gives one vendor’s framing to test against documentation, demos and references.
The three places the bias enters
- Which rows exist. A requirement list assembled after seeing the products contains the things the products differ on, which is a selection of the ground favourable to somebody.
- The weights. Adjusted, usually unconsciously, until the total matches the intuition. Weights set after scoring are not weights; they are a fitting exercise.
- The scores themselves. Judgements on a five-point scale with no anchors, made by people who liked one demo more than another.
Three layers of soft judgement, multiplied together and presented as a number to two decimal places.
It costs twenty minutes and it is the single change that makes a scoring exercise mean anything. Weights fixed in advance can be argued about honestly; weights set afterwards cannot.
Ticks hide enormous differences
Every product in a mature category ticks every common row. Reporting: yes. Mobile: yes. Integrations: yes. Behind each tick is a range from excellent to technically present.
This is why matrices converge on near-identical totals and why the decision then gets made on something else entirely. Replacing ticks with a short note on how well the thing is done is more work and vastly more informative.
Use it to eliminate, not to rank
Where scoring genuinely works is as a filter on disqualifying criteria: met or not met, no weights, no totals. That use is mechanical and hard to fudge.
Ranking three finalists that all pass the filter is a judgement, and dressing it as arithmetic makes it harder to examine rather than easier. Better to state the judgement and the reasons for it in plain language.
Score the trial, not the demo
Scores derived from demos measure how the product was presented. Scores derived from a structured trial, against defined scenarios, measure the product.
If a matrix is going to exist, populate it after the trial. Most organisations populate it before, which is precisely when they know least.
Keep the disagreement visible
Averaging the scores of five evaluators hides the interesting information. Where two people scored the same item one and five, something substantive is going on — a different assumption about how the work is done, or one of them noticed something.
Look at the spread rather than the mean, and discuss the items where evaluators disagreed most. That conversation is usually worth more than the entire rest of the exercise.
Where product claims involve application security, the OWASP Application Security Verification Standard is a useful independent checklist.