How to read validity documentation
A practical walk-through of norming, reliability and validity, so you can assess a test without being a psychometrician.
Norming
A test score means nothing on its own. It only acquires meaning when set against a norm group, the population the test was calibrated on. So the first question to put to any vendor is who the norm group consists of, how large it is, and when the data was collected.
Three things are worth checking. Size first: a norm group of a few hundred people carries wider margins of uncertainty than one of several thousand. Then composition. If the norm group is largely students but the test is used for experienced specialists, the candidate is being compared against the wrong reference frame. Finally age. Labour markets change, and a norm group collected fifteen years ago does not necessarily describe today's applicant pool.
Local norming deserves particular attention in the Danish market. Many international tests are normed on English-speaking populations and translated afterwards. That is not a problem in itself, but the manual should state whether a Danish or Nordic norm group exists and how large it is. If nothing is stated, the answer is usually no.
Reliability
Reliability is about how stably a test measures. If the same person takes it twice a few weeks apart, how much does the result move? And do the items belonging to one scale actually measure the same thing?
Manuals typically report two figures. Internal consistency, often as Cronbach's alpha, shows how closely the items within a scale hang together. Test-retest reliability shows stability over time. Both are reported on a scale from 0 to 1, and higher is better, but a high figure is not automatically a mark of quality. A scale can reach very high internal consistency by asking essentially the same question ten times, which makes the scale narrow rather than strong.
The practical advice is to look at reliability scale by scale, not at an average for the whole instrument. A test can have solid figures on its main scales and weak ones on the subscales that the report nonetheless uses to draw conclusions about the candidate.
Validity
Validity is the question that actually matters: does the test measure what it claims to measure, and does the result say anything about performance on the job afterwards?
Predictive validity is the strongest form of evidence. It requires the vendor to have followed a group of candidates from the point of testing and compared their scores against a later measure of job performance. Studies of that kind are expensive and slow, which is why they do not exist for every test. Construct validity is easier to document and shows that the test correlates with other established measures of the same construct.
When reading the documentation, three questions are enough to separate the well-evidenced from the unevidenced. Do validity studies exist for this specific test, or does the manual simply point to research on the construct in general? Were the studies conducted on independent samples, or solely by the vendor? And do the job types in those studies resemble the roles you are actually hiring for?
If the answers are missing, that does not make it a bad test. But it does make it one where you carry the risk that it works.
Source: Based on publicly available test manuals and industry standards for psychometric documentation.