Last updated
Accuracy is the wrong word for one thing and the right word for three. A test can be solid on two of them and hopeless on the third, which is how a result ends up feeling right and meaning nothing.
A thermometer is accurate because there is a real temperature to be right about. Personality has no equivalent: there is no true extraversion sitting in somebody waiting to be read off, so there is nothing for a number to match.
What can be asked instead is narrower and more answerable. Do the questions come from somewhere defensible. Do they hold together. Is there anything to compare an answer against. Those three are what separate a serious instrument from a quiz, and any of them can be checked before trusting a result.
This is the cheapest thing to check and the most often skipped. A question set is either published, so that anybody can inspect it, or it is private and the reader is asked to take it on trust.
Public item pools exist precisely so this is answerable. Items that have been used and re-used by researchers for decades come with a paper trail: what they were written to measure, how they behaved, and what got dropped. A question written last month by a marketing team has none of that, and no way for anybody outside to tell the difference.
The word to look for is not scientific. It is the name of the item pool, and a test that will not say is answering the question by declining it.
A scale is a group of questions that are supposed to be getting at one thing. Whether they are is measurable, and the usual measure is how consistently somebody who agrees with one of them agrees with the rest.
This is where a great many quizzes quietly fail. Five questions that sound related to the person writing them can turn out to have almost nothing in common in the answers, in which case adding them up produces a number about nothing.
It is also the figure most often borrowed rather than earned. A reliability figure belongs to one specific set of questions. Quoting the figure from a published instrument while asking different questions is a claim about somebody else's work.
This is the one that decides what a result can say, and it is almost never mentioned.
A score of 34 means nothing by itself. It means something once it can be placed among the answers of many other people who took the same questions. Then the position is real: a percentile says where somebody sits in a measured distribution rather than on an invented curve.
Without that sample there is still an honest reading available, and it is a smaller one. The result can describe how somebody answered, which of several things sounded most like them, or where they landed on the answer scale itself. What it cannot do is say where they sit relative to anybody, and a test doing that without a sample behind it is making the number up.
Most of what is sold as a personality test has no reference sample at all. That does not make those tests worthless. It makes the sentence about being in the top ten per cent worthless.
A well-built instrument measures how somebody answered on the day they answered. That is a real fact and it is a narrow one, and there are several ordinary reasons the description can miss.
A hard week moves some scales genuinely. Answering as the person somebody is at work produces a different result from answering as the person they are at home. And self-report can only see what a person can see about themselves, which on some questions is not much.
The useful response to a result that does not fit is not to discard the instrument or to accept the label. It is to check whether the description is wrong or merely unwelcome, which only the reader can do.
None of these needs a background in statistics, and any instrument worth taking answers all four in public.
Or see everything at Learn.
Every assessment here is free to take and gives a real result, with the sourcing for it set out in public. The Big Five is the usual place to start.
50 questions · about 5 min · free
Start your free assessment