Home Work & Purpose Your Myers-Briggs type is four coin flips in a trench coat

Your Myers-Briggs type is four coin flips in a trench coat

0
184
A blank identification card attached to a lanyard
Photo by jim on Unsplash (unsplash.com/@skill01). Used under the Unsplash License.

The Myers-Briggs measures four real traits with decent reliability — and then throws away most of what it measured. The problem is not that it is fake. It is the arithmetic of the cut-point.

Around two million people take the Myers-Briggs Type Indicator each year. It is used in team-building, careers advice, management training and, despite the publisher’s explicit instruction not to, in hiring.

The usual debunk — that it is astrology for offices, that it measures nothing — is wrong, and being wrong about it is why the debunk keeps getting rebutted. The real criticism is narrower, more technical, and much harder to answer.

It does measure something

In a 1989 study of 468 adults, Robert McCrae and Paul Costa correlated MBTI scores against the Big Five personality dimensions. Extraversion–Introversion correlated with Big Five Extraversion at around −0.70. Sensing–Intuition correlated with Openness at around 0.72. Those are strong — roughly half the variance shared.

Internal consistency is genuinely good too: a 2025 synthesis of 193 studies found figures between 0.845 and 0.921.

So the instrument is picking up real, well-established trait variance. Two caveats worth holding: Thinking–Feeling and Judging–Perceiving correlate far more weakly (around 0.45), and there is no MBTI scale at all corresponding to Neuroticism — arguably the single most consequential personality dimension for wellbeing and job outcomes.

McCrae and Costa’s verdict was that the four-letter codes “simply summarize four independent main effects.” Four continuous traits. Not sixteen kinds of person.

The cut-point problem

This is the heart of it.

For a type system to work, the underlying trait has to be genuinely categorical — two clumps with a gap between them. If scores are normally distributed, the line you draw falls in the crowded middle, and the people nearest it are sorted almost at random.

Researchers tested this using taxometric methods designed specifically to detect latent categories. Their conclusion: “there is not a true, non-arbitrary taxon underlying Jungian preferences measured by any of these measures… the preferences appear to manifest as continuous dimensions.”

This is not an MBTI-specific scandal. A review of 177 taxometric studies covering 533,377 people found only 14 per cent of constructs were genuinely categorical, and that normal personality yielded “little persuasive evidence of taxa.” Almost nothing in personality comes in kinds.

So the line between I and E runs through the middle of a bell curve. Someone scoring marginally on one side and someone scoring marginally on the other receive opposite four-letter identities, opposite profile booklets and opposite career suggestions — and the difference between them is smaller than the instrument’s own margin of error.

Why the scales are steady but the type is not Each letter is reasonably reliable on its own. All four have to be right at once. E / I.76 S / N.75 T / F.61 J / P.78 test–retestreliability All four letters must hold simultaneously — ~50% change a letter on retest within weeks. Scale reliabilities: Randall, Isaacson & Ciro, 2017 (pooled). Type-change figure: Pittenger 1993, based on Form G, which the publisher notes was retired in the 1990s.

Which explains the retest problem

The famous statistic is that around half of people get a different type when they retake it weeks later. The publisher pushes back — fairly — that this traces to a 1979 study of a form retired in the 1990s, and reports correlations of 0.81 to 0.86 for the current version.

But the mechanism is not in dispute, and it is arithmetic. Pooled test-retest reliabilities for the four scales run from 0.61 to 0.78. Requiring all four letters to land the same way compounds that error four times over. Even at a generous 0.90 per dichotomy, you would expect only about 66 per cent of people to get the identical four-letter code twice.

The scales are stable. The type is not. And it is the type that gets printed on the lanyard.

What it does not predict

A 2023 study of 464 business students modelled the four MBTI dichotomies against five dimensions of leadership practice. The personality-to-leadership path yielded R² = 0.01 — about one per cent of the variance.

Fairness requires the comparison. The best personality predictor of job performance that exists is Conscientiousness, at a corrected validity of around 0.20 — roughly four per cent of variance. The ceiling for any personality measure here is low. The MBTI is not being held to a standard its rivals sail past.

In 1991 the US National Research Council concluded that “at this time there is not sufficient, well-designed research to justify the use of the Myers-Briggs Type Indicator (MBTI) in career counseling programs,” and called for longitudinal studies comparing successful and unsuccessful careers. Those studies do not appear to have been done.

The publisher agrees with the critics on hiring

This is the part that surprises people. The Myers-Briggs Company’s own published position is that “it’s unethical to use the MBTI assessment in recruitment or selection,” because doing so “confuses preferences with skills and abilities.” Practitioners found misusing it can be cut off from the product.

If your employer screened you with it, they were going against the vendor’s written instruction.

The fair case for it

Its defenders make an argument worth taking seriously, in peer-reviewed form: selection and development are different enterprises with different standards. If an instrument exists to prompt self-reflection and give a team shared vocabulary, judging it by hiring-style predictive validity is judging it against a job it does not claim.

And the non-evaluative framing is a real design achievement. Every pole is described positively; no type is worse. In an office, a vocabulary nobody is ashamed to be assigned has obvious practical advantages over one that tells a colleague they scored low on emotional stability. Even hostile reviewers concede its value as a teaching tool.

Whether it actually delivers durable self-knowledge is untested. The National Research Council flagged in 1991 that research needed to separate people’s satisfaction with feedback from real effects. That gap is still open.

What is genuinely unsettled

The convergent-validity correlations are not fully consistent across studies — a 2022 analysis of 9,487 adults reported substantially weaker MBTI–Big Five correlations than the classic 1989 figures, and we cannot tell from the published tables whether that reflects different scoring or a genuine failure to replicate.

And there is a striking gap: the 2025 synthesis found that across 193 studies published since 1999, no structural-validity studies and no test-retest studies met inclusion criteria. The best-known reliability evidence is about a form no longer sold, and in a quarter of a century nobody has published a public check on the current one.

What to do with your four letters

Treat them as four dials, not a label. If a description resonates, that is worth something — the traits underneath are real. If you sat near a boundary, you would have received a different description that also resonated, which is worth knowing too.

The question is not whether Katharine Briggs and Isabel Myers had psychology credentials. They did not, and the official handbook says so. The question is why an instrument sold on the premise of sixteen kinds of person has gone twenty-five years without anyone publishing a check on whether the kinds hold still.