You have read that a psychologist can watch a couple argue for a few minutes and predict divorce with 94 per cent accuracy. The number is real. What it measures is not prediction — and the institute’s own defence of it gives the game away.
It is one of the most repeated findings in popular psychology. John Gottman, in a “love lab,” observing couples for fifteen minutes, forecasting the fate of a marriage with better than ninety per cent accuracy.
The studies exist. The percentages appear in them. And the claim still does not mean what almost every article says it means.
The numbers come from different studies than you think
Start by separating three papers that popular coverage routinely fuses into one.
The famous 94 per cent comes from Buehlman, Gottman and Katz (1992). It was not based on watching a couple argue — it used an oral history interview, in which couples were asked about their relationship’s past. The sample was 52 couples with seven divorces, and the model used nine variables. Nine predictors, seven events.
The 1992 conflict-discussion study that people usually have in mind reports no such figure. Its prediction of dissolution yielded a canonical correlation of .52. The 97.7 and 94.3 per cent figures in its tables classify couples as “regulated” or “nonregulated” — a different outcome variable entirely.
And the paper actually titled “Predicting divorce among newlyweds from the first three minutes of a marital conflict discussion” — Carrère and Gottman, 1999 — reports no overall accuracy percentage at all. It compares group means. The percentage in the headline is imported from somewhere else.
None of it was prediction
Here is the structural problem, and it applies to every study above.
In each, the researchers already knew which couples had divorced. They then built a statistical model to separate the divorced from the intact, using variables measured earlier. Then they scored that model on the same couples it was built from.
That procedure cannot fail to produce a high number. It is not forecasting; it is describing. A model with nine variables fitted to seven divorces will separate those seven beautifully, and will tell you almost nothing about the next couple through the door. This is why each paper produces a somewhat different equation, and why each equation “predicts” its own sample so well.
The test that distinguishes a real predictive model from a well-drawn description is straightforward: fit it on one sample, then score it on a fresh one. It is called cross-validation, and it is not optional.
What happens when someone actually runs that test
In 2001, Richard Heyman and Amy Smith Slep published a paper in the Journal of Marriage and Family with the flat title “The hazards of predicting divorce without crossvalidation.”
They did something clever. Rather than reanalyse Gottman’s data, they built a deliberately unglamorous model from national survey data on 528 people — using nothing but education, employment, substance use and number of children. They fitted it on half the sample and tested it on the other half.
On its own data it hit 90 per cent accuracy. On fresh data, accuracy fell to 69 per cent and sensitivity collapsed from 92 to 46 per cent. Most damning, the positive predictive value — how often the model is right when it says a couple will divorce — fell from 65 per cent to 29 per cent. Adjusted to the real population divorce rate of about 16 per cent, it drops to roughly 21 per cent.
A model that looked 90 per cent accurate was wrong about four times out of five when it actually flagged a marriage.
The point is not that demographics predict divorce better than behaviour. The point is the opposite: 90 per cent in-sample accuracy is cheap. You can get it from four boring variables. It is not evidence of anything until it survives fresh data.
Heyman and Smith Slep added a sentence that has aged pointedly: “No published study predicting divorce with general population couples has done this to date.”
The one independent test
In 2007, researchers at the Oregon Social Learning Center took the models from Gottman’s 1998 newlywed study and tested them on 85 couples they had recruited themselves — a lower-income, at-risk sample.
Their conclusion: “the major findings of Gottman et al. failed to replicate.” The headline predictors — men rejecting their partner’s influence, men failing to de-escalate negative affect, women’s harsh start-up — did not predict whether couples separated. Only 2 of 22 affective processes did.
In fairness, this is a genuinely different population from Gottman’s white, middle-class, volunteer samples, and Gottman and Coan published a reply in the same issue arguing exactly that. It is not a foolish objection, and it means the failure to replicate is confounded with real population differences. But it is the only independent test that has been run, and it did not go well.
The defence that proves the point
The Gottman Institute’s research FAQ addresses the criticism. It argues that achieving 90 per cent accuracy by chance “could only happen by chance 1 x 10⁻¹⁹ times.”
This is the crux. That calculation rebuts an accusation nobody has made. No critic claims the in-sample fit arose by luck. The objection is that fitting a model to data and then scoring it on the same data tells you how well the model describes those couples, not how it performs on new ones — and a p-value against chance cannot substitute for a holdout test.
The FAQ does contain a real concession, worth quoting because it is more honest than most of the coverage: the claim is that “a particular couple is behaving like the couples that were in the group that got divorced.” That is a statement about group resemblance. It is a reasonable thing to say. It is also not a prediction, and it is not 94 per cent of anything.
What actually survives
This is not a case for dismissing the work, and it would be lazy to read it that way.
The observational coding systems are a genuine methodological achievement — the Oregon team used the same framework successfully on a completely different population. Their paper is a failure to replicate specific predictive equations, not a repudiation of the method.
The substance largely holds up as correlation. Contempt, criticism, defensiveness and stonewalling really do track marital distress. In the Oregon replication, the affect measures predicted relationship satisfaction among intact couples broadly as expected, including the positive-to-negative ratio findings. Sustaining 4-, 6- and 14-year follow-ups with high retention is hard, and the labs did it.
There is even a finding in this body of work more interesting than the percentages: in a 14-year follow-up, high negative affect predicted early divorce, while the absence of positive affect predicted later divorce. That is genuinely non-obvious, and it gets almost no attention, because it does not come with a number.
What remains unsettled
The decisive study — fit the equations on one sample, score them on an independent one — still appears not to have been published. How much of the 2007 replication failure is overfitting and how much is population difference is genuinely unresolved. Whether communication drives satisfaction or mostly moves with it is contested; recent work suggests the two largely covary rather than one predicting the other. And none of this settles whether Gottman Method couples therapy works, which is a separate question with its own evidence.
What can be said is narrower: the coding works, contempt is corrosive, and the percentage was the marketing. It is the one part of the claim that has never survived contact with data it had not already seen.












