Longhand Press

Method

The four communication behaviours, and the cross-validation problem

Few findings in relationship research travel as far outside the academy as the observational categories developed by John Gottman and Robert Levenson. Fewer still have been contested as directly, or as early, inside it.

Beginning in the 1980s, Gottman and Levenson recorded married couples discussing points of disagreement in a laboratory setting, later in an apartment-style facility at the University of Washington. Sessions were coded second by second using the Specific Affect Coding System, which assigns observed behaviour to categories of emotional expression. Physiological measures — heart rate, skin conductance — were recorded alongside the video.

From this material, Gottman described four behaviours he treated as especially informative: criticism, in which a complaint is framed as a defect of character; contempt, expressed through mockery, sneering, or a tone of superiority; defensiveness, meaning counter-attack or the refusal to accept any share of a problem; and stonewalling, the withdrawal of a listener from the exchange. He labelled the set with a biblical metaphor, and the label is now common outside the research literature.

The descriptive work is not the disputed part. Coding systems of this kind produce reasonable inter-rater agreement, and the four behaviours can be observed and counted. What has been disputed is the second claim: that models built on these observations identified which couples would later separate, with accuracy figures often reported above ninety percent.

The objection

In 2001, Richard Heyman and Amy Slep published an analysis in the Journal of Marriage and Family under the title "The hazards of predicting divorce without crossvalidation." Their argument was statistical rather than substantive. Several of the high-accuracy figures came from models whose parameters had been selected using the same sample the models were then tested against.

A model fitted and evaluated on one dataset will describe that dataset well by construction. The test of prediction is performance on data the model has not seen. When Heyman and Slep applied the procedure to an independent sample, classification accuracy fell substantially below the figures in circulation.

The sample sizes compound the issue. The best-known studies followed dozens of couples rather than thousands, drawn largely from one region of the United States, predominantly white, and married at the time of recruitment. Small samples produce unstable parameter estimates; homogeneous samples limit what can be said about anyone outside them. Neither is a flaw in the conduct of the research — laboratory observation of this depth is expensive and slow — but both constrain the conclusions the data can carry.

A 2010 magazine investigation by the journalist Laurie Abraham brought the cross-validation question to a general readership, tracing how the accuracy figures had been generated and how they had been reported. The peer-reviewed objection had by then been available for nearly a decade.

Where does this leave the four categories? Roughly where the coding work itself places them: as descriptions of behaviour that researchers can identify reliably in recorded conversation, and that correlate with reported dissatisfaction across a number of studies. Correlation across a sample is a weaker statement than prediction for a case, and considerably weaker than the accuracy figures that accompany the metaphor in popular retellings.

It is worth separating the two claims when either is encountered. The observational categories have held up as categories. The forecasting apparatus built on top of them has not held up in the same way, and the researchers who raised that objection did so within the discipline, in print, and were not answered by better cross-validated figures.

Sources

This article summarises published research. It is not professional advice, does not assess any individual situation, and makes no claim about outcomes. See the disclaimer.