Process / pipelineSociologyNetwork composition analysisPipeline

Homophily Analysis

Also known as: homophily measurement, assortative mixing analysis, birds-of-a-feather analysis, tie-similarity analysis

OriginatorLazarsfeld & Merton (concept); McPherson, Smith-Lovin & Cook (synthesis)Year1954 (concept); 2001 (synthesis)Sources2Related methods15

Homophily analysis quantifies the tendency of similar individuals to form ties — the principle that 'birds of a feather flock together'. It compares the rate at which people connect with others who share an attribute (race, gender, age, education, attitudes) against what would be expected by chance, distinguishing the homophily that arises merely from group sizes from the genuine, behavior-driven preference for similar others.

Key highlights

  • Quantifies a fundamental and pervasive social mechanism with simple, interpretable indices.
  • Distinguishes baseline homophily (from group sizes) from genuine inbreeding preference.
  • Scales from descriptive mixing matrices to inferential ERGM/SAOM tests of similarity effects.
  • Applicable to categorical and continuous attributes and to whole networks, groups, or individual nodes.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use homophily analysis when you have network data with measured nodal attributes and want to know whether and how strongly similarity drives tie formation — in friendship, marriage, collaboration, or communication networks. Descriptive measures (mixing matrix, assortativity, E–I index) summarize observed patterns; model-based terms test homophily net of structural confounds; longitudinal models separate selection from influence. It requires accurate attribute data and a defined network boundary. It is not appropriate when attributes are unmeasured, when group sizes are ignored (baseline homophily will be mistaken for preference), or when cross-sectional data are used to claim selection over influence, which they cannot identify.

Strengths & limitations

Strengths
  • Quantifies a fundamental and pervasive social mechanism with simple, interpretable indices.
  • Distinguishes baseline homophily (from group sizes) from genuine inbreeding preference.
  • Scales from descriptive mixing matrices to inferential ERGM/SAOM tests of similarity effects.
  • Applicable to categorical and continuous attributes and to whole networks, groups, or individual nodes.
Limitations
  • Descriptive measures do not distinguish whether similarity caused ties (selection) or ties caused similarity (influence).
  • Failing to account for group sizes conflates baseline with inbreeding homophily, overstating preference for majorities.
  • Results depend on how attribute categories are defined and on the network boundary.
  • Missing ties or attribute data bias mixing matrices and assortativity estimates.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

What is the difference between baseline and inbreeding homophily?

Baseline homophily is the level of same-group tie formation expected purely from the relative sizes of groups, even under random mixing — a large majority will have many same-group ties by chance alone. Inbreeding homophily is the additional same-group bias beyond that baseline, reflecting a genuine preference for similar others. Only inbreeding homophily is evidence of homophilous choice, so analyses must net out the baseline.

Can homophily analysis distinguish selection from social influence?

Not from cross-sectional data. An observed clustering of similar people connected to each other is equally consistent with selection (similar people choosing to connect) and influence (connected people becoming similar). Separating the two requires longitudinal network-and-behavior data and models such as stochastic actor-oriented models that estimate selection and influence simultaneously.

What does the assortativity coefficient measure?

It is essentially the correlation between the attributes of nodes at the two ends of a tie, normalized to lie between −1 and 1. A value near 1 means strongly assortative mixing (ties overwhelmingly within categories), near 0 means mixing indistinguishable from random, and negative values mean disassortative mixing where opposites tend to connect. It condenses the whole mixing matrix into one comparable number.

Sources

  1. 1.
    McPherson, M., Smith-Lovin, L., & Cook, J. M. (2001). Birds of a feather: homophily in social networks. Annual Review of Sociology, 27, 415–444.
  2. 2.
    Newman, M. E. J. (2003). Mixing patterns in networks. Physical Review E, 67(2), 026126.

You have read it. What now?

Cite this page

ScholarGate. (2026, June 22). Homophily Analysis. ScholarGate. https://scholargate.app/sociology/homophily-analysis