Collostructional Analysis
Also known as: Collexeme Analysis, Distinctive Collexeme Analysis, Co-varying Collexeme Analysis
Collostructional analysis is a family of corpus-based methods, introduced by Anatol Stefanowitsch and Stefan Th. Gries in 2003, that quantify the mutual attraction or repulsion between specific words (lexemes) and the grammatical constructions they occur in. Rooted in construction grammar, it treats a construction — such as the ditransitive "V NP NP" or the "into-causative" — as a meaningful unit and asks which words are statistically drawn to it or kept from it. The core technique, simple collexeme analysis, cross-tabulates how often a lexeme appears in the construction against how often each appears elsewhere, and measures the strength of association, conventionally with a Fisher–Yates exact test. Two extensions handle near-synonymous constructions (distinctive collexeme analysis) and the joint behavior of two slots within one construction (co-varying collexeme analysis), making the method a rigorous quantitative window onto the lexis–grammar interface.
Key highlights
- Quantifies the attraction between words and grammatical constructions, going beyond raw frequency to expected co-occurrence.
- Grounded in construction grammar, giving its measures a clear theoretical interpretation about meaning and selection.
- The Fisher–Yates exact test handles the skewed, low-frequency counts typical of construction data without distributional assumptions.
- Its distinctive and co-varying variants extend the same logic to alternations and to interactions between two construction slots.
Intuition
This section is available to Pro members. Upgrade to Pro
How it works
This section is available to Pro members. Upgrade to Pro
When to use it
Use collostructional analysis when you want quantitative, corpus-based evidence about how words and grammatical constructions select one another — to characterize the meaning of a construction from its preferred verbs, to adjudicate alternations such as the dative or particle-placement alternation, or to test construction-grammar and usage-based hypotheses about the lexis–grammar interface. It is well suited to second-language acquisition, lexicography, and semantic research that needs to move beyond raw frequency to attraction. It is less appropriate when the construction cannot be reliably identified in the corpus, when counts are too sparse for stable estimates, or when the research question concerns simple co-occurrence rather than the word–construction relationship.
Strengths & limitations
- Quantifies the attraction between words and grammatical constructions, going beyond raw frequency to expected co-occurrence.
- Grounded in construction grammar, giving its measures a clear theoretical interpretation about meaning and selection.
- The Fisher–Yates exact test handles the skewed, low-frequency counts typical of construction data without distributional assumptions.
- Its distinctive and co-varying variants extend the same logic to alternations and to interactions between two construction slots.
- Results depend heavily on how the construction is defined and on the accuracy of identifying its instances in the corpus.
- Computing cell d (everything else) requires a defensible estimate of the construction's complement, which is often contested.
- Association strength reflects the corpus and may conflate frequency artifacts with genuine semantic attraction if not checked.
- The single-number ranking can obscure the qualitative reasons a word is attracted, so it must be paired with concordance inspection.
Common pitfalls
This section is available to Pro members. Upgrade to Pro
Applications
This section is available to Pro members. Upgrade to Pro
Frequently asked
How is collostructional analysis different from collocation analysis?
Collocation analysis measures the statistical association between a word and other words in its textual window. Collostructional analysis measures the association between a word and a grammatical construction — an abstract pattern such as the ditransitive — rather than a neighboring word. The contingency-table logic is similar, but the second variable is a construction defined by the analyst, which is why the method is tied to construction grammar and requires identifying construction instances in the corpus.
What are the three kinds of collostructional analysis?
Simple collexeme analysis measures how strongly individual words are attracted to or repelled by a single construction. Distinctive collexeme analysis compares two functionally similar or alternating constructions to find which words prefer each. Co-varying collexeme analysis examines the association between the words filling two different slots within one construction. All three share the same association-statistic machinery applied to different contingency tables.
Why use the Fisher exact test for collostruction strength?
Construction data are typically skewed and sparse — a few very frequent verbs and many rare ones — conditions under which chi-square and log-likelihood approximations can be unreliable. The Fisher–Yates exact test computes the probability directly from the contingency table without distributional assumptions, and collostruction strength is reported as the negative log of that p-value so that larger numbers indicate stronger, more significant association.
Sources
- 1.Stefanowitsch, A., & Gries, S. T. (2003). Collostructions: Investigating the interaction of words and constructions. International Journal of Corpus Linguistics, 8(2), 209–243.
- 2.Gries, S. T., & Stefanowitsch, A. (2004). Extending collostructional analysis: A corpus-based perspective on alternations. International Journal of Corpus Linguistics, 9(1), 97–129.
You have read it. What now?
Cite this page
ScholarGate. (2026, June 22). Collostructional Analysis. ScholarGate. https://scholargate.app/linguistics/collostructional-analysis