Multilevel Content Validity
Multilevel Content Validity Assessment · Also known as: hierarchical content validity, nested-data content validity, multilevel scale content evaluation, MCV
Multilevel content validity extends the classical content validity framework to settings where items, raters, or respondents are nested within hierarchical structures — such as students within schools, patients within clinics, or items rated by panels from distinct cultural or professional groups. It ensures that scale content is relevant and representative at every level of the hierarchy, not just in the aggregate.
Read the full method
Sign in with a free account to read this section.
Method map
The neighbourhood of related methods — select a node to explore.
When to use it
Use multilevel content validity when developing or adapting a scale for use in inherently hierarchical contexts: students nested in classrooms, employees in organisations, patients in healthcare units, or community members in cultural groups. It is especially important when the construct itself operates at multiple levels (e.g., individual wellbeing influenced by workplace climate) or when the scale will be used for cross-level comparisons. Do not apply this approach when data are purely individual-level with no meaningful nesting, when expert panels cannot be organised to represent distinct hierarchical levels, or when sample sizes at higher levels are so small that ICC estimates would be unstable.
Strengths & limitations
- Ensures scale content is appropriate and interpretable at every level of a nested data structure, not only in the aggregate.
- Reveals level-specific content gaps early in scale development, before costly data collection.
- Combines qualitative expert judgment with quantitative CVI and ICC indices, providing both statistical and substantive evidence.
- Aligns content evaluation with subsequent multilevel analytical plans, enhancing coherence across the validation pipeline.
- Applicable across disciplines — education, organisational psychology, health sciences — wherever hierarchical data structures arise.
- Requires recruiting and coordinating multiple expert panels representative of different levels, which is resource-intensive.
- ICC estimates of rater agreement clustering are unstable when the number of higher-level units is small (fewer than 10–15 groups).
- No universally accepted CVI threshold exists for multilevel settings; practitioners must justify cut-points based on panel size and context.
- The approach evaluates content relevance but does not assess response process validity or construct representation beyond expert judgment.
Frequently asked
How does multilevel content validity differ from standard content validity?
Standard content validity uses a single expert panel to evaluate whether items represent the construct for one target population. Multilevel content validity organises the evaluation around hierarchical levels, ensuring items are relevant and interpretable at each level — individual, group, organisation — and uses ICC to detect whether content adequacy varies across higher-level units.
What sample size of expert panels is needed for each level?
A minimum of five to six experts per level is generally recommended to produce stable CVI estimates. With five raters, the minimum acceptable I-CVI to keep an item with a significance threshold of .05 is 0.80; with six raters it is 0.83. Larger panels reduce the CVI threshold needed for statistical defensibility.
Should multilevel content validity always precede multilevel CFA?
Yes, ideally. Content validity is a logical precondition for structural validity: if items do not adequately represent the construct at the relevant levels, a well-fitting multilevel CFA model may still be measuring the wrong thing. Establishing content validity first reduces the risk of retaining psychometrically convenient but substantively inappropriate items.
What ICC value indicates that content validity varies problematically across higher-level units?
There is no fixed threshold, but an ICC above 0.10 for rater-group membership suggests meaningful clustering in content relevance ratings — meaning different expert groups disagree systematically about which items are relevant. Such findings call for level-targeted item revision rather than a single uniform content validity conclusion.
Can multilevel content validity be applied retrospectively to an existing scale?
Yes, though retrospective evaluation is less efficient. A panel review of existing items using multilevel CVI and ICC can identify which items carry unintended level-specificity or are poorly suited to higher-level interpretation. Findings guide item revision or the development of level-specific subscale scoring.
Sources
- Lynn, M. R. (1986). Determination and quantification of content validity. Nursing Research, 35(6), 382–385. DOI: 10.1097/00006199-198611000-00017 ↗
- Wilson, M. (2005). Constructing Measures: An Item Response Modeling Approach. Lawrence Erlbaum Associates. ISBN: 978-0805847857
How to cite this page
ScholarGate. (2026, June 3). Multilevel Content Validity Assessment. ScholarGate. https://scholargate.app/en/psychometrics/multilevel-content-validity
Which method?
Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.
- Construct ValidityPsychometrics↔ compare
- Content ValidityPsychometrics↔ compare
- Discriminant ValidityPsychometrics↔ compare
- Multilevel CFAPsychometrics↔ compare
- Multilevel Measurement InvariancePsychometrics↔ compare
- Multilevel Scale DevelopmentPsychometrics↔ compare