Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Psychometrics›Multilevel Scale Development
Latent structureScale / measurement

Multilevel Scale Development

Also known as: multilevel measurement modeling, hierarchical scale development, MLSEM scale construction, nested data scale development

Multilevel scale development constructs and validates measurement instruments for data collected from individuals nested within higher-level units such as classrooms, organizations, or clinics. It partitions item variance into within-group and between-group components, ensuring that reliability and factor structure are evaluated at both levels simultaneously.

ScholarGate
  1. Latent structure
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Multilevel Scale Development
Confirmatory factor anal…Multilevel CFAMultilevel Measurement I…Multilevel Reliability A…Scale developmentLongitudinal scale devel…Multilevel Content Valid…

When to use it

Use multilevel scale development when your sample is hierarchically nested (pupils within schools, employees within firms, patients within clinics) and when the ICC for the items is non-trivial, meaning that group membership explains a meaningful portion of item variance. It is especially warranted when the scale will be used to make inferences at both the individual and the group level. Do not use it when data come from simple random samples with no natural grouping, when cluster sizes are very small (fewer than five members per group), or when there are too few groups (fewer than roughly 30) to estimate between-level parameters reliably. Also avoid the multilevel approach if your theoretical model posits that the construct operates only at one level.

Strengths & limitations

Strengths
  • Correctly partitions item variance into within-person and between-group components, avoiding biased reliability and factor loading estimates from ignoring clustering.
  • Allows the researcher to test whether the same factor structure holds at both individual and group levels, revealing whether a scale measures the same construct across levels.
  • Supports inferences at the group level (e.g., school-average climate) without ecological fallacy, because the between-level model is explicitly specified.
  • Provides level-specific reliability estimates so that scale quality can be evaluated independently for individual- and group-level research questions.
  • Compatible with modern multilevel structural equation modeling (MLSEM) software, enabling simultaneous estimation of measurement and structural parameters.
Limitations
  • Requires sufficient numbers of groups (typically at least 30) and sufficient group sizes to estimate between-level parameters with acceptable precision.
  • Model complexity increases substantially compared to single-level CFA or EFA, demanding larger overall sample sizes and greater statistical expertise.
  • Between-level factor structures can be difficult to interpret when group sizes are unequal or when ICCs are very small, yielding unstable parameter estimates.
  • Software implementation (e.g., Mplus, R packages such as lavaan with cluster correction or lme4) is more demanding than standard factor analysis tools.
  • Existing scale development guidelines were largely developed for single-level data; direct application of conventional criteria (e.g., loading thresholds, fit index cutoffs) to multilevel models requires caution.

Frequently asked

How do I decide whether multilevel scale development is necessary?

Calculate the intraclass correlation (ICC) for each item. If ICC values exceed roughly 0.05–0.10 and your research questions involve both individual and group levels, a multilevel model is warranted. If ICCs are negligible and all inferences concern only individual-level variation, standard single-level methods with cluster-robust standard errors may be sufficient.

How many groups and how large must they be?

A common minimum is around 30 groups for stable between-level estimates, though more groups improve precision. Cluster sizes of at least five to ten members per group are recommended; very small clusters yield unreliable between-level estimates. Simulation studies and power analysis tools (e.g., the R package powerlmm) should be used to plan adequate sampling designs.

Can the factor structure differ across levels?

Yes, and this is one of the most important insights of multilevel scale development. A two-factor within-level solution may collapse to a single between-level factor, or items may load on entirely different dimensions at the group level. Always test and report the factor structure at each level separately rather than assuming isomorphism.

What software can I use?

Mplus is the most widely used platform for multilevel CFA and MLSEM. In R, the lavaan package supports multilevel CFA via the sem() function with cluster arguments, and the lme4 or nlme packages handle random-effects models. Commercially, LISREL and SAS PROC CALIS also offer multilevel options.

How do I report between-level reliability?

Common between-level reliability indices include the between-level omega (omega_B) from MLSEM, the rwg or rwg(j) inter-rater agreement index for group-mean scores, and the ICC(2) (average rater reliability). Report the index appropriate to your aggregation rationale and cite its formula explicitly, as conventions vary across disciplines.

Sources

  1. Hox, J. J. (2010). Multilevel Analysis: Techniques and Applications (2nd ed.). Routledge. ISBN: 978-1848728462
  2. Raudenbush, S. W. & Bryk, A. S. (2002). Hierarchical Linear Models: Applications and Data Analysis Methods (2nd ed.). Sage Publications. ISBN: 978-0761919049

How to cite this page

ScholarGate. (2026, June 3). Multilevel Scale Development. ScholarGate. https://scholargate.app/en/psychometrics/multilevel-scale-development

Related methods

Confirmatory factor analysisMultilevel CFAMultilevel Measurement InvarianceMultilevel Reliability AnalysisScale development

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • Confirmatory factor analysisPsychometrics↔ compare
  • Multilevel CFAPsychometrics↔ compare
  • Multilevel Measurement InvariancePsychometrics↔ compare
  • Multilevel Reliability AnalysisPsychometrics↔ compare
  • Scale developmentPsychometrics↔ compare
Compare side by side →

Referenced by

Longitudinal scale developmentMultilevel Content Validity

Similar methods

Multilevel Measurement InvarianceMultilevel Reliability AnalysisMultilevel CFAMultilevel EFAMultilevel Convergent ValidityMultilevel Discriminant ValidityMultilevel Test-Retest ReliabilityMultilevel McDonald's omega

Related reference concepts

Structural and Latent Variable ModelsItem Response TheoryStructural Equation ModelingPsychometrics & Statistics & MethodologyHierarchical Linear ModelingDevelopmental Scales & Schedules

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Multilevel Scale Development (Multilevel Scale Development). Retrieved 2026-07-21 from https://scholargate.app/en/psychometrics/multilevel-scale-development · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Raudenbush, Bryk, Hox and colleagues
Year
1990s–2000s
Type
Hierarchical measurement / scale construction
DataType
Ordinal or continuous items from nested (clustered) samples
Subfamily
Scale / measurement
Related methods
Confirmatory factor analysisMultilevel CFAMultilevel Measurement InvarianceMultilevel Reliability AnalysisScale development
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account