Hypothesis testStatisticsTest

Descriptive Statistics

Also known as: summary statistics, exploratory data summary, Betimsel İstatistik

OriginatorJohn W. TukeyYear1977Sources1Related methods10

Descriptive statistics is a set of procedures that numerically and visually summarises the essential characteristics of a dataset: central tendency (mean, median, mode), spread (standard deviation, interquartile range), shape (skewness, kurtosis), and frequency distributions. Systematised for applied data analysis by John W. Tukey in his 1977 work on Exploratory Data Analysis, descriptive statistics serves as the indispensable first step before any inferential or modelling procedure.

Key highlights

  • Universal applicability — works with any variable type and any data structure.
  • Requires no distributional assumptions, making it the safest starting point for any analysis.
  • Immediately reveals data quality issues such as outliers, missing values, and implausible ranges.
  • Provides the factual basis — means, standard deviations, sample sizes — needed to justify inferential method choices.
  • Lowest difficulty and mathematical load of any analytical procedure, ensuring broad accessibility.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Descriptive statistics should be applied at the very start of any empirical study, before inferential testing or predictive modelling. It is appropriate for any variable type — continuous, ordinal, categorical, binary, or count data — and for any data structure (cross-sectional, longitudinal, panel). A minimum sample of around ten observations is needed to make spread measures stable, but there is no formal upper limit. The only substantive requirement is that summaries must match the measurement level of each variable: means for ratio and interval scales, medians for ordinal data, and frequencies or proportions for nominal data.

Strengths & limitations

Strengths
  • Universal applicability — works with any variable type and any data structure.
  • Requires no distributional assumptions, making it the safest starting point for any analysis.
  • Immediately reveals data quality issues such as outliers, missing values, and implausible ranges.
  • Provides the factual basis — means, standard deviations, sample sizes — needed to justify inferential method choices.
  • Lowest difficulty and mathematical load of any analytical procedure, ensuring broad accessibility.
Limitations
  • Purely descriptive — no causal or inferential claims can be drawn from summary statistics alone.
  • In skewed or multi-modal distributions a single summary figure (such as the mean) can be misleading if reported without accompanying spread or shape measures.
  • Does not assess statistical significance of observed differences or associations between variables.
  • When applied to very small samples (fewer than ten observations), spread measures such as the standard deviation are highly unstable.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

When should I use the median instead of the mean?

Use the median when the distribution is visibly skewed (absolute skewness > 1) or when outliers are present. The median is a resistant measure of centre — it is not pulled by extreme values the way the mean is. Alongside the median, report the interquartile range rather than the standard deviation as the measure of spread.

Do descriptive statistics require normally distributed data?

No. Descriptive statistics make no distributional assumption. They summarise whatever data you have. Checking for normality using the skewness and kurtosis values, or running a Shapiro-Wilk test, is itself a part of the descriptive phase and informs subsequent inferential choices.

What is the difference between descriptive and inferential statistics?

Descriptive statistics characterise the sample you actually have — they make no claims about any broader population. Inferential statistics use that sample to make probabilistic statements about a population, quantify uncertainty through p-values and confidence intervals, or test hypotheses. Descriptive statistics always comes first; it is the foundation on which valid inference rests.

Which summaries should I always report?

For continuous variables, report the mean and standard deviation if the distribution is approximately symmetric, or the median and interquartile range if it is skewed. Also include the sample size (n) and the range or 95% confidence interval around the mean. For categorical variables, report counts and percentages for each category.

Sources

  1. 1.
    Tukey, J.W. (1977). Exploratory Data Analysis. Addison-Wesley.
    ISBN 978-0201076165

You have read it. What now?

Cite this page

ScholarGate. (2026, June 1). Descriptive Statistics. ScholarGate. https://scholargate.app/statistics/descriptive-statistics

Descriptive Statistics | ScholarGate