Process / pipelineClinical PsychologyMood-disorder-assessment-clinician-ratedPipeline

Hamilton Depression Rating Scale (HAM-D)

Also known as: HAM-D, HDRS, Hamilton Rating Scale for Depression

OriginatorMax HamiltonYear1960Sources3Related methods9

The Hamilton Depression Rating Scale, published by Max Hamilton in 1960, is a clinician-administered interview assessment of depressive symptom severity. The most common version contains 17 items (HAM-D-17), though 21-item and 24-item versions exist. It is considered the gold standard outcome measure in antidepressant drug trials and remains the most cited depression rating scale in the psychiatric literature. Unlike self-report measures, HAM-D requires clinician judgment and observation, making it particularly valuable in research settings where standardized measurement by trained raters is essential.

Key highlights

  • Gold standard in antidepressant trials—deeply embedded in drug efficacy research, FDA submissions, and regulatory decision-making, with 60+ years of precedent
  • Clinician-based judgment captures behavioral cues self-report cannot—psychomotor changes, lack of insight, subtle suicidal communication, and affect incongruence detected through interview and observation
  • Extensively validated across decades—robust literature documenting reliability, validity, and responsiveness to medication and psychotherapy in diverse populations
  • Item diversity—covers cognitive, affective, motivational, vegetative, and suicidal domains across 17 dimensions, providing multidimensional symptom assessment
  • Sensitive to treatment response—detects medication efficacy and psychotherapy gains with good effect sizes in controlled trials
  • Publicly available—no licensing fee or copyrighted restrictions; widely accessible for researchers and clinicians

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Primary use: baseline and endpoint assessment in antidepressant efficacy trials and FDA new drug applications, where regulatory standards mandate clinician-rated outcomes. Also appropriate for research studies requiring objective, standardized measurement resistant to self-report bias. Suitable for hospitalized depressed patients where direct behavioral observation complements interview data. Less appropriate for routine primary care screening (too time-intensive) or resource-limited settings; PHQ-9 or BDI-II are more practical alternatives. Requires trained, reliable raters; single-rater use defeats the purpose of standardized clinician assessment.

Strengths & limitations

Strengths
  • Gold standard in antidepressant trials—deeply embedded in drug efficacy research, FDA submissions, and regulatory decision-making, with 60+ years of precedent
  • Clinician-based judgment captures behavioral cues self-report cannot—psychomotor changes, lack of insight, subtle suicidal communication, and affect incongruence detected through interview and observation
  • Extensively validated across decades—robust literature documenting reliability, validity, and responsiveness to medication and psychotherapy in diverse populations
  • Item diversity—covers cognitive, affective, motivational, vegetative, and suicidal domains across 17 dimensions, providing multidimensional symptom assessment
  • Sensitive to treatment response—detects medication efficacy and psychotherapy gains with good effect sizes in controlled trials
  • Publicly available—no licensing fee or copyrighted restrictions; widely accessible for researchers and clinicians
Limitations
  • Requires trained clinicians—administration is time-consuming (15–30 min) and demands rater training; inter-rater reliability varies (ICC 0.50–0.75 even with training), limiting reliability in busy clinical settings
  • Not appropriate for self-administration—differs fundamentally from self-report instruments; patient alone cannot complete HAM-D
  • Variable item anchors—inconsistent scaling across items (0–2, 0–3, 0–4, half-points) complicates scoring and interpretation; some items are ambiguously defined
  • Rater drift over time—clinicians' scoring standards may shift with experience or fatigue, particularly in long-term studies; regular rater calibration is essential but often omitted
  • Emphasis on vegetative symptoms—insomnia, appetite, sexual interest represent 5+ items, potentially overweighting somatic depression in patients with medical comorbidities
  • Factor structure debated—psychometric analyses suggest multiple overlapping dimensions rather than a unitary depression construct, raising questions about item coherence

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

Who can administer the HAM-D?

The HAM-D is designed for trained clinicians: psychiatrists, psychologists, psychiatric nurses, or research coordinators with formal training. Untrained administration yields unreliable scores. Training typically involves studying the interview guide, practicing with standardized vignettes, and achieving reliability on supervised patient assessments. For research trials, formal rater certification is expected.

How does HAM-D differ from self-report scales like PHQ-9?

HAM-D is clinician-administered and incorporates behavioral observation; PHQ-9 is self-report. HAM-D requires clinical judgment and training; PHQ-9 requires only literacy. HAM-D is the gold standard in research trials; PHQ-9 is more practical for routine screening. HAM-D is longer (15–30 min) and costlier in clinician time; PHQ-9 takes 2–5 min.

What is the difference between response and remission on HAM-D?

Response is typically a ≥50% reduction in HAM-D score from baseline (e.g., 24 down to 12). Remission is a much lower absolute score, usually <7 or <8, indicating near-normal mood and function. A patient may respond (improve substantially) without remitting (returning to non-depressed state). Remission is the clinical goal; response alone may leave the patient symptomatic.

Why do HAM-D scores sometimes vary between raters even with training?

HAM-D items contain subjective anchors and require clinical judgment. Different raters may weight patient report versus observed behavior differently, interpret ambiguous responses variably, or apply different thresholds for 'moderate' vs 'severe' symptoms. Regular rater calibration (comparing ratings on video vignettes) and supervision minimize drift but cannot eliminate inter-rater variation entirely.

Is HAM-D appropriate for use in primary care?

HAM-D is not ideal for routine primary care due to length (15–30 min), rater training demands, and cost. PHQ-9 or other brief self-report measures are more practical for screening. HAM-D is reserved for specialized settings: psychiatric research, antidepressant trials, and complex inpatient cases where clinician-based assessment adds diagnostic or therapeutic value.

Sources

  1. 1.
    Hamilton, M. (1960). A rating scale for depression. Journal of Neurology, Neurosurgery & Psychiatry, 23(1), 56–62.
  2. 2.
    Bagby, R. M., Ryder, A. G., Schuller, D. R., & Marshall, M. B. (1997). The Hamilton Depression Rating Scale: has the gold standard become a lead weight? American Journal of Psychiatry, 161(12), 2163–2177.
  3. 3.
    Williams, J. B. (1988). A structured interview guide for the Hamilton Depression Rating Scale. Archives of General Psychiatry, 45(8), 742–747.

You have read it. What now?

Cite this page

ScholarGate. (2026, June 3). Hamilton Depression Rating Scale. ScholarGate. https://scholargate.app/clinical-psychology/hamilton-depression-rating-scale