Process / pipelineOrganizational BehaviorOrganizational behavior / performance appraisalPipeline

Behaviorally Anchored Rating Scales

Also known as: BARS, Behavioral Expectation Scales, Smith-Kendall Scales, Behaviorally Anchored Scales

OriginatorPatricia Cain Smith & L. M. KendallYear1963Sources1Related methods8

Behaviorally anchored rating scales (BARS) are performance-appraisal instruments whose scale points are defined by concrete examples of job behavior rather than by vague adjectives like 'good' or 'excellent.' Patricia Cain Smith and L. M. Kendall introduced the method in 1963 with their technique of retranslation of expectations, a procedure for constructing unambiguous behavioral anchors. The core problem they tackled is that ordinary rating scales leave raters to guess what each numerical point means, so that one supervisor's 4 is another's 2, fatally undermining reliability and fairness. BARS solves this by attaching specific behavioral descriptions, drawn from critical incidents and vetted by independent expert judges, to each level of each performance dimension. The construction process is deliberately participatory and quantitative, which both improves measurement and builds rater understanding. BARS became one of the most influential and widely studied formats in performance appraisal.

Key highlights

  • Replaces ambiguous numeric anchors with concrete behavioral examples, giving scale points a shared meaning across raters.
  • Grounds the scale in real job behavior through critical incidents and the retranslation consensus filter, improving content validity.
  • Engages the eventual users in construction, increasing rater understanding and acceptance of the appraisal system.
  • Provides behaviorally specific, defensible documentation useful for feedback, development, and legal scrutiny.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use behaviorally anchored rating scales when you need defensible, behavior-based performance appraisal and rater consistency matters, for example in legally sensitive evaluation, promotion decisions, or any setting where vague adjective scales have produced inconsistent or contested ratings. BARS are well suited to jobs with identifiable, observable behaviors that recur enough to be captured as critical incidents, and where the considerable development effort can be justified by repeated use. They are less appropriate for highly idiosyncratic or rapidly changing jobs whose key behaviors cannot be stably described, for one-off evaluations where the construction cost cannot be amortized, or where the relevant performance is better captured by objective output metrics than by behavioral observation.

Strengths & limitations

Strengths
  • Replaces ambiguous numeric anchors with concrete behavioral examples, giving scale points a shared meaning across raters.
  • Grounds the scale in real job behavior through critical incidents and the retranslation consensus filter, improving content validity.
  • Engages the eventual users in construction, increasing rater understanding and acceptance of the appraisal system.
  • Provides behaviorally specific, defensible documentation useful for feedback, development, and legal scrutiny.
Limitations
  • Construction is labor-intensive and time-consuming, requiring incident collection, retranslation, and scaling by multiple judges.
  • Scales are job-specific and must be rebuilt when jobs change or for each new role, limiting transferability.
  • Real behavior may fall between anchors or span several, forcing raters to interpolate and reintroducing some judgment error.
  • Empirical gains in reliability and reduced rating errors over simpler formats have proven modest and inconsistent across studies.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

What is the retranslation step and why does it matter?

Retranslation is the heart of the BARS method. After developers tentatively assign critical incidents to performance dimensions, a separate, independent group of expert judges is asked to sort the same incidents into dimensions on their own. Only incidents that a high proportion of these judges independently place in the same dimension are kept. Smith and Kendall called this 'retranslation of expectations' because the incident must survive being translated back into a dimension by fresh judges. It matters because it filters out ambiguous incidents: if independent experts cannot agree what an incident exemplifies, it would confuse raters, so removing it is what gives the anchors their shared, unambiguous meaning.

Do BARS actually produce more accurate ratings than ordinary scales?

The theoretical case is strong, but the empirical evidence is mixed. Decades of research comparing BARS with graphic rating scales and other formats found that the improvements in reliability and reductions in errors such as leniency and halo were generally modest and not always present. The construction process is also costly. That said, BARS offer real advantages beyond raw psychometrics: the behavioral anchors improve the clarity and defensibility of ratings, support better feedback, and engage users in development. So BARS are valued less for dramatically higher accuracy than for the shared behavioral language and procedural legitimacy they bring to appraisal.

How do BARS differ from behavioral observation scales?

Both use behavioral content, but they ask raters different questions. Classic BARS, originally behavioral expectation scales, ask the rater to judge what an employee could be expected to do and to locate them against behavioral anchors arranged from poor to excellent. Behavioral observation scales (BOS), a later variant, instead list specific behaviors and ask how frequently the employee actually performed each one, then sum or average the frequencies. BARS thus emphasize expected level of effectiveness anchored by examples, while BOS emphasize observed frequency of concrete behaviors. Both descend from the critical incident technique and the goal of grounding appraisal in behavior rather than vague trait judgments.

Sources

  1. 1.
    Smith, P. C., & Kendall, L. M. (1963). Retranslation of expectations: An approach to the construction of unambiguous anchors for rating scales. Journal of Applied Psychology, 47(2), 149-155.

You have read it. What now?

Cite this page

ScholarGate. (2026, June 23). Behaviorally Anchored Rating Scales. ScholarGate. https://scholargate.app/organizational-behavior/behaviorally-anchored-rating-scales

Behaviorally Anchored Rating Scales | ScholarGate