Government Performance Measurement
Also known as: Public Sector Performance Measurement, Government Performance Management, Public Performance Metrics, Agency Performance Measurement
Government performance measurement is the systematic, ongoing collection of quantitative and qualitative indicators about what public agencies put in, do, and achieve. Rather than treating measurement as a single number that grades an agency, the discipline — crystallised by Robert Behn's argument that different managerial purposes require different measures — asks first what a measure is for: evaluating, controlling, budgeting, motivating, promoting, celebrating, learning or improving. It draws heavily on Harry Hatry's practical handbook tradition of distinguishing inputs, outputs and outcomes and building measurement into routine operations. The output is not a verdict but a feedback system that ties day-to-day activity to public results.
Key highlights
- Forces explicit articulation of purpose, so agencies collect data that decisions will actually use rather than measuring for its own sake.
- Distinguishes inputs, outputs and outcomes, preventing the common error of mistaking activity counts for public results.
- Provides continuous, low-cost feedback through routine administrative systems rather than expensive one-off studies.
- Supports both internal management improvement and external accountability and transparency to oversight bodies and citizens.
Intuition
This section is available to Pro members. Upgrade to Pro
How it works
This section is available to Pro members. Upgrade to Pro
When to use it
Use government performance measurement when an agency has reasonably stable objectives, a logic linking its activities to public outcomes, and managers or overseers who will act on the results. It is most valuable for routine, repeatable services — benefits processing, sanitation, permitting, public health programs — where indicators can be tracked over time and compared across units. The approach assumes that meaningful outcomes can be operationally defined and measured at acceptable cost, and that consequences attached to measures will not perversely distort behaviour. It is less appropriate for one-off interventions, highly novel programs lacking a settled theory of change, or contexts where outcomes are so attributable to external forces that agency effort cannot be isolated; there, evaluation or experimental designs are better suited than ongoing indicator tracking.
Strengths & limitations
- Forces explicit articulation of purpose, so agencies collect data that decisions will actually use rather than measuring for its own sake.
- Distinguishes inputs, outputs and outcomes, preventing the common error of mistaking activity counts for public results.
- Provides continuous, low-cost feedback through routine administrative systems rather than expensive one-off studies.
- Supports both internal management improvement and external accountability and transparency to oversight bodies and citizens.
- Outcomes are often shaped by factors outside agency control, making attribution of results to agency effort difficult.
- Important public values such as equity, dignity or deterrence are hard to quantify and risk being neglected in favour of countable activities.
- High-stakes measures invite gaming, goal displacement and 'teaching to the metric' that can degrade the very service being measured.
- Sustaining accurate, comparable data over years demands organisational discipline that frequently erodes once initial enthusiasm fades.
Common pitfalls
This section is available to Pro members. Upgrade to Pro
Applications
This section is available to Pro members. Upgrade to Pro
Frequently asked
What is the difference between an output and an outcome?
An output is what an agency produces directly — forms processed, inspections conducted, people trained — and is largely within the agency's control. An outcome is the result that matters to the public — claimants receiving correct benefits on time, restaurants becoming safer, trainees finding jobs — and is influenced by factors beyond the agency. Hatry's framework insists on tracking outcomes because high output volume can coexist with poor public results, and conflating the two is the most common measurement error.
Why does Behn say different purposes require different measures?
Behn identifies eight managerial purposes — evaluate, control, budget, motivate, promote, celebrate, learn and improve — and argues no single measure serves all of them. A measure designed to hold a unit accountable may be a poor motivator for frontline staff, and a budgeting measure may be useless for diagnosing why a process failed. Naming the purpose first prevents agencies from collecting a generic pile of statistics that fits no actual decision.
How do you stop performance measures from being gamed?
Gaming intensifies when a single quantified measure carries high-stakes rewards or sanctions, so robust systems use balanced sets of indicators, mix quantitative metrics with qualitative review, audit data quality, rotate or refresh measures, and treat numbers as a starting point for managerial conversation rather than an automatic verdict. The aim is to keep measurement in service of learning and improvement rather than turning it into a target that distorts the behaviour it was meant to track.
Sources
- 1.Behn, R. D. (2003). Why Measure Performance? Different Purposes Require Different Measures. Public Administration Review, 63(5), 586–606.
- 2.Hatry, H. P. (2006). Performance Measurement: Getting Results (2nd ed.). Washington, DC: Urban Institute Press.ISBN 9780877667346
You have read it. What now?
Cite this page
ScholarGate. (2026, June 22). Government Performance Measurement. ScholarGate. https://scholargate.app/public-administration/performance-measurement-government