Process / pipelineField MethodsDomain-specific humanities/social sciencePipeline

Program Evaluation — Systematic Assessment of Program Merit and Worth

Also known as: evaluation research, program assessment, educational evaluation, systematic program evaluation

OriginatorMichael Scriven; Daniel Stufflebeam; Peter RossiYear1960s–1970s (Scriven 1967; Stufflebeam CIPP model 1971)Sources2Related methods23

Program evaluation is a systematic, empirically grounded process of collecting and analyzing information about a program to determine its merit, worth, or significance. Applied across education, public health, social services, and policy, it addresses questions such as whether a program is reaching its target population, whether it is being implemented as designed, and whether it is producing the intended outcomes. It draws on both quantitative and qualitative methods and serves accountability, improvement, or knowledge-generation purposes.

Key highlights

  • Produces defensible, evidence-based judgments that go beyond description to inform consequential decisions.
  • Flexible: accommodates quantitative, qualitative, and mixed-method data collection depending on the evaluation question.
  • Engages stakeholders systematically, improving the likelihood that findings will actually be used.
  • Can serve multiple purposes simultaneously — accountability, learning, and knowledge generation — when designed carefully.
  • Established models (CIPP, logic models, realistic evaluation) provide structured frameworks that are widely understood by funders and policymakers.
  • Applicable across a wide range of program types and sectors including education, public health, social policy, and international development.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use program evaluation when a decision-maker needs defensible evidence about whether a program should be continued, expanded, modified, or discontinued. It is the appropriate method when accountability to funders or policy-makers is required, when a program is mature enough for outcome assessment, or when implementation problems are suspected. Program evaluation suits education, health, social services, and government policy contexts. Do not use it as a substitute for formative user research in early program design (needs assessment or participatory design is better then), and do not treat a simple satisfaction survey as a program evaluation — the latter requires explicit evaluative criteria, rigorous data collection, and a documented judgment about merit or worth.

Strengths & limitations

Strengths
  • Produces defensible, evidence-based judgments that go beyond description to inform consequential decisions.
  • Flexible: accommodates quantitative, qualitative, and mixed-method data collection depending on the evaluation question.
  • Engages stakeholders systematically, improving the likelihood that findings will actually be used.
  • Can serve multiple purposes simultaneously — accountability, learning, and knowledge generation — when designed carefully.
  • Established models (CIPP, logic models, realistic evaluation) provide structured frameworks that are widely understood by funders and policymakers.
  • Applicable across a wide range of program types and sectors including education, public health, social policy, and international development.
Limitations
  • Rigorous outcome evaluation requires an adequate comparison strategy; in many real-world settings random assignment is politically or practically infeasible.
  • Results are context-specific: a program that works well in one community may not generalize to another without re-evaluation.
  • Evaluation findings can be ignored or selectively used by stakeholders if the political environment is hostile to negative findings.
  • Comprehensive evaluations are costly and time-consuming; under-resourced evaluations often lack the design rigor to support causal attribution.
  • Tensions between accountability purposes (which may create incentives to report positive findings) and improvement purposes (which require honest identification of failures) can compromise objectivity.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

What is the difference between program evaluation and academic research?

Academic research aims to generate generalizable knowledge and tests theory; it is judged primarily by methodological rigor and contribution to a discipline. Program evaluation aims to inform a specific decision about a specific program; it is judged by utility, feasibility, and the accuracy of its evaluative conclusions. Evaluators must balance rigor with practicality and must explicitly render a judgment — not just report findings — whereas researchers typically do not.

When should I do a process evaluation versus an outcome evaluation?

A process (implementation) evaluation asks whether the program is being delivered as designed and is reaching the intended population — it is appropriate at any stage but especially valuable in the first one to two years of a program. An outcome evaluation asks whether the program is producing the intended changes — it requires that the program be sufficiently mature and consistently delivered. In most cases, process evaluation should precede or accompany outcome evaluation; a clean null outcome result is uninterpretable without knowing whether the program was actually implemented.

Do I need a control group to evaluate my program?

Not necessarily, but attribution of effects to the program requires some form of comparison. Randomized control trials provide the strongest causal evidence but are not always feasible. Quasi-experimental designs — matched comparison groups, regression discontinuity, interrupted time series — can provide reasonable causal estimates. For formative or process evaluations, or when causal claims are not the primary purpose, simpler pre-post or descriptive designs are often sufficient and appropriate.

Who should conduct the evaluation — internal or external evaluators?

External evaluators bring independence and credibility, which is important for accountability-focused evaluations. Internal evaluators have greater program knowledge and are better positioned for ongoing formative evaluation and learning. Many programs use a combination: an internal evaluation team for continuous monitoring and improvement, with periodic external evaluations for summative accountability judgments. The choice should be driven by the evaluation's primary purpose and the need for independence versus deep contextual knowledge.

What is a logic model and is it required?

A logic model (also called a theory of change) is a diagram or narrative that maps the program's intended causal chain from inputs and activities through outputs to short-, medium-, and long-term outcomes. It is not legally required but is considered best practice by most funders and evaluation standards. Without an explicit logic model, evaluators cannot specify what outcomes to measure, at what time points, or how to interpret findings — making the evaluation much harder to design and more likely to produce uninterpretable results.

Sources

  1. 1.
    Rossi, P. H., Lipsey, M. W., & Freeman, H. E. (2004). Evaluation: A Systematic Approach (7th ed.). Sage.
    ISBN 978-0761908944
  2. 2.
    Stufflebeam, D. L., & Shinkfield, A. J. (2007). Evaluation Theory, Models, and Applications. Jossey-Bass.
    ISBN 978-0787977566

You have read it. What now?

Cite this page

ScholarGate. (2026, June 3). Program Evaluation. ScholarGate. https://scholargate.app/field-methods/program-evaluation

Program Evaluation — Program Evaluation Research