Audit Experiment
Also known as: Correspondence study, Field audit study, Discrimination audit, Responsiveness audit
An audit experiment, also called a correspondence or field audit study, sends matched but fictitious requests to real-world targets — such as legislators, landlords, or employers — while randomizing a single treatment cue, then compares the rate and quality of responses. In political science the canonical design follows Butler and Broockman's 2011 study of U.S. state legislators, which varied the putative race signaled by a constituent's name to measure discrimination in responsiveness.
Key highlights
- Measures actual discriminatory behavior in a natural setting, avoiding the social-desirability bias that plagues self-reported attitudes.
- Randomization of the cue delivers strong internal validity, isolating the causal effect of the signaled attribute.
- Captures consequential, real-world responsiveness from genuine officials or gatekeepers rather than hypothetical scenarios.
- Scalable and relatively inexpensive, allowing large samples and many cue variations across institutions or regions.
Intuition
This section is available to Pro members. Upgrade to Pro
How it works
This section is available to Pro members. Upgrade to Pro
When to use it
Use an audit experiment when you want to measure discrimination or differential treatment in a real-world interaction that you can plausibly initiate, and where the target cannot easily detect that the request is experimental. It excels at studying bias in responsiveness by legislators, bureaucrats, landlords, employers, or service providers, capturing revealed behavior rather than stated attitudes. It is inappropriate when the deception would impose serious costs or harm, when the request is implausible enough to be flagged, when the channel cannot deliver matched messages, or when the behavior of interest cannot be triggered by an unsolicited contact.
Strengths & limitations
- Measures actual discriminatory behavior in a natural setting, avoiding the social-desirability bias that plagues self-reported attitudes.
- Randomization of the cue delivers strong internal validity, isolating the causal effect of the signaled attribute.
- Captures consequential, real-world responsiveness from genuine officials or gatekeepers rather than hypothetical scenarios.
- Scalable and relatively inexpensive, allowing large samples and many cue variations across institutions or regions.
- Involves deception of real targets, raising ethical concerns about wasted time, consent, and burden on public officials.
- Typically captures only an initial response, missing downstream behavior and the full quality of treatment over time.
- The signaled cue (e.g., a name) may carry unintended associations beyond the intended construct, confounding interpretation.
- External validity is bounded by the specific request, channel, and population, and effects may not generalize to other interactions.
Common pitfalls
This section is available to Pro members. Upgrade to Pro
Applications
This section is available to Pro members. Upgrade to Pro
Frequently asked
How does an audit experiment differ from a survey experiment?
A survey experiment randomizes a stimulus among respondents and measures their self-reported attitudes or intentions, whereas an audit experiment randomizes a cue in a real request sent to genuine targets and measures their actual behavior. The audit captures consequential, unprompted behavior in a natural setting and sidesteps social-desirability bias, but it relies on deception and observes only the targets' reactions, not their internal reasoning or stated views.
Are audit experiments ethical given that they deceive real people?
Ethics is the central tension. Audit studies typically cannot obtain informed consent because that would defeat the design, so researchers minimize harm by keeping requests realistic and low-burden, often forgoing debriefing when contact would itself impose cost, and securing institutional review board approval. Debate continues about the burden placed on public officials' time and whether the scientific value of measuring discrimination justifies the deception.
How do you make sure the name or cue measures only the intended attribute?
This is the key validity concern. Names chosen to signal race may also signal socioeconomic status, region, or generation, so the estimated effect could conflate several constructs. Careful designs pretest cues to confirm what they actually signal, use multiple names per condition to average over idiosyncrasies, and sometimes add explicit cues or factorial variations to disentangle the dimensions, following the cautions raised in the correspondence-study literature.
Sources
- 1.Butler, D. M., & Broockman, D. E. (2011). Do Politicians Racially Discriminate Against Constituents? A Field Experiment on State Legislators. American Journal of Political Science, 55(3), 463–477.
- 2.Bertrand, M., & Mullainathan, S. (2004). Are Emily and Greg More Employable Than Lakisha and Jamal? A Field Experiment on Labor Market Discrimination. American Economic Review, 94(4), 991–1013.
You have read it. What now?
Cite this page
ScholarGate. (2026, June 22). Audit Experiment. ScholarGate. https://scholargate.app/political-science/audit-experiment