Process / pipelinePolitical ScienceContent analysis / computational textPipeline

Event Data Analysis

Also known as: Event data coding, Political event data, Conflict event data, CAMEO event coding

OriginatorConflict-studies and computational-social-science traditions (McClelland, Schrodt, King)Sources3Related methods5

Event data analysis converts streams of news reports into structured records of political interactions — who did what to whom, when — and aggregates them into time series of cooperation and conflict between actors. Each event is coded as a source actor, an action type drawn from an ontology such as CAMEO, a target actor, and a date. Modern systems extract these events automatically from millions of news stories, enabling near-real-time measurement of interstate and intrastate behavior for forecasting and analysis.

Key highlights

  • Produces high-frequency, fine-grained time series of political interactions at a scale impossible for hand coding.
  • Standardized ontologies (CAMEO) and intensity scales make events comparable across actors, places, and time.
  • Automated coding enables near-real-time, continuously updated datasets for early warning and forecasting.
  • Validated against human coders, with automated systems shown to match trained coders on conflict data.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use event data analysis when you need fine-grained, high-frequency measures of interactions between political actors over time — interstate conflict and cooperation, protest and repression, or crisis escalation — and when the relevant behavior is reported in news media. It is well suited to forecasting, early warning, and dynamic analysis of dyadic relations. It is less appropriate when the behavior of interest is poorly covered by media, when media bias systematically distorts which events are reported, or when the research question requires deep contextual interpretation that structured codes cannot capture.

Strengths & limitations

Strengths
  • Produces high-frequency, fine-grained time series of political interactions at a scale impossible for hand coding.
  • Standardized ontologies (CAMEO) and intensity scales make events comparable across actors, places, and time.
  • Automated coding enables near-real-time, continuously updated datasets for early warning and forecasting.
  • Validated against human coders, with automated systems shown to match trained coders on conflict data.
Limitations
  • Subject to media coverage and reporting bias: under-reported regions and events are systematically missing.
  • Automated extraction makes coding errors, misidentifying actors, actions, or targets in ambiguous text.
  • Duplicate reporting can inflate counts if deduplication is imperfect, distorting intensity measures.
  • Structured codes strip context, so the same code can conflate substantively different actions.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

What is CAMEO and how does it relate to Goldstein scores?

CAMEO (Conflict and Mediation Event Observations) is a standardized ontology of political actions — verbal cooperation, material conflict, and so on — used to code event verbs into a common hierarchy. Goldstein-type scores assign each action code a numeric value on a cooperation-to-conflict scale, so that, for example, signing an agreement scores strongly cooperative and an armed attack scores strongly conflictual. Together they let heterogeneous news verbs be turned into comparable, quantitative intensity measures for aggregation.

How reliable is automated event coding compared to human coders?

King and Lowe's 2003 study, using a rare-events evaluation design, found that a well-built automated extraction system could match the performance of trained human coders on international conflict data at far lower cost. Reliability nonetheless varies with text quality, actor disambiguation, and the ontology, so best practice validates automated output against human-coded samples and known benchmarks, and treats coding error as a measurement issue to be modeled rather than ignored.

What is media bias in event data and why does it matter?

Event data are only as complete as the news that reports the events. Regions, actors, and event types that receive little media attention are systematically under-recorded, and source outlets may emphasize certain kinds of events. This reporting bias means event counts reflect both real activity and coverage patterns. Analysts mitigate it by using diverse and local sources, modeling under-reporting, and interpreting trends rather than treating raw counts as a census of behavior.

Sources

  1. 1.
    Schrodt, P. A. (2012). Precedents, Progress, and Prospects in Political Event Data. International Interactions, 38(4), 546–569.
  2. 2.
    King, G., & Lowe, W. (2003). An Automated Information Extraction Tool for International Conflict Data with Performance as Good as Human Coders: A Rare Events Evaluation Design. International Organization, 57(3), 617–642.
  3. 3.
    Boschee, E., Lautenschlager, J., O'Brien, S., Shellman, S., Starz, J., & Ward, M. (2015). ICEWS Coded Event Data. Harvard Dataverse.

You have read it. What now?

Cite this page

ScholarGate. (2026, June 22). Event Data Analysis. ScholarGate. https://scholargate.app/political-science/event-data-analysis