Process / pipelineInternational RelationsInformation extraction / NLP for political sciencePipeline

Event Data Analysis of Conflict

Also known as: Political Event Data, Machine-Coded Conflict Event Data, Conflict Event Extraction, Who-Did-What-to-Whom Event Coding

OriginatorPhilip Schrodt (KEDS/TABARI); ICEWS team (Boschee et al.)Year1994Sources2Related methods10

Event data analysis is the automated extraction of structured records of political interactions — who did what to whom, when, and where — from large volumes of news text, for the quantitative study of conflict and cooperation. Pioneered for machine coding by Philip Schrodt with the KEDS and TABARI systems and scaled in projects such as ICEWS and GDELT, it turns unstructured reporting into dated actor-action-target triples coded to an ontology like CAMEO, which can then be aggregated into time series of interstate or intrastate hostility.

Key highlights

  • Scales to millions of reports, producing near-real-time, high-frequency measures impossible to hand-code.
  • Yields standardized, reproducible codes (CAMEO) and intensity scores (Goldstein) comparable across studies.
  • Covers the full conflict-cooperation spectrum and any pair of actors the news mention, not just wars.
  • Supports forecasting and early-warning models because new events are coded as soon as they are reported.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use event data analysis when you need fine-grained, high-frequency measures of political interaction across many actors and a long time span — for conflict forecasting, early warning, studying escalation dynamics, or testing theories of cooperation and conflict at daily or weekly resolution. It excels where hand-coding is infeasible because of volume. It is less suitable when the phenomenon is rarely reported or systematically underreported, when fine semantic distinctions exceed the coder's accuracy, or when the research question requires the contextual judgment that structured triples discard; in those cases curated datasets like the Correlates of War militarized disputes may be preferable.

Strengths & limitations

Strengths
  • Scales to millions of reports, producing near-real-time, high-frequency measures impossible to hand-code.
  • Yields standardized, reproducible codes (CAMEO) and intensity scores (Goldstein) comparable across studies.
  • Covers the full conflict-cooperation spectrum and any pair of actors the news mention, not just wars.
  • Supports forecasting and early-warning models because new events are coded as soon as they are reported.
Limitations
  • Entirely dependent on media coverage: events that go unreported do not exist in the data, and coverage is geographically and politically uneven.
  • Automated parsing makes systematic errors — misattributed actors, miscoded actions, duplicate reports inflating counts.
  • The CAMEO ontology and Goldstein weights impose a fixed schema that may not match a study's theoretical categories.
  • High event volume can create an illusion of precision that masks substantial measurement error and source bias.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

What is CAMEO and why is it central to conflict event data?

CAMEO (Conflict and Mediation Event Observations) is the standardized ontology that classifies political actions into a hierarchy ranging from verbal cooperation to material conflict, with a parallel scheme for actors. It is central because it makes machine-coded events comparable across projects and time and because the Goldstein conflict-cooperation weights are defined over its categories, turning categorical events into a continuous hostility scale.

How does this differ from datasets like the Correlates of War militarized disputes?

Militarized interstate dispute data are carefully hand-coded, low-frequency records of confrontations involving the threat or use of force, curated for accuracy and theoretical precision. Event data are automatically coded, high-frequency, broad-spectrum records derived from news, trading some accuracy and contextual judgment for scale, timeliness, and coverage. They are complementary: see the related Militarized Interstate Dispute Analysis entry.

Why is event data analysis treated as a natural language processing task?

Because its core technical work is information extraction from text: parsing sentences, resolving named entities (actors), classifying the relation (action) between them, and geolocating and timestamping the result. Modern pipelines use the full NLP toolkit — syntactic parsing, named-entity recognition, coreference resolution, and increasingly neural language models — which is why it sits in the NLP tasks leaf.

Sources

  1. 1.
    Schrodt, P. A., Davis, S. G., & Weddle, J. L. (1994). Political science: KEDS — A program for the machine coding of event data. Social Science Computer Review, 12(4), 561–588. See also Gerner, Schrodt et al. (1994), Machine coding of event data using regional and international sources, International Studies Quarterly, 38(1), 91–119.
  2. 2.
    Boschee, E., Lautenschlager, J., O'Brien, S., Shellman, S., Starz, J., & Ward, M. (2015). ICEWS Coded Event Data. Harvard Dataverse.

You have read it. What now?

Cite this page

ScholarGate. (2026, June 22). Event Data Analysis of Conflict. ScholarGate. https://scholargate.app/international-relations/event-data-conflict