Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Text mining›Automated Essay Scoring (AES)
Process / pipeline

Automated Essay Scoring (AES)

Also known as: AES, automated writing evaluation, AWE, Otomatik Deneme Puanlaması

Automated Essay Scoring (AES) is a natural-language-processing task in which a computational model assigns scores to student-written essays across dimensions such as grammatical correctness, coherence, content richness, and organisation — replicating, at scale, what a human rater would do. The approach was formalised as a research field by Shermis and Burstein (2013) and has been transformed since 2019 by transformer language models, particularly BERT, which allow AES systems to leverage deep contextual representations of text.

ScholarGate
  1. Process / pipeline
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Automated Essay Scoring
BERT EmbeddingsReadability AnalysisSentiment AnalysisText Classification

When to use it

AES is appropriate when a labelled corpus of scored essays exists, the scoring rubric is clearly defined and stable, and the volume of essays to be scored is too large for manual scoring alone. The method suits large-scale standardised testing, intelligent tutoring systems that provide formative feedback, and educational research that requires consistent scoring across many participants. A minimum of around 30 scored essays is the bare threshold; in practice, several hundred to thousands of labelled essays are needed for reliable models. If no labelled scored data is available, or the rubric is ill-defined, AES cannot be applied.

Strengths & limitations

Strengths
  • Scales to arbitrarily large volumes of essays with consistent application of the scoring rubric.
  • Eliminates inter-rater drift and fatigue that affect repeated human scoring.
  • Transformer-based systems capture contextual, grammatical, and semantic signals simultaneously.
  • Enables rapid formative feedback loops in intelligent tutoring contexts.
Limitations
  • Requires a labelled corpus of scored essays; the system cannot score without prior human-rated training data.
  • Model performance is bounded by the quality and size of the labelled corpus.
  • Systems can be gamed by surface features such as essay length or rare vocabulary without genuine quality.
  • Bias toward dominant dialects or writing styles present in the training data is a documented risk.

Frequently asked

How many scored essays do I need to train an AES model?

The minimum threshold is around 30, but this is rarely sufficient for reliable predictions. In practice, several hundred labelled essays are needed for a traditional feature-based model, and several thousand for fine-tuning a transformer. Corpus size is the single strongest predictor of model quality.

Can AES replace human raters entirely?

AES is best used alongside human raters rather than as a complete replacement. High-stakes decisions — such as pass/fail in a standardised exam — should involve at least one human score, with AES serving as a second rater or as a filter for borderline cases. In low-stakes formative contexts, AES alone can provide useful feedback.

What metric should I use to evaluate an AES model?

Quadratic weighted kappa (QWK) is the standard metric because it accounts for ordinal agreement and penalises large disagreements more than small ones. Pearson correlation and mean squared error are reported alongside QWK. Accuracy alone is not appropriate for ordinal score scales.

Does BERT always outperform traditional feature-based methods?

On large, well-labelled corpora BERT-based models generally do outperform hand-crafted feature pipelines. On small corpora (under a few hundred essays) the advantage shrinks, and regularised regression on interpretable features can be competitive while offering greater transparency into which dimensions drive the score.

Sources

  1. Shermis, M.D. & Burstein, J. (2013). Handbook of Automated Essay Evaluation. Routledge. link ↗
  2. Devlin, J., Chang, M.-W., Lee, K. & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL-HLT, 4171-4186. DOI: 10.18653/v1/N19-1423 ↗

How to cite this page

ScholarGate. (2026, June 1). Automated Essay Scoring (AES). ScholarGate. https://scholargate.app/en/text-mining/automated-essay-scoring

Related methods

BERT EmbeddingsReadability AnalysisSentiment AnalysisText Classification

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • BERT EmbeddingsText mining↔ compare
  • Readability AnalysisText mining↔ compare
  • Sentiment AnalysisText mining↔ compare
  • Text ClassificationText mining↔ compare
Compare side by side →

Similar methods

Text Coherence ScoringAutomatic Text EvaluationFine-Tuned TransformerLinguistic Acceptability AssessmentFine-Tuned BERT-based ClassificationBERT Fine-TuningBERT-based ClassificationFine-Tuned Question Answering

Related reference concepts

Evaluation and AnnotationQuestion Answering and Dialogue SystemsAutomatic Speech RecognitionInformation ExtractionStylometry and Authorship AttributionMachine Translation

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Automated Essay Scoring (Automated Essay Scoring (AES)). Retrieved 2026-07-21 from https://scholargate.app/en/text-mining/automated-essay-scoring · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Shermis & Burstein (eds.); landmark consolidation 2013; deep-learning era from Devlin et al. 2019
Year
1966 (Project Essay Grade); modern deep-learning era from 2019
Type
Supervised text-regression / text-classification task
Input
Labelled student essays with human-assigned scores
Output
Numeric score or score-band per essay
Dimensions
Grammar, coherence, content richness, organisation
MinSample
30
Difficulty
2 / 3
Related methods
BERT EmbeddingsReadability AnalysisSentiment AnalysisText Classification
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account