Skip to contentScholarGate
LibraryBookshelfDeskReview StudioAssistant
Sign in
On this page
IntuitionHow it worksWhen to use itStrengths & limitationsCommon pitfallsApplicationsFrequently asked🔒 Read the full methodSourcesRelated methods
Cite this pageSpotted an issue on this page? Report or suggest a fix →
Home›Text mining›Linguistic Acceptability Assessment — Grammaticality Judgment
Process / pipeline

Linguistic Acceptability Assessment — Grammaticality Judgment

Linguistic Acceptability Assessment (Grammaticality Judgment) · Also known as: grammaticality judgment, acceptability judgment, CoLA task, Dilbilgisel Kabul Edilebilirlik Değerlendirme

Linguistic acceptability assessment is a natural-language-processing task that automatically estimates whether a sentence would be judged grammatically acceptable by a native speaker of the target language. Grounded in Chomsky's (1957) distinction between grammatical and ungrammatical utterances, the task was formalised as a neural benchmark by Warstadt, Singh and Bowman (2019) through the Corpus of Linguistic Acceptability (CoLA). It is used in language-learning research, linguistics studies, and quality auditing of natural-language-generation (NLG) systems.

ScholarGate
  1. Process / pipeline
  2. v1
  3. 2 Sources
  4. PUBLISHED
Cite this page →
Tools & resources
Download slides
Learn & explore

Read the full method

Members only

Sign in with a free account to read this section.

Sign in

Method map

The neighbourhood of related methods — select a node to explore.

Linguistic Acceptability Assessment
BERT EmbeddingsSentiment AnalysisText ClassificationTF-IDF

When to use it

Linguistic acceptability assessment fits when you need to automatically score the grammatical well-formedness of sentences in a target language and a model trained for that language is available. Typical applications include auditing NLG pipeline output, providing feedback in language-learning tools, and testing linguistic hypotheses computationally. The method requires a target-language model or annotated data; it cannot be applied to a language for which no trained model exists. The acceptability criterion must be measurable as binary or continuous, and the sentence sample should contain at least 20 items for meaningful evaluation.

Strengths & limitations

Strengths
  • Automates native-speaker acceptability intuitions at scale, enabling evaluation of large sentence sets without manual annotation.
  • Neural models trained on CoLA capture subtle grammatical contrasts that rule-based parsers miss.
  • Output can be binary (acceptable/unacceptable) or graded, making it flexible for both classification and regression tasks.
  • Applicable in language-learning feedback, NLG quality auditing, and computational linguistics without domain-specific rule writing.
Limitations
  • A trained model for the target language must exist; coverage for low-resource languages is limited.
  • Acceptability is gradient and context-dependent in natural language; a binary label may oversimplify borderline cases.
  • Neural models reflect the biases of their training data, which may not match the dialect or register of the sentences under evaluation.
  • A minimum of approximately 20 sentences is needed for stable evaluation; very small sets yield unreliable aggregate metrics.

Frequently asked

What is CoLA and why is MCC used instead of accuracy?

CoLA (Corpus of Linguistic Acceptability) is a benchmark of 10,657 English sentences from linguistics publications, labelled by expert judges as acceptable or unacceptable. Because roughly 69% of sentences are labelled acceptable, a model that always predicts 'acceptable' achieves 69% accuracy without learning anything. Matthews Correlation Coefficient (MCC) accounts for class imbalance and is therefore the standard metric on CoLA.

Can this method be applied to languages other than English?

Yes, but a model trained on acceptability judgments in the target language must exist. For languages beyond English, resources are limited; the analyst should verify that the model's training domain and register match the sentences under evaluation before drawing conclusions.

What is the difference between linguistic acceptability and grammaticality?

Grammaticality is a formal property of whether a sentence is generated by a given grammar. Acceptability is an empirical, gradient notion reflecting native-speaker intuitions, which can be influenced by frequency, processing difficulty, and context. The CoLA benchmark targets acceptability as judged by expert native speakers rather than formal grammaticality, making it a more realistic target for NLP systems.

How many sentences do I need to get meaningful results?

The method requires at least 20 sentences for stable aggregate metrics. With fewer sentences, individual label confidence matters more than aggregate MCC or accuracy. For robust evaluation — especially when comparing conditions — aim for several hundred sentences and report confidence intervals around aggregate metrics.

Sources

  1. Warstadt, A., Singh, A. & Bowman, S. (2019). Neural Network Acceptability Judgments. Transactions of the Association for Computational Linguistics, 7, 625–641. DOI: 10.1162/tacl_a_00290 ↗
  2. Chomsky, N. (1957). Syntactic Structures. Mouton, The Hague. ISBN: 978-9027933249

How to cite this page

ScholarGate. (2026, June 1). Linguistic Acceptability Assessment (Grammaticality Judgment). ScholarGate. https://scholargate.app/en/text-mining/linguistic-acceptability

Related methods

BERT EmbeddingsSentiment AnalysisText ClassificationTF-IDF

Which method?

Set this method beside its closest kin and read them side by side — the library lays the books on the table; the choice is yours.

  • BERT EmbeddingsText mining↔ compare
  • Sentiment AnalysisText mining↔ compare
  • Text ClassificationText mining↔ compare
  • TF-IDFText mining↔ compare
Compare side by side →

Similar methods

Grammaticality Judgment TaskAcceptability Judgment TaskPart-of-Speech TaggingDependency ParsingPOS TaggingAutomated Essay ScoringHallucination DetectionConstituency Parsing

Related reference concepts

Foundations of Computational LinguisticsEvaluation and AnnotationTreebanks and Annotated CorporaNatural Language ProcessingSyntactic ParsingMachine Translation

Spotted an issue on this page? Report or suggest a fix →

ScholarGate — Linguistic Acceptability Assessment (Linguistic Acceptability Assessment (Grammaticality Judgment)). Retrieved 2026-07-21 from https://scholargate.app/en/text-mining/linguistic-acceptability · Dataset: https://doi.org/10.5281/zenodo.20539026
Quick facts
Originator
Noam Chomsky (theoretical foundations, 1957); Warstadt, Singh & Bowman (neural formulation, 2019)
Year
1957 (theory); 2019 (neural benchmark — CoLA)
Type
NLP binary/continuous classification task
Benchmark
Corpus of Linguistic Acceptability (CoLA)
Output
Acceptability label (acceptable / unacceptable) or continuous acceptability score
MinimumSample
20
Related methods
BERT EmbeddingsSentiment AnalysisText ClassificationTF-IDF
ScholarGate

A content-first reference library for research methods — what each one is, how it works, and where it comes from.

Open data (CC-BY)

Explore

  • Library
  • Search the library…
  • Browse by field
  • Fields
  • Journey
  • Compare
  • Which method?

Reference

  • Subjects
  • Atlas
  • Glossary
  • Methodology
  • Philosophy

Your tools

  • Bookshelf
  • Desk
  • Chat

Company

  • About
  • Pricing
  • Contact
  • Suggest a method

Entries are compiled from published sources for reference. Verifying the accuracy and suitability of any information for your own use remains your responsibility.

© 2026 ScholarGate · A research-method reference library
  • Privacy
  • Cookies
  • Terms
  • Delete account