Acceptability Judgment Task
Also known as: Acceptability Judgement Task, AJT, Sentence Acceptability Rating, Experimental Syntax Judgment Task
The acceptability judgment task is the modern, quantified successor to informal grammaticality judgments: instead of a single linguist marking a sentence grammatical or not, many participants rate carefully controlled sentences on a graded scale, and the ratings are analyzed statistically. Built on factorial designs with fillers and counterbalancing, and on response formats from Likert scales to magnitude estimation to forced choice, it turns intuition into replicable, gradient data. The approach anchors the experimental-syntax program associated with Jon Sprouse and colleagues, which tests grammatical hypotheses with the same methodological rigor as psycholinguistic experiments.
Key highlights
- Captures gradient acceptability and quantifies the size of a contrast, not just its direction.
- Replicable and statistically testable, with effect sizes, confidence, and crossed random effects for participants and items.
- Factorial designs with fillers and counterbalancing control confounds and mask the target, reducing strategic responding.
- Largely validates the traditional syntax literature while objectively flagging the minority of unreliable contrasts.
Intuition
This section is available to Pro members. Upgrade to Pro
How it works
This section is available to Pro members. Upgrade to Pro
When to use it
Use an acceptability judgment task when you want quantified, replicable evidence about gradient well-formedness and the size and reliability of a grammatical contrast — the standard for theoretical syntax claims, for adjudicating between competing analyses, and for cross-linguistic or cross-population comparisons. It is the right instrument when a contrast is subtle or contested, when you need effect sizes and statistics, or when reviewers expect formal confirmation of an intuition. It is less necessary for sharp, uncontroversial contrasts where informal judgments already converge, and it does not measure processing — for that, pair it with reading-time or eye-tracking methods.
Strengths & limitations
- Captures gradient acceptability and quantifies the size of a contrast, not just its direction.
- Replicable and statistically testable, with effect sizes, confidence, and crossed random effects for participants and items.
- Factorial designs with fillers and counterbalancing control confounds and mask the target, reducing strategic responding.
- Largely validates the traditional syntax literature while objectively flagging the minority of unreliable contrasts.
- Like all judgments, it taps acceptability — grammar plus performance — so processing and plausibility still color the ratings.
- Requires many participants, well-built materials, and statistical expertise, making it slower and costlier than informal judgments.
- Scale choice (Likert, magnitude estimation, forced choice) affects sensitivity and comparability across studies.
- Cannot localize the source of an effect in real time; it measures the endpoint, not the moment-by-moment processing difficulty.
Common pitfalls
This section is available to Pro members. Upgrade to Pro
Applications
This section is available to Pro members. Upgrade to Pro
Frequently asked
How does the acceptability judgment task differ from the older grammaticality judgment task?
They share a goal but differ in rigor and ontology. The grammaticality judgment task asks for a binary grammatical/ungrammatical verdict and was historically given by one linguist about the abstract grammar (competence). The acceptability judgment task asks many naive participants for graded ratings of how acceptable a sentence sounds, over factorially designed materials with fillers, and analyzes the ratings statistically. Because what people actually report is acceptability — grammar plus performance — the modern task is named accordingly, and grammaticality is inferred from the measured pattern.
Is this the same as the NLP 'linguistic acceptability' task?
No. The ScholarGate method 'linguistic-acceptability' refers to the computational NLP task in which a machine-learning classifier (trained, for example, on the CoLA benchmark) predicts whether a sentence is acceptable. The acceptability judgment task here is a behavioral, experimental method in which human participants rate sentences and researchers test grammatical hypotheses statistically. They are related in spirit — both about sentence well-formedness — but one models human judgments with an algorithm while the other elicits and analyzes the human judgments themselves.
Which response scale should I use — Likert, magnitude estimation, or forced choice?
It depends on the contrast and the goal. Magnitude estimation, introduced by Bard, Robertson, and Sorace, treats acceptability as a continuous ratio scale and can capture fine gradience, but it is demanding for participants and its ratio assumptions are debated. Likert scales are simple and, in method-comparison studies, prove highly sensitive. Forced-choice (two-alternative) tasks maximize statistical power for a single targeted contrast. Many experimental syntacticians now default to Likert or forced choice, reserving magnitude estimation for questions specifically about fine-grained gradience.
Sources
- 1.Sprouse, J., Schütze, C. T., & Almeida, D. (2013). A comparison of informal and formal acceptability judgments using a random sample from Linguistic Inquiry 2001–2010. Lingua, 134, 219–248.
- 2.Bard, E. G., Robertson, D., & Sorace, A. (1996). Magnitude estimation of linguistic acceptability. Language, 72(1), 32–68.
- 3.Schütze, C. T. (2016). The Empirical Base of Linguistics: Grammaticality Judgments and Linguistic Methodology. Language Science Press.
You have read it. What now?
Cite this page
ScholarGate. (2026, June 22). Acceptability Judgment Task. ScholarGate. https://scholargate.app/linguistics/acceptability-judgment-task