Process / pipelineDisability StudiesCorpus linguistics / signed languagesPipeline

Sign Language Corpus Analysis

Also known as: Signed Language Corpus Linguistics, ID-Glossed Sign Corpus, Sign Language Annotation Pipeline, Deaf Corpus Linguistics

Sign language corpus analysis is the methodology for building and studying machine-readable, multimedia collections of signed languages, the natural visual-spatial languages of deaf communities. Because signed languages have no widely used written form, a corpus cannot be a body of text; it must be a structured collection of video recordings of deaf signers, layered with time-aligned annotation that makes the language searchable and analyzable. Trevor Johnston's 2010 account of moving from archive to corpus set out the central methodological principles, most notably ID-glossing, in which every instance of a sign is annotated with a single stable identifier linked to a lexical database so that all tokens of the same sign can be found and counted. The pipeline records signing on video, segments and ID-glosses it, links those glosses to a lexicon, aligns multiple annotation tiers to the video timeline, and then supports quantitative, corpus-based analysis of frequency, variation, and grammar. The result turns a fragile collection of recordings into reusable empirical evidence about how a signed language is actually used.

Key highlights

  • Makes a language with no written form searchable and countable through consistent ID-glossing linked to a lexicon.
  • Captures the multilinear nature of signing by aligning manual and non-manual annotation tiers to the video timeline.
  • Grounds claims about the language in real usage, replacing individual intuition with frequencies from authentic data.
  • Produces reusable, extensible resources that serve research, lexicography, education, and the deaf community over time.

Intuition

This section is available to Pro members. Upgrade to Pro

How it works

This section is available to Pro members. Upgrade to Pro

When to use it

Use sign language corpus analysis when you want to study a signed language empirically — its lexicon, grammar, or sociolinguistic variation — on the basis of how it is actually used rather than elicited intuitions, and when you can record and annotate authentic signing. It is the appropriate methodology for building reusable language resources, for documenting endangered or under-described signed languages, and for any quantitative claim about sign frequency, distribution, or variation. The approach demands substantial investment in video recording, consistent ID-glossing, a linked lexical database, and trained annotators, so it is not suited to quick or small projects, nor to questions answerable from a dictionary alone. It also requires deaf community involvement and ethical handling of identifiable video, and it is less applicable where the research question concerns sign perception or production at a level finer than lexical tokens, which calls for specialized phonetic or kinematic methods.

Strengths & limitations

Strengths
  • Makes a language with no written form searchable and countable through consistent ID-glossing linked to a lexicon.
  • Captures the multilinear nature of signing by aligning manual and non-manual annotation tiers to the video timeline.
  • Grounds claims about the language in real usage, replacing individual intuition with frequencies from authentic data.
  • Produces reusable, extensible resources that serve research, lexicography, education, and the deaf community over time.
Limitations
  • Annotation is extremely labor-intensive, and building a usable corpus from raw video requires large investments of expert time.
  • Consistency depends on disciplined ID-glossing and a stable lexical database; inconsistent glossing undermines all quantitative analysis.
  • Video data are identifiable, raising privacy and ethical demands that constrain recording, storage, and sharing.
  • Lexical-token-level glossing captures words and grammar but not fine phonetic or kinematic detail, which needs additional methods.

Common pitfalls

This section is available to Pro members. Upgrade to Pro

Applications

This section is available to Pro members. Upgrade to Pro

Frequently asked

Why can't a sign language corpus just be a collection of videos?

Because raw video is not searchable as language. A computer cannot find every instance of a particular sign in unannotated footage, so a mere archive of recordings supports no counting, searching, or quantitative analysis. To become a corpus, the video must be annotated so that the language it contains is identifiable and aggregable. Johnston's central point is that the move from archive to corpus is precisely this annotation step, with consistent ID-glossing linked to a lexicon being what turns a pile of recordings into reusable empirical evidence about the language.

What is ID-glossing and why does it matter so much?

ID-glossing means assigning each sign a single, stable identifier that is used for every occurrence of that lexeme throughout the corpus, drawn from a fixed lexical database rather than improvised. It matters because consistency is the foundation of quantitative analysis: if the same sign is glossed different ways by different annotators, its tokens cannot be aggregated and frequencies become meaningless. The ID-gloss is an identifier, not a translation, so it should not change with context. This discipline is what allows a sign language corpus to be searched and counted reliably.

Why does sign language annotation use multiple time-aligned tiers?

Because signing is multilinear: meaning is carried simultaneously by the hands and by non-manual features such as facial expression, head movement, and eye gaze. A single stream of annotation cannot represent things that happen at the same time, so the corpus uses parallel tiers — for each hand, for translations, for non-manual markers, and more — all aligned to the video timeline. This time-aligned, multi-tier structure preserves the simultaneity of the language and lets analysts study how manual and non-manual elements combine, which a flat transcription would lose.

Sources

  1. 1.
    Johnston, T. (2010). From archive to corpus: transcription and annotation in the creation of signed language corpora. International Journal of Corpus Linguistics, 15(1), 106-131.

You have read it. What now?

Cite this page

ScholarGate. (2026, June 23). Sign Language Corpus Analysis. ScholarGate. https://scholargate.app/disability-studies/sign-language-corpus