Abstract and Keywords
This chapter analyzes the specific characteristics of corpora of endangered languages from a corpus linguistic perspective. Therefore it starts with a definition of the central notions of corpus and text and then investigates how the heterogeneous language documentation corpora may fit into a general typology of corpora. The third section looks at the genres and registers that for methodological and theoretical reasons are typical for language documentations, whereas the fourth section deals with the structure of corpora and how texts of a particular content, genre or register can be accessed in archives. The format of the texts, which are typically annotated audio and video recordings, is described in the fifth section and deals with metadata, transcription, orthography, translation, glossing, and syntactic annotation. How annotated corpora can be analyzed for grammatical and lexical research is shown in the sixth section. The last section summarizes the specific features of language documentation corpora.
Access to the complete content on Oxford Handbooks Online requires a subscription or purchase. Public users are able to search the site and view the abstracts and keywords for each book and chapter without a subscription.
If you have purchased a print title that contains an access token, please see the token for information about how to register your code.