Chapter 10Interlinear Glossing, Abbreviations, and Citation Conventions
How to read the grammar's examples — the Leipzig-based interlinear format, morpheme segmentation, the gloss and abbreviation inventory, and the source-attribution sigla.
10.1 The interlinear format
Examples of Ainu data follow the three-line interlinear format of the Leipzig Glossing Rules (Comrie, Haspelmath & Bickel 2015). The morphemic line segments each Ainu word at morpheme and person-marker boundaries; the gloss line aligns one label with each morpheme; the free translation gives an idiomatic English rendering. A fourth line, the literal translation, is supplied where the morpheme-by-morpheme reading illuminates an opaque or partially grammaticalized form.
Two further lines may precede the morphemic line. A surface line gives the unsegmented written form when phonological alternation produces a surface string that differs from the underlying segmentation — for example, when a stem-initial vowel triggers reduction of a preceding person prefix. A source-script line gives the original katakana (or other) orthography when an example is transcribed from a katakana source, and is accompanied by the romanized segmentation immediately below it. The Latin phonemic transcription is described in Chapter 11 (The Latin Phonemic Transcription) and the katakana conventions in Chapter 12 (Katakana Orthography and the Extended Small-Kana Codas).
The morphemic and gloss lines must contain the same number of space-delimited tokens; each token groups one Ainu word together with its prefixes and suffixes, and the corresponding gloss token mirrors that internal structure using boundary symbols described below. A mismatch in token count causes a build error.
‘I heard it.’
constructed example
Illustrates the format: the morphemic line segments at person-marker (=) and affix (-) boundaries; the gloss line aligns one label per morpheme. The possessed evidential formal noun ruwe (ru-we, 'the trace of —') is discussed in the evidential chapters.
10.2 Morpheme boundary conventions
Three boundary symbols appear within tokens in the morphemic and gloss lines: Comrie, Haspelmath & Bickel (2015)
| symbol | use | morphemic | gloss |
|---|---|---|---|
- | affix boundary (derivational and nominal) | ru-we | track-poss |
= | person-marker boundary (orthographic convention, all personal indexes) | ku=nu | 1sg.a=hear |
. | portmanteau: two grammatical values fused in one gloss atom | a= | 4.a= |
The = symbol is used for all personal-index morphemes as an orthographic
convention of the Nakagawa lineage, not as Leipzig clitic notation — Nakagawa states
explicitly that his = is unrelated to the general-linguistic affix/clitic
display convention, and adopted it to keep inflectional person marking apart from
derivation and the dictionary headword recoverable Nakagawa (2024: 58–60).
Readers coming from general linguistics should register the departure: in the Leipzig
Glossing Rules = marks a clitic boundary (Comrie, Haspelmath & Bickel 2015), and
this grammar's = makes no such claim. Yu states the point directly, noting
that the Ainu equals sign is widely but wrongly assumed to descend from the Leipzig
clitic sign Yu (2025: §2.5.1):
人称接辞には,一般的な言語のハイフンとは異なり,慣習に従ってイコール「=」で表されることが多く、この表記法は言語学的に接語(clitic)を表す習慣であるイコールサインあるいはダブルハイフンに由来すると考えられることが多いが、実際には中川裕氏により広まった、接語とは無関係に人称接辞とその他派生接辞と区別するためだけの表記法であり、特に接辞と接語の違いを主張したわけではないと指摘されている。ただし実際には接辞か接語かは議論の余地がある。 The convention is not universal, and part of the field writes against it. Satō notates
the same morphemes with hyphens — ku-ye, a-kor Satō (2009: 53–54) — as does Ijäs's teaching
grammar Ijäs (2023), and Yu records both practices side by side Yu (2025: §2.5.1). Morphologically the set is mixed:
at least two members — the intransitive-subject markers =an and =as — show degrees of independence that set them apart from typical
inflectional affixes Nakagawa (2024: 49, 183); Bugaeva (2012: 472–473), while the prefixal markers pattern with ordinary
affixes. The choice of = throughout is thus deliberately neutral on which
markers are true clitics and which are affixes; that classification is examined in Chapter 69 (The Personal-Affix Template: Position Classes and Affix Ordering) ‹contested›.
The dot records a portmanteau — two grammatical values encoded in a single exponence — without claiming that they form a morphologically complex unit. In the fourth-person paradigm the prefix a= encodes the fourth person acting as the agent of a transitive verb; it is glossed 4.a, where 4 marks person and a marks the S/A/O role. Each component of every portmanteau is listed independently in the abbreviation set, so that 4.a, 1sg.a, and 2pl.o are all resolvable atom by atom.
10.3 Grammatical glosses and the abbreviation set
Every morpheme in the gloss line receives either a lexical gloss in lower case —
an ordinary English word for a stem or root (e.g. hear, house, track) — or a grammatical gloss in small capitals. Each small-capitals item is an abbreviation atom drawn from the fixed set tabulated in Chapter 174 (Abbreviations and Glossing-Symbol Conventions); no ad-hoc
or unlisted atom is permitted. When a grammatical category appearing in the data lacks a
registered atom, the atom is added to that table before the example enters any chapter. The
set is validated at build time: an unregistered capital-letter sequence in a gloss line causes
a build error.
The role labels s, a, and o follow the usage of Bugaeva (2012: 471): s is the single argument of an intransitive verb, a is the agent-like argument of a transitive verb, and o is the patient-like argument. Person is identified by the numerals 1, 2, 3, and 4 in the sense of Nakagawa (2024: 163): the numerals are form-class labels, not an inherent claim about the semantic content of each form. The fourth-person set — whose referential range covers the indefinite agent-defocusing use, the narrative first person, the addressee-inclusive 'we', the honorific second person, and the logophoric subject of reported speech — is the subject of Chapter 67 (The Indefinite/Fourth Person: Forms, Reference, and Agent-Defocusing) and Chapter 68 (Honorific and Logophoric Uses of the Fourth Person).
Evidential glosses follow the abbreviation set directly: evid marks a generic evidential category, infr the inferential, and rep the reportative. The formal nouns that build the nominalization-plus-copula evidential constructions (ruwe, siri, hawe, humi) are glossed by their compositional morpheme labels (poss, cop) rather than by a holistic evidential tag, since the grammaticalization of those forms is partial and the morpheme-level analysis remains transparent Dal Corso (2018: 24–25); this choice is consistent with the convention of Nakagawa (2024: 258). The evidential system is described in Chapter 120 (The Nominalization-plus-Copula Evidential Schema).
10.4 Dialect tags and example attribution
Every attested example carries a dialect tag drawn from the fixed set below. Finer locality — village, speaker, or archive signum — is recorded in a separate place field and does not modify the tag itself. The tag hk is used when the source does not specify a variety more precisely. Sakhalin (sa) and Kuril (ku) examples appear exclusively as labelled contrasts; Hokkaido is the variety under description throughout. The corpus and the dialect sample are described in Chapter 8 (The Oral-Literature Corpus and Spoken-Language Data) and Chapter 9 (The Dialect Sample and the Corpus-Quantification Method).
| tag | variety |
|---|---|
| HK | Hokkaido (variety not further specified) |
| SAR | Saru |
| CHI | Chitose |
| ISH | Ishikari |
| TOK | Tokachi |
| HOR | Horobetsu |
| SHI | Shizunai |
| ASA | Asahikawa |
| YAK | Yakumo |
| SA | Sakhalin (labelled contrast) |
| KU | Kuril (labelled contrast) |
Each attested example cites the underlying source — the narrator, edition, or author whose data it contains — rather than any aggregating database. A corpus example is attributed to the edition (Nakagawa (2024), Dal Corso (2018), Tamura (1996)) or to the narrator and archive slot directly. Works consulted only at second hand are flagged as reported in the bibliography and render with a distinguishing badge. An example constructed to illustrate a contrast or paradigm that is not directly quotable is marked constructed example; every example is either attributed or constructed, with no third category.
10.5 Source citation and the bibliography
In-text citations use the author–year format, with a colon and page number when a specific location matters: Nakagawa (2024: 163) or (Bugaeva 2012: 471). All cited keys are registered in a central bibliography; a key that does not appear in the registry causes a build error, so the citation apparatus is closed under the set of registered works.
The key scheme is lastnameYEAR: all-lowercase ASCII, no diacritics, no spaces.
Japanese authors are keyed by romanised surname with macrons stripped
(中川裕 → nakagawa2024; 佐藤知己 → sato2008). Multi-author works key on the first
author's surname. Collisions within the same author and year receive a lower-case suffix letter
in order of first citation (sato2023a, sato2023b); a topical suffix
is used instead when it aids orientation (bugaeva2021poss, bugaeva2021antip). The complete bibliography is gathered in Chapter 175 (Consolidated References and Bibliography).
Page references on examples follow the same colon-delimited syntax used in in-text citations, and may include section numbers, example numbers, or volume references where pagination is absent (e.g. Nakagawa (2024: §13.2)). Multiple sources may be joined with a semicolon within a single citation field.
10.6 Evidence-grade labels
Claims that go beyond the well-established consensus, or that represent genuinely contested positions in the literature, are tagged with a closed set of evidence-grade labels. The label is placed at the end of the sentence or clause it scopes over, in guillemets: ‹corpus-confirmed›. Unmarked prose is treated as consensus-level; only claims that depart from that default require a label.
| grade | meaning |
|---|---|
| consensus | all major sources agree; well-attested and uncontroversial |
| contested | sources genuinely disagree; the positions are presented without settling the question |
| corpus-confirmed | a claim verified against the corpus; the data bear it out |
| corpus-suggested | the corpus points toward the claim but the evidence is thin or indirect |
| speculative | plausible but unverified, e.g. a diachronic hypothesis or a gap-filling inference |
| original-needs-review | new analysis advanced in this grammar; flagged for editorial review |
The grade original-needs-review is restricted to the five domains in which new analysis is permitted: alignment, applicatives, noun incorporation, evidentiality, and clause linkage. The corpus-based grades presuppose the corpus described in Chapter 9 (The Dialect Sample and the Corpus-Quantification Method).
References cited in this chapter
Bugaeva (2012) ·Comrie, Haspelmath & Bickel (2015) ·Dal Corso (2018) ·Ijäs (2023) ·Nakagawa (2024) ·Satō (2009) ·Tamura (1996) ·Yu (2025)