Chapter 9The Dialect Sample and the Corpus-Quantification Method
Which Hokkaido dialects are sampled, how Sakhalin/Kuril contrast is framed, and exactly how the grammar computes and reports corpus frequencies.
9.1 The Hokkaido dialect sample
Hokkaido Ainu was spoken in regionally differentiated forms whose systematic divergences bear on morphology, phonology, and lexical semantics alike. The grammar takes the Saru (SAR) and Chitose (CHI) varieties as its descriptive core. Saru is the most extensively documented Hokkaido dialect: Tamura's 1996 dictionary Tamura (1996) supplies a dense lexical record; Nakagawa's annotated oral-literature text collections from Nibutani and Biratori narrators give a genre-diverse discourse corpus Nakagawa (2000); and Nakagawa's 2024 grammar takes Saru as the default Hokkaido reference throughout Nakagawa (2024). Chitose is the second depth-documented variety: Nakagawa's 1995 dictionary provides the foundational Chitose lexical record Nakagawa (1995); Satō's 2008 grammar integrates Chitose systematically alongside Saru comparisons Satō (2008); and the NINJAL Glossed Audio Corpus of Folklore contributes additional Chitose material Nakagawa et al. (2016). Six further dialect areas enter the description where sources allow: Shizunai (SHI), documented most fully by Refsing (1986); Horobetsu (HOR), the subject of Takahashi's morphosyntactic studies Takahashi (2017)Takahashi (2018); Tokachi (TOK), whose dialect is documented in Takahashi's Tokachi example collection Takahashi (2014) and evidentiality and deixis studies Takahashi (2013)Takahashi (2011); Ishikari (ISH), described by Asai (1969); and the more sparsely attested Asahikawa (ASA) and Yakumo (YAK) areas. The tag HK marks data not assignable to a specific locality. Table 1 summarises the sample, its primary sources, and the relative depth of corpus coverage.
| Tag | Variety / area | Primary descriptive sources | Coverage depth |
|---|---|---|---|
| SAR | Saru valley (Nibutani, Biratori) | Tamura (1996); Nakagawa (2000); Nakagawa (2024); NINJAL corpus NINJAL (2003) | Descriptive core |
| CHI | Chitose | Nakagawa (1995); Satō (2008); Nakagawa et al. (2016); Bugaeva topical dictionary Bugaeva et al. (2015) | Descriptive core |
| SHI | Shizunai (Shinhidaka) | Refsing (1986); ILCAA Saru/Shizunai materials ILCAA Ainu materials (1976) | Supplementary |
| HOR | Horobetsu (Iburi) | Takahashi (2017); Takahashi (2018); Chiba University materials Chiba University (eds.) (2015) | Supplementary |
| TOK | Tokachi (Obihiro, Honbetsu) | Takahashi (2014); Takahashi (2013); Takahashi (2015) | Supplementary |
| ISH | Ishikari | Asai (1969) | Supplementary |
| ASA | Asahikawa | Scattered attestation; no dedicated grammar | Cited where documented |
| YAK | Yakumo | Scattered attestation; no dedicated grammar | Cited where documented |
The Hattori dialect dictionary Hattori (1964) and Satō's 2008 grammar provide comparative paradigm data for several of these areas even where no monographic description exists; those sources supply the supplementary attestation cited in phonological and morphological chapters. Tamura Masashi's grammatical description of the Shiranuka dialect Tamura, M. (2011) adds a further locality for isolated phenomena.
9.2 Dialect classification and the isogloss basis
The classification of Hokkaido Ainu dialects rests on a sequence of lexicostatistic and dialectometric studies beginning with Hattori and Chiri's 1960 founding survey Hattori & Chiri (1960), which applied the 200-item Swadesh list to 19 localities (13 Hokkaido, 6 Sakhalin) and established the basic quantitative structure of inter-dialect similarity. Asai's 1974 cluster analysis of the same matrix, extended with Chitose data, produced the widely cited major bipartition of Hokkaido dialects into a southwestern group centred on Saru and an eastern-northern group Asai (1974). Tamura's 1988 encyclopedia article formalised this as the NE/SW division Tamura (1988), a terminology adopted in subsequent descriptive literature. Ono's reanalysis of the Asai matrix with updated statistical methods partially confirms but also complicates that bipartition Ono (2020), and Kirikae's place-name isogloss study documents the geographical boundary of the par/car alternation across Hokkaido Kirikae (1994). The current state-of-the-art synthesis is Nakagawa and Fukazawa's chapter in the 2022 Handbook, which evaluates the evidence for proposed boundaries and recommends a cautious multi-feature approach Nakagawa & Fukazawa (2022). Fukazawa and Ono's 2024 study reconsiders those boundaries with additional data Fukazawa & Ono (2024), and Kasuga's 2026 dialectometric analysis extends coverage to 75 localities Kasuga (2026) ‹corpus-suggested›. The full treatment of Hokkaido dialect classification and its dialectometric basis is given in Chapter 159 (Hokkaido Dialect Classification and Dialectometry); the phonological microvariation documented in the corpus is discussed in Chapter 160 (Hokkaido Dialect Microvariation: Phonology and Morphosyntax).
For the purposes of chapter-level analysis, dialectal variation is introduced where a phenomenon behaves differently across the tagged varieties and where the difference is attested in the primary sources. Sporadic attestation for a peripheral area is noted but not treated as systematic evidence. Where no source distinguishes dialects, the description applies to the Saru–Chitose southern baseline and should be understood as such.
9.3 Sakhalin and Kuril Ainu as external contrast
Sakhalin Ainu (sometimes called Enciw; tag SA) and Kuril Ainu (tag KU) are not objects of description in this grammar; they appear as systematic external contrast points. Murasaki's 1979 grammar of Sakhalin Ainu Murasaki (1979), Dal Corso's 2021 re-edition with translation and grammatical notes Dal Corso (2021), Dal Corso's 2024 study of Piłsudski's corpus phonetics and morphosyntax Dal Corso (2024), and Tangiku's 2022 Handbook chapter on HA–SA differences Tangiku (2022) supply the Sakhalin data used for comparison. Piłsudski's 1912 text collection Piłsudski (1912) remains the primary philological witness for early-twentieth-century Sakhalin speech and contributes the SA examples cited in several chapters. Kuril material is attested only in historical sources and enters the grammar for specifically historical comparisons. Cross-dialect contrasts (SA or KU examples) are presented in tagged contrast display; their comparative status is indicated by the dialect tag alone, without repeated prose announcement. The dedicated comparative treatment is in Chapter 162 (Sakhalin and Kuril Ainu: The External Comparison).
9.4 The attested text corpus
Examples in this grammar are drawn from a body of attested Hokkaido Ainu discourse totalling approximately 167,890 sentences. The Saru sub-corpus is the largest (around 77,964 sentences), followed by the Shizunai/southern materials (approximately 43,295 sentences), a Mukawa component (approximately 13,424 sentences), and the Chitose sub-corpus (approximately 4,420 sentences), with smaller contributions from Tokachi, Horobetsu, and other localities; numbers reflect the alignment state at the time of authoring and will grow as further editions are added. The genre composition of the corpus is described in Chapter 8 (The Oral-Literature Corpus and Spoken-Language Data).
The component sources are:
- The NINJAL Corpus of Ainu Oral Literature NINJAL (2003), an aligned Hokkaido narrative corpus with audio, covering Saru and Chitose narrators (including 木村きみ and 小田イト).
- Nakagawa's annotated oral-literature text collection, volumes 1–24 Nakagawa (2000), covering uwepeker (prose tales) and kamuy yukar from Saru and Chitose narrators principally at Nibutani.
- The NINJAL Glossed Audio Corpus of Ainu Folklore Nakagawa et al. (2016), an online morphologically annotated corpus of Chitose and Saru folktales.
- The ILCAA aligned Saru/Shizunai audio-text materials from recordings of 1976–1984 ILCAA Ainu materials (1976).
- The Biratori Town Ainu oral-literature archive, Saru-dialect oral literature held by the National Ainu Museum Biratori Ainu oral literature (1969).
- Chiba University Ainu language materials, covering Horobetsu and adjacent areas, with narrators including Kanenari Matsu Chiba University (eds.) (2015).
- Kayano's conversational Ainu recordings Kayano (1987) and the Hokkaidō Utari Kyōkai Akor Itak textbook Hokkaidō Utari Kyōkai (1994), contributing conversational and pedagogical register.
- The Ainu Times (Ainu Koraci) learner newsletter Ainu Koraci (1991), providing Saru and Chitose texts by speakers including 神崎雅好 and 萱野志朗.
Each sentence in the corpus is associated with a source identifier, a narrator/author when
known, and a dialect tag. Where a specific document is cited in a chapter,
the example carries the appropriate cite attribute pointing to the underlying
published edition, never to an aggregating platform (see Chapter 10 (Interlinear Glossing, Abbreviations, and Citation Conventions)).
9.5 Corpus-frequency methodology
Frequency claims in this grammar rest on string searches over the aligned corpus described above. The primary unit of count is the clause or sentence boundary as aligned in each source edition; where alignment is unavailable, the typographic sentence of the published text serves as proxy. Counts are reported as raw token frequencies (attestation counts) alongside the total search space (number of sentences or clauses in the relevant sub-corpus), to allow readers to assess the density and dialect distribution independently.
Several caveats govern interpretation. String searches return all surface occurrences of a form, including homophonous lexical items; chapters note when a form is polysemous and flag the proportion of counts that can be unambiguously assigned to the grammatical use under discussion. The corpus is not uniformly dialect-balanced: Saru and Shizunai materials predominate, so raw frequency figures reflect that weighting. Where a claim requires genre-internal comparison — for instance, comparing evidential token frequencies in yukar versus conversational registers — the counts are restricted to the appropriate genre sub-corpus and the size of that sub-corpus is stated. Claims based on corpus counts are graded ‹corpus-confirmed› when the evidence is unambiguous across a sufficient token base, and ‹corpus-suggested› when the count supports the claim but the evidence is thin or restricted to one dialect or genre. The evidence-grade system is described in full in Chapter 10 (Interlinear Glossing, Abbreviations, and Citation Conventions).
Where a chapter reports a specific count, the format is: form, raw count, sub-corpus size, and source or genre label (for example, ruwe ne: 16,350 tokens in the 167,890-sentence Hokkaido corpus). Normalisation to occurrences per thousand sentences is supplied when comparing across sub-corpora of substantially different sizes.
9.6 Attribution: the original-source rule
Every attested example in this grammar cites its underlying source — the specific published edition in which the sentence appears, named with its narrator or author where identified — and never a secondary aggregation platform. A sentence extracted from the NINJAL corpus cites the underlying text edition and narrator; a sentence from the Biratori archive cites the specific document held by the National Ainu Museum; a dictionary gloss cites the dictionary (Tamura 1996, Nakagawa 1995, Kayano 1996) in whose pages it appears. Citing an aggregating database or web platform in place of the underlying published edition obscures evidentiary weight and fails the attribution rule.
Constructed examples — sentences composed to illustrate a grammatical point and not drawn
from any attested source — are explicitly marked constructed. There is no
intermediate category: an example either carries a cite attribute pointing to
a resolved bibliography key, or it is marked constructed. A grammaticality
judgment about a constructed sentence is the author's, derived from established patterns in
the descriptive literature; where the judgment is uncertain it is marked and sourced to
the nearest attested parallel. Narrator and locality information, when available, is given
in the place field of the example and is not repeated in the prose surrounding
it. The full glossing and citation conventions are set out in Chapter 10 (Interlinear Glossing, Abbreviations, and Citation Conventions).
9.7 Evidence grading
Every nontrivial descriptive or analytic claim in the grammar carries an explicit evidence grade drawn from a closed set. Uncontroversial, well-attested claims with consensus across the major sources carry no inline grade (the default is ‹consensus›). Claims that the major sources disagree on are marked ‹contested›, with the positions named and sourced. Claims checked against the corpus data are marked ‹corpus-confirmed› (evidence robust) or ‹corpus-suggested› (evidence thin or indirect). Plausible but unverified ideas, typically concerning diachrony or cross-dialectal inference beyond the attested record, are marked ‹speculative›. Novel re-analyses by this grammar's authors, permitted only within the five theory-magnet domains identified in Chapter 1 (Aims, Scope, and Design Philosophy of This Grammar), are marked ‹speculative› with the supporting reasoning attached. Comprehensive grading conventions and examples of each grade are given in Chapter 10 (Interlinear Glossing, Abbreviations, and Citation Conventions).
9.8 Community observations from the aynu-corpora archive
The aynu-corpora Discord community has accumulated a substantial body of grammatical observation — morpheme analyses, form-class proposals, productivity assessments, distributional generalisations — since approximately 2023 aynu-corpora Discord (2023–2026). These community observations enter the grammar when three conditions hold: the item bears on a grammatical point covered in a chapter, it can be responsibly stated within the grammar's evidentiary framework, and no prior peer-reviewed treatment addresses the specific point. Each incorporated observation is cited with the member handle and date, graded according to the confidence level recorded in the archive (asserted items corroborated by the corpus or literature receive a standard grade; proposed items are marked ‹contested›; speculative items are marked ‹speculative›; disputed items are noted as community-contested). Community observations are leads, not authorities; they supplement rather than displace the peer-reviewed descriptive literature.
References cited in this chapter
Ainu Koraci (1991) ·Asai (1969) ·Asai (1974) ·aynu-corpora Discord (2023–2026) ·Biratori Ainu oral literature (1969) ·Bugaeva et al. (2015) ·Chiba University (eds.) (2015) ·Dal Corso (2021) ·Dal Corso (2024) ·Fukazawa & Ono (2024) ·Hattori (1964) ·Hattori & Chiri (1960) ·Hokkaidō Utari Kyōkai (1994) ·ILCAA Ainu materials (1976) ·Kasuga (2026) ·Kayano (1987) ·Kirikae (1994) ·Murasaki (1979) ·Nakagawa (1995) ·Nakagawa (2000) ·Nakagawa (2024) ·Nakagawa & Fukazawa (2022) ·Nakagawa et al. (2016) ·NINJAL (2003) ·Ono (2020) ·Piłsudski (1912) ·Refsing (1986) ·Satō (2008) ·Takahashi (2011) ·Takahashi (2013) ·Takahashi (2014) ·Takahashi (2015) ·Takahashi (2017) ·Takahashi (2018) ·Tamura (1988) ·Tamura (1996) ·Tamura, M. (2011) ·Tangiku (2022)