Skip to content

Chapter 17The Vowel Inventory and Vowel Realization

The five-vowel triangular system /a e i o u/, the absence of phonemic length in Hokkaido, and allophonic centralization/devoicing.

17.1 The Five-Vowel System

Hokkaido Ainu has five vowel phonemes: /a/, /e/, /i/, /o/, /u/ Nakagawa (2024: 26); Satō (2008: §2.1); Shiraishi (2022: §2). The inventory is triangular, symmetric, and stable: no Hokkaido dialect reduces it to four vowels or expands it to six Nakagawa (2024: 25). Phonemic vowel length is absent from all Hokkaido dialects, distinguishing them from Sakhalin Ainu, which maintains contrastive length Nakagawa (2024: 27). The length/accent diachrony is discussed in Chapter 23 (The Pitch-Accent Placement Rule and the Accented/Accentless Dialect Split) and Chapter 24 (The Prosodic Unit and the Accent-versus-Tone Analysis).

Vowel phonemes of Hokkaido Ainu
phonemephonetic targetapproximate articulatory description
/a/[a]open, unrounded; cross-dialect stable
/e/[e̞]–[ɛ̝]mid front, slightly lowered; phonetically equivalent to Japanese え
/i/[i]close front, unrounded
/o/[o̞]–[ɔ̝]mid back, slightly lowered; phonetically equivalent to Japanese お
/u/[ɯ̹]close back, with stronger lip rounding than Japanese /ɯ/; auditorily near [o]

17.2 Phonetic Realization

Formant measurements of the five vowels are available from Obata & Teshima (1933; three male speakers, Asahikawa and Sakhalin, the first acoustic study of Ainu) and Ōhashi (1985; two female Saru speakers, spectrographic analysis), consolidated and discussed in Shiraishi (2022: §2). A nine-speaker study (Saru, Mukawa, Kushiro) by Tamura & Motohashi (2002a), also reported in Shiraishi (2022: §2), confirms the main findings: /i/ and /e/ have closely spaced F1–F2 values, and /u/ has a markedly lower F2 than Japanese [ɯ].

17.2.1 The mid vowels e and o

Neither /e/ nor /o/ occupies the higher position often assumed in traditional transcriptions. At least in the well-recorded Saru (Hioka Sadamo) speech, /e/ is realized as [e̞] to [ɛ̝] and /o/ as [o̞] to [ɔ̝] — positions equivalent to the standard Japanese mid vowels え and お aynu-corpora Discord (2023–2026) (nukopoli, 2024-11-08; corroborated by the Obata & Teshima and Ōhashi formant data, per Shiraishi (2022: §2)). The close F-space proximity of /i/ and /e/ accounts for the observation, recurring in older descriptions, that the two vowels are auditorily difficult to distinguish Shiraishi (2022: §2). Similarly, /u/ and /o/ are hard to separate in unaccented, lax positions Satō (2008: §2.1).

17.2.2 The back vowel u

/u/ is a close back vowel with lip rounding distinctly stronger than Japanese /ɯ/, and a lower F2, placing it further back in the vowel space than Japanese /u/ Nakagawa (2024: 26); Satō (2008: §2.1); Shiraishi (2022: §2). Nakagawa characterizes it as 円唇後舌母音 ('rounded back vowel'), and Satō (2008: §2.1) notes that /o/ is also protruded, making the distinction between /o/ and /u/ difficult in casual speech. To Japanese-speaking ears, /u/ sounds close to [o], which is why early Japanese records transcribed aynu 'person' as アイノ and inaw 'ritual stick' as イナオ Nakagawa (2024: 26).

(1)

[ái̯nʊ]

aynu person

‘person; human being’

Nakagawa 2024: 26; Hokkaido (dialect not further specified)

The historical Japanese transcription as アイノ reflects the back-rounded quality of /u/, which Japanese listeners perceived as near [o].

17.3 Vowel Devoicing

/i/ and /u/ devoice when flanked by voiceless consonants; /e/ is occasionally affected in the same environment. The phenomenon is well documented across Hokkaido dialects, with the earliest systematic account in Kindaichi (1931) and detailed treatment in Tamura (1998), both discussed in Shiraishi (2022: §2.1). Representative forms: ci̥se 'house', su̥tu 'root', se̥ta 'dog'.

The degree of devoicing shows heavy inter-speaker and inter-dialectal variation. Northern dialects devoice more frequently. A longitudinal comparison reported by Tamura (1998) — two Saru speakers from different generations — showed the younger speaker (b. 1912) devoicing approximately 25 times more often than the older (b. 1890) and in a wider range of phonological environments, including closed syllables under accent and consecutive vowels (eci̥ci̥snupehe 'your tears'); this is discussed in Shiraishi (2022: §2.1). The syllable types most prone to devoicing are si and ci; the types tu, pi, and pu are reported never to devoice — a contrast with Japanese, which devoices and in the same environment Shiraishi (2022: §2.1). A Chitose-dialect recording shows eci=kay realized as [et͡sɨ̥kaj], with devoiced [ɨ̥] in the ci sequence aynu-corpora Discord (2023–2026) (nukopoli, 2024-12-20) ‹corpus-confirmed›. L2 learners influenced by kana spelling tend to over-devoice, applying devoicing beyond its natural distribution (Tamura 1998: 55, cited in Shiraishi (2022: §2.1)).

The interaction of devoicing with consonant voicing and the /s/ palatalization pattern is discussed in Chapter 16 (The Consonant Inventory and Its Phonetic Realization) and Chapter 18 (/s/ and the s ~ š Palatalization Alternation).

17.4 Phonemic Length and Allophonic Lengthening

Hokkaido Ainu has no phonemic vowel length. All length observed in performance is allophonic, arising from two distinct sources Nakagawa (2024: 26–27); Shiraishi (2022: §2.2).

First, monosyllabic CV words lengthen the vowel in isolation or in pause-final position: kaː 'thread', toː 'lake', niː 'tree'. Second, the vowel of the accented syllable lengthens in polysyllabic words: ˈnuːman 'yesterday', poˈroː 'big'. Both processes are automatic and "not conducted consciously" by speakers (Kindaichi 1931: 4, cited in Shiraishi (2022: §2.2)); this accent-conditioned lengthening is why the yukar epic was formerly written ユーカラ (accent on yu) Nakagawa (2024: 26).

A grammaticalized emphatic lengthening, distinct from the allophonic process, breaks the extended vowel with a glottal onset: wen 'bad' → weʔen 'very bad'; to 'lake' → toʔo 'faraway' (Kindaichi 1931; Shiraishi (2022: §2.2)). The glottal element here marks the emphatic lexical expansion, not a phonemic long vowel.

Sakhalin Ainu, by contrast, maintains genuine vowel-length oppositions (e.g. ‹SA› niisah vs. nisah) corresponding to the Hokkaido pitch-accent contrast (HA nísap 'suddenly' vs. nisáp 'shin') Nakagawa (2024: 27); Nakagawa & Fukazawa (2022: §3.1). The reconstructed diachronic relationship — vowel length as the older opposition, later replaced by pitch-accent in Hokkaido — is argued at length in Itabashi (2001) and discussed in the prosody chapters Chapter 23 (The Pitch-Accent Placement Rule and the Accented/Accentless Dialect Split).

17.5 Vowel Sequences and Coda Glides

Sequences such as ay, aw, uy, ew, and oy are phonetically falling diphthongs — [ai̯], [au̯], [ui̯], [eu̯], [oi̯] — but their phonological analysis is not as diphthongs in the classical sense. The elements -y and -w behave as syllable-coda consonants in morphological tests (they trigger the consonant allomorph of the nominalizer -pe rather than the vowel allomorph -p: okay-pe, siw-pe) and cannot carry accent independently Nakagawa (2024: 29); Shiraishi (2022: §3). The full analysis — including the phonemic status of /y/ and /w/, glide insertion at hiatus, and the diphthong vs. coda-glide question — is given in Chapter 22 (Glides, Vowel Hiatus, and the Diphthong Question). Morphophonological consequences of vowel hiatus at morpheme boundaries are treated in Chapter 29 (Glide Epenthesis and Vowel-Hiatus Resolution).

17.6 Vowel Frequency

/a/ is the most frequent vowel in the Hokkaido corpus, a point in common with Japanese. /e/ is markedly more frequent than in Japanese, a characteristic that persists when person-affix tokens (a=, =an, and related forms) are excluded from the count aynu-corpora Discord (2023–2026) (nukopoli, 2024-09-25) ‹corpus-suggested›. The corpus tools available for this grammar count word tokens rather than phoneme occurrences, so these frequency ranks cannot be reproduced from per-segment counts at this writing; the underlying pattern is, however, consistent with the high frequency of the participial suffix -e, the applicative prefixes, and the evidential formal nouns all containing /e/.

17.7 The Vowel Co-occurrence Restriction

In two morphological contexts — the suffix that derives transitive verbs from intransitives, and the vowel suffix that forms the possessed (affiliative) form of nouns — the choice of suffix vowel is conditioned by the vowel of the root: kaykay-e 'break (tr.)', mosmos-o 'wake (tr.)', yakyak-u 'crush'; possessed tektek-e 'hand', nannan-u 'face' Chiri (1952); Shiraishi (2022: §4.5). Chiri (1952) named this pattern 母音調和 ('vowel harmony'), hypothesizing that the suffix was originally -i (the 3sg possessive), which underwent full assimilation to the root vowel Chiri (1952).

The characterization as vowel harmony is contested. The pattern is confined to two derivational contexts; it is absent from the plural suffix -pa (kom-pa, not *kom-po), from personal prefixes (ku-kor, not *ko-kor), and from underived forms; and no backness or tongue-root feature uniformly characterizes the alternating set Shibatani (1990: 15); Vovin (1993: 43); Shiraishi (2022: §4.5). The phenomenon is therefore better described as a morphologically restricted vowel co-occurrence pattern than as productive harmony ‹contested›. The morphophonological analysis of the possessed-form suffix, including the competing accounts of the assimilation mechanism, is pursued in Chapter 45 (Morphophonology of the Affiliative Suffix and Echo/Copy Vowels).

17.8 Dialect Variation in Vowel Quality

While the five-vowel inventory is constant across Hokkaido dialects, vowel quality shifts are attested in individual dialect records. In recordings of the Mukawa dialect (Katayama corpus), /i/ is realized as a lowered, fronted vowel approximating [e], and /e/ as a more open [ɛ]; the shift is general in the recording, though /i/ retains a canonical [i] realization outside the affiliative form in some tokens aynu-corpora Discord (2023–2026) (nukopoli, 2024-03-18) ‹corpus-suggested›. The pattern is consistent with a lowering chain in which the entire front series shifts one step downward; whether it is a stable feature of Mukawa or reflects idiolectal or recording conditions is not settled by the available evidence.

The Shizunai (静内) and Samani dialect areas are accentless (Chapter 23 (The Pitch-Accent Placement Rule and the Accented/Accentless Dialect Split)), which affects the realization of vowels under accent but does not alter the underlying inventory. Systematic acoustic comparisons of vowel space across Hokkaido dialect regions remain a gap in the literature Nakagawa & Fukazawa (2022: §3).

References cited in this chapter

aynu-corpora Discord (2023–2026) ·Chiri (1952) ·Itabashi (2001) ·Nakagawa (2024) ·Nakagawa & Fukazawa (2022) ·Satō (2008) ·Shibatani (1990) ·Shiraishi (2022) ·Vovin (1993)