Chapter 21Syllable Structure, Phonotactics, and Word-Edge Constraints
The canonical (C)V(C) syllable and obligatory onset, the permitted-coda inventory and onset/coda asymmetry, heterosyllabic clusters, and word-edge constraints including the citation-vs-combining shape mismatch (loanword adaptation treated in Part XXI).
21.1 The (C)V(C) template
The Hokkaido Ainu syllable has the shape (C)V(C): an optional onset consonant, a single vowel nucleus, and an optional coda consonant Nakagawa (2024: 41); Satō (2008: §2.3); Shiraishi (2022: §4.1). Four syllable types result:
| shape | template | example | IPA | gloss |
|---|---|---|---|---|
| bare vowel | V | a | [a] | perfective particle; 'oh' |
| open | CV | ka | [ka] | 'thread' |
| checked (VC) | VC | ik | [ik̚] | 'go up' |
| closed (CVC) | CVC | pet | [pet̚] | 'river' |
Every syllable contains exactly one vowel. Two or more consonants in sequence occur only at a syllable boundary, never within a single syllable: the shapes CCVC, CVCC, and CVVC are all excluded Nakagawa (2024: 42); Shiraishi (2022: §4.1). The word toop 'far away', which looks like CVVC in the Latin transcription, parses as two syllables to.op (CV.VC), the syllable boundary falling between the two successive vowels Nakagawa (2024: 42). The phonemic vowel inventory is described in Chapter 17 (The Vowel Inventory and Vowel Realization) and the consonant inventory in Chapter 16 (The Consonant Inventory and Its Phonetic Realization).
Nakagawa (2024: 25) characterizes Hokkaido Ainu as having relatively few phonemes and a straightforward syllable structure, noting that the five-vowel inventory is constant across all Hokkaido dialects — dialect differences are comparatively small. In contrast, Sakhalin Ainu (SA) allows a long-vowel nucleus alongside a coda consonant, producing CVVC syllables absent from Hokkaido: SA aa, ruu correspond to Hokkaido a, ru, and the SA length contrast niisah : nisah corresponds to the Hokkaido pitch-accent contrast nísap : nisáp Nakagawa (2024: 42–43). The diachronic SA-length ↔ HA-pitch correspondence is treated in Chapter 24 (The Prosodic Unit and the Accent-versus-Tone Analysis).
Whether every vowel-initial syllable has an underlying glottal stop /ʔ/ is disputed. Under the /ʔ/-as-phoneme view (Hattori 1961; Tamura, earlier work) the template reduces to just CV and CVC, eliminating V and VC shapes. Nakagawa rejects this on the grounds that no minimal pair exists in monomorphemic forms and that positing /ʔ/ requires many costly deletion rules Nakagawa (2024: 31–33), a judgement shared by Refsing Refsing (1986) and Chiri (1942) as reported by Shiraishi (2022: §4.3). The debate is covered in Chapter 20 (The Laryngeals: Glottal Stop, /h/, and the Final-h Question); the majority position — /ʔ/ as a phonetic boundary element rather than a phoneme — is followed here, retaining V and VC syllable shapes in the template ‹contested›.
21.2 Syllabification
Nakagawa (2024: 43) provides a three-step algorithm for assigning syllable boundaries in the Latin transcription:
- Split between two adjacent vowels: VV → V.V.
- Split between two adjacent consonants: CC → C.C.
- In a VCV sequence, split before the consonant: VCV → V.CV.
The algorithm applied to several representative words: Nakagawa (2024: §1.3)
| word | syllabification | shape sequence | gloss |
|---|---|---|---|
| teeta | te.e.ta | CV.V.CV | 'long ago' |
| aynu | ay.nu | CVC.CV | 'human' |
| kamuy | ka.muy | CV.CVC | 'deity' |
| cepkoyki | cep.koy.ki | CVC.CVC.CV | 'swim upstream' |
| irankarapte | i.ran.ka.rap.te | V.CVC.CV.CVC.CV | 'greet (someone)' |
(Nakagawa 2024: 43.) The algorithm treats /y/ and /w/ as consonants, so aynu = ay.nu (CVC.CV) and kamuy = ka.muy (CV.CVC). The arguments for treating /-y/ and /-w/ as coda consonants rather than as the second element of a diphthong are presented in Chapter 22 (Glides, Vowel Hiatus, and the Diphthong Question).
A minimal pair that turns on syllable count exists in Saru and Chitose: yayrayke 'commit suicide' = yay.ray.ke (CVC.CVC.CV, three syllables) versus yairayke 'give thanks' = ya.i.ray.ke (CV.V.CVC.CV, four syllables). The sequence -ay- in the first word is a single closed syllable (CVC); the -ai- in the second is two syllables separated by a morpheme boundary. These words carry different accent patterns that follow from the syllable count Nakagawa (2024: 42), a distribution governed by the iambic placement rule in Chapter 23 (The Pitch-Accent Placement Rule and the Accented/Accentless Dialect Split).
Syllabification applies across morpheme boundaries through resyllabification: a coda consonant on one morpheme is re-parsed as the onset of a following vowel-initial morpheme. The compound mat-ikor (woman + possessions) thus surfaces as macikor [ma.ʨi.kor], with the underlying coda /t/ shifting to the onset of the second syllable and palatalizing before /i/ Kindaichi & Chiri (1936: 1); Nakagawa (2024: 36):
macikor
‘women's possessions’
Nakagawa 2024: 36; Saru
The underlying boundary mat|ikor resyllabifies to ma.ci.kor [ma.ʨi.kor] (CV.CV.CVC). The coda /t/ shifts to onset position and undergoes the obligatory t+i → [ʨi] palatalization. The same surface form is attested throughout the Saru corpus, e.g. usa macikor usa tamasay 'various women's treasures and bead-necklaces'.
The palatalization rule t+i → [ʨi] is categorically obligatory wherever the coda /t/ of one morpheme and the onset /i/ of the next come into contact; it is covered in detail in Chapter 28 (Assimilation and Cluster Simplification: Coda /r/, Nasals, and Clusters). Resyllabification is blocked when the following vowel bears secondary accent or when the boundary separates a reduplicant from its base, where a glottal boundary element appears instead Shiraishi (2022: §4.4). The interaction of resyllabification with the morphophonology of personal affixes is treated in Chapter 31 (Personal-Affix Junctural Sandhi, =an/a= Allomorphy, and Connected-Speech Reduction).
21.3 Permitted coda consonants
Nine of the eleven consonant phonemes occur in coda position. The two excluded from coda position in Hokkaido Ainu are the affricate /c/ [ʨ] and the laryngeal /h/; the remaining nine — stops /p t k/, fricative /s/, nasals /m n/, rhotic /r/, and glides /y w/ — all appear as codas Nakagawa (2024: 27); Satō (2008: §2.3); Shiraishi (2022: §3).
| phoneme | coda realization | example | gloss |
|---|---|---|---|
| /p/ | [p̚] unreleased | cep | 'fish' |
| /t/ | [t̚] unreleased | pet | 'river' |
| /k/ | [k̚] unreleased | ik | 'go up' |
| /s/ | [ɕ] after /i/; [sʲ] elsewhere | as, cis | 'stand'; 'cry' |
| /m/ | [m] | hum | 'sound' |
| /n/ | [n]; [ŋ] before /k/; [m] before /p/, /m/ | pon, honkor | 'small'; 'be pregnant' |
| /r/ | [ɾ] with vocalic release | kor | 'have' |
| /y/ | [i̯] | kay, kamuy | 'break'; 'deity' |
| /w/ | [u̯] | ohaw, paw | 'soup'; 'mouth' |
The coda stops /p t k/ are realized as unreleased [p̚ t̚ k̚] — closure is maintained but no burst follows, a pattern similar to Korean final stops and unlike English Nakagawa (2024: 29); Shiraishi (2022: §3). In older recordings, a full release was sometimes heard; the degree of implosion also varies across speakers aynu-corpora Discord (2023–2026) (aomidori\_cs\_, 2023-12-05) ‹corpus-confirmed›. The coda allophony of /s/ — [ɕ] after the vowel /i/ versus a more front [sʲ] after other vowels — is treated in Chapter 18 (/s/ and the s ~ š Palatalization Alternation).
Coda /n/ undergoes place assimilation: it surfaces as [ŋ] before /k/ and as [m] before /p/ or /m/ Nakagawa (2024: 30). So honkor 'be pregnant' is pronounced [hoŋkor] and anpe 'true thing' is [ampe]. The orthographic convention records the underlying /n/ throughout. Coda /r/ carries a vocalic release — usually a copy of the preceding vowel — that was long confused with a final full vowel in early transcriptions; its phonemic status and sandhi behavior are covered in Chapter 19 (The Rhotic /r/ and Coda-r).
21.4 Lexical statistics
The frequency figures in this section are drawn from a syllable count over the headwords of three dictionaries of southern Hokkaido dialects: Tamura's Saru dictionary Tamura (1996), Nakagawa's Chitose dictionary Nakagawa (1995), and Kayano's Saru (Nibutani) dictionary Kayano (1996). After removing bound affixes, clitics, multi-word headwords, and forms containing letters outside the Latin transcription of the Hokkaido phoneme inventory, the three sources yield 13,891 distinct single-word types; 13,872 of these parse exhaustively into (C)V(C) syllables, giving 51,003 syllables in all. The 19 residues are transcription errors in the source files. Headwords include derived and compound stems, so the figures describe the lexicon as dictionaries record it; cluster counts in particular reflect morpheme seams as well as root-internal sequences.
| shape | count | share |
|---|---|---|
| CV | 29,247 | 57.3% |
| CVC | 14,199 | 27.8% |
| V | 5,905 | 11.6% |
| VC | 1,652 | 3.2% |
Open syllables outnumber closed ones by a little over two to one (68.9% against 31.1%), and syllables with an onset outnumber onsetless ones nearly six to one. Words run from one to twelve syllables, with three- and four-syllable words together making up over half the types; the mean is 3.7 syllables per word. About a third of word types (4,927 of 13,872) begin with a vowel.
Among nuclei, /a/ leads with 29.7% of all syllables, followed by /e/ (21.0%), /o/ (17.2%), /i/ (16.7%), and /u/ (15.5%). The most frequent onset is /k/ (9,277 of 43,446 filled onsets, 21.4%), with /r s p t n/ each between 4,600 and 5,200; /w/ is the rarest onset at 1,058. Coda frequency splits sharply by position in the word. Word-finally /r/ dominates (1,119 of 4,664 final codas), ahead of /p/ and /k/; word-internally /n/ leads (2,322 of 11,178 internal codas), followed by /y/ — a reflection of the many -Vn- and -Vy- sequences in derived stems. The sample contains no coda /h/ whatsoever, and the nine orthographic coda-/c/ tokens all trace to transcription errors in the source files, so the count confirms the categorical exclusion of /c h/ from coda position stated above.
Of the 99 heterosyllabic coda-onset combinations the nine codas and eleven onsets permit, 98 are attested at least once. The single systematic gap is *t.y, consistent with the affrication of /t/ + /y/ to /c/ at morpheme boundaries (see Chapter 28 (Assimilation and Cluster Simplification: Coda /r/, Nasals, and Clusters)). The distribution is heavily skewed: the ten commonest junctures (n.k 625, r.k 578, n.n 553, y.k 524, k.k 488, and so on) account for over 40% of all cluster tokens, while the rarest attested combinations (p.w, w.h) occur once each, at compound seams. Vowel hiatus (V.V) is likewise common — 2,376 sequences — with e.a the most frequent pairing (237) and i.i the least (24).
Weighting by running text instead of dictionary types changes the picture in instructive ways. A token count over the 168,302 Hokkaido sentences of the corpus described in Chapter 9 (The Dialect Sample and the Corpus-Quantification Method) — 1,057,229 parsed word tokens, 2,201,575 syllables, after excluding tokens in legacy orthographies whose digraph spellings syllabify falsely — gives CV 58.1%, CVC 19.6%, V 15.8%, VC 6.4%. The CV share is identical to the type count, but closed syllables drop from 31.1% to 26.0% and onsetless syllables rise from 14.8% to 22.2%, both driven by the high-frequency vowel-initial grammatical words (a=, an, or, oka). Word-final coda ranking reverses: /n/ leads with 117,224 tokens against 78,781 for /r/, where the type count has /r/ first — a frequency effect of words like an, wen, and pon against the many but individually rarer /r/-final content stems. Word-final /-m/ is likewise secure at token level: 12,605 tokens (1.2% of parsed tokens), led by isam (3,838), kam, kewtum, and hum. Residual coda-/c/ and coda-/h/ tokens number about 130 (0.006% of syllables) and trace to transcription noise, so the categorical coda ban holds at token level as well.
21.5 Phonotactic gaps in the syllable inventory
Several syllable shapes that the (C)V(C) template formally permits are absent or near-absent from the native lexicon. The best-established gaps, from Nakagawa's onset-×-vowel table (2024: 43–45), are the following:
| prohibited or rare shape | status | notes and sources |
|---|---|---|
| *ti | absent from root inventory | Nakagawa 2024: 43, 45 (Table 1 shows the ti cell empty). The surface affricate [ʨi] always derives from underlying /t/ + /i/ across a morpheme boundary — see Chapter 28 (Assimilation and Cluster Simplification: Coda /r/, Nasals, and Clusters). No HA root with stem-internal ti is attested. The three-dictionary headword sample (see the lexical statistics above) contains ti only in the transparent compound petikuswa (pet-ikus-wa), where the /t/ + /i/ contact sits on a morpheme seam. |
| *yi, *wi, *wu | absent as root-internal syllables | Nakagawa 2024: 43, 45: no word has these as its root; they arise only at morpheme boundaries by resyllabification or glide insertion Nakagawa (2024: 43, 45); Shiraishi (2022: §4.2). Whether they constitute independent phonemic syllables is contested between Tamura and Nakagawa — see Chapter 22 (Glides, Vowel Hiatus, and the Diphthong Question) ‹contested›. |
| cuC, ceC, coC | rare ‹corpus-suggested› | The corpus contains very few tokens of /c/ before /u e o/ followed by a coda; Nakagawa (2024: 43) notes that /c/ before /a/ or /u/ aligns more naturally with native-speaker perception than before other vowels, suggesting a partial front-vowel affinity for the affricate. The headword sample bears the asymmetry out: onset /c/ occurs before /i/ in 1,131 syllables and before /a/ in 316, against 128 for /u/, 116 for /e/, and 86 for /o/; with a coda added the counts fall to 81 (cuC), 53 (ceC), and 42 (coC) — rare but attested. |
| -ow | near-absent ‹corpus-suggested› | Two word-final -ow types occur in the 13,872-word headword sample, and a corpus survey returns almost no /-ow/ tokens aynu-corpora Discord (2023–2026) (nukopoli, 2024-12-19). A diachronic account proposes that earlier *-aw(e) shifted to -ew(e) in most environments aynu-corpora Discord (2023–2026) (antitwilight, 2024-12-19) ‹speculative›. |
| -m word-finally | not a gap ‹corpus-confirmed› | Orthographic /-m/ before /p/ or /m/ is an assimilation notation for underlying /-n/ (see above); at token level a stem-final /-m/ that stays /-m/ before other segments is very rare in the corpus aynu-corpora Discord (2023–2026) (nukopoli, 2024-12-19). Dictionary types tell against a categorical gap: 276 of the 13,872 headwords in the lexical sample end in -m, among them amam 'grain', isam 'not exist', and hum 'sound'. The token count settles it: 12,605 word-final /-m/ tokens occur in the Hokkaido corpus (1.2% of parsed tokens), 3,838 of them isam alone, so word-final /-m/ is unremarkable at both type and token level. |
Among the starred onset+vowel combinations, /-c/ and /-h/ are also categorically excluded from coda position (see the preceding section), and the /-c/ exclusion from coda means that *-ac, *-ec, etc., are all absent. The full discussion of the affricate /c/ and its distribution belongs to Chapter 16 (The Consonant Inventory and Its Phonetic Realization).
21.6 The *CVVC constraint and glide-blocking
The exclusion of CVVC syllables has a direct morphophonological consequence. In southern Hokkaido dialects (Saru, Chitose, Mukawa), the vowel /u/ of the first-person agent prefix ku= deletes before a stem-initial /a e u o/, and the vowel /i/ of an /i/-initial stem weakens to the glide /y/ in the same environment. The weakening applies when the stem's first syllable is open (CV onset + vowel): ku=ipe 'I eat a meal' surfaces as ku=ype [kúype], since the first syllable of ipe is open (V) and weakening to /y/ gives the licit shape CVC + CV (ku.ype).
When the stem's first syllable is closed (CVC), weakening is blocked. The form is ku=ikra [ku.ík.ra], not *ku=ykra: if /i/ weakened to /y/, the resulting sequence /kuykra/ would be unsyllabifiable under (C)V(C) — the three-consonant cluster /ykr/ cannot be distributed between two (C)V(C) syllables without creating either a CC onset or a CC coda Nakagawa (2024: 52–53); aynu-corpora Discord (2023–2026) (nukopoli, 2024-11-28) ‹corpus-confirmed›. Nakagawa's published example for this blocking is kuínkar (not *kuynkar) 'I look' (2024: 53): the first syllable of inkar is closed (/in/), so the hiatus is retained rather than resolved by weakening.
The contrast between the two environments:
| prefix | stem | first σ of stem | surface form | parse |
|---|---|---|---|---|
| ku= | ipe 'eat a meal' | open: /i/ (V) | ku=ype [kúy.pe] | CVC.CV — licit |
| ku= | inkar 'look' | closed: /in/ (VC) | ku=inkar [ku.ín.kar] | CV.VC.CVC — licit |
| ku= | ikra 'send' | closed: /ik/ (VC) | ku=ikra [ku.ík.ra] | CV.VC.CV — licit |
| ku= | ikra 'send' | closed: /ik/ (VC) | *ku=ykra | *unsyllabifiable — illicit |
The same principle holds for the second-person prefix e=: e=itak → eytak 'you speak' (open stem, weakening permitted), but e=ikra [e.ík.ra] 'you send' (closed stem, hiatus retained) Nakagawa (2024: 40, 52–53). The full allomorphy paradigm for ku=, e=, and ci= under vowel deletion and glide-weakening is set out in Chapter 31 (Personal-Affix Junctural Sandhi, =an/a= Allomorphy, and Connected-Speech Reduction).
21.7 Loanword syllable adaptation
Loanwords that contain consonant clusters or CVVC structures are adapted to fit the (C)V(C) template by vowel epenthesis. The epenthetic segment is typically /o/ or a copy of the nearest adjacent vowel aynu-corpora Discord (2023–2026) (nukopoli, 2024-12-11) ‹corpus-suggested›. In Hokkaido the donor language is Japanese; Russian borrowings belong to the Sakhalin varieties, so the clearest illustration of the repair comes from there ‹SA›: potoloko (← Rus. потолок 'ceiling') breaks up both a word-initial CC onset (/pt/) and a word-final CC cluster (/lk/) with epenthetic /o/. Its single corpus attestation is Sakhalin, in Sentoku Tarōji's letters — keta potoloko kasiketa okay=ahci ‹corpus-confirmed›. A hypothetical English loanword like Christmas would enter as something like kurisumas, with each cluster split by an inserted vowel Nakagawa (2024: 42).
The identity of the epenthetic vowel is not always /o/: a copy of the adjacent vowel also occurs, and the choice may reflect the phonological neighborhood rather than a single underlying default. Loan adaptation across contact periods — Japanese into Hokkaido Ainu, Russian into the Sakhalin varieties — shows the (C)V(C) template operating as a consistent filter, with epenthesis as the principal repair strategy ‹corpus-suggested›.
References cited in this chapter
aynu-corpora Discord (2023–2026) ·Kayano (1996) ·Kindaichi & Chiri (1936) ·Nakagawa (1995) ·Nakagawa (2024) ·Refsing (1986) ·Satō (2008) ·Shiraishi (2022) ·Tamura (1996)