Chapter 9 explores prosodic structure as an integral component of linguistic structure. Prosodic structure specifies how phonological constituents are to be grouped to form larger units within a given utterance; this is known as their delimitative function. Prosodic structure also helps determine which of phonological constituents are produced with prominence relative to the other constituents; this is known as its culminative function. These functions entail strengthening of segmental realization (prosodic strengthening), often leading to linguistic enhancement of syntagmatic and paradigmatic contrast. Theories of the phonetics-prosody interface assume that phonetic realization of the spoken utterance is fine-tuned according to prosodic structure. In turn, crucial aspects of phonetic realization signal higher-order prosodic structure for listeners.
o/√ merger; and (4) at least for the urban Cheju speakers, the merger is best accounted for by the merger-by-transfer model, a unidirectional change in which one phonemic category becomes another (cf. Labov, 1994). Further, when our data are compared with other acoustic data available (including studies of the standard Korean in the 1960s and 1990s), it suggests that the directionality of the diachronic sound change is guided by both auditorily and articulatorily based principles such as contrast maximization and effort minimization principles.
This study examines articulatory characteristics of the three-way contrast in labial stops / p p h p */ (lenis, aspirated, fortis, respectively) in Korean in phrase-initial and phrase-medial prosodic positions with a two-fold goal. First, it investigates supralaryngeal articulatory reflexes of the stops and explores articulatory invariance of these stops across prosodic positions. Second, it investigates Korean stops in kinematic terms from the perspective of domain-initial strengthening, and explores the nature of prosodically-conditioned speech production from a dynamical perspective. Results showed that the articulatory reflex of the three-way contrast was invariantly observed across prosodic positions with lip constriction degree (/ p /</ p h /</ p */), while lip constriction duration showed a binary distinction (/ p /</ p h p */). Kinematically, there was only very weak articulatory evidence for the contrast across prosodic positions: The V-to-C lip closing movement tended to be faster for / p h / than for / p /, and the C-to-V lip opening movement tended to be larger for / p h p */ than for / p /. As for domain-initial strengthening, the consonantal lip closing gesture was characterized by a larger, longer and slower articulation, whereas the vocalic lip opening gesture (after the release) was larger and faster, but not longer. Kinematic relations indicated that the lip closing movement is most likely controlled by a rate of the clock (possibly modulated by a temporal modulation gesture, or π-gesture) comparable to boundary effects in English, but the boundary-induced lip opening movement was better accounted for by a change in target (possibly modulated by a spatial modulation gesture, or μ-gesture) which was comparable to prominence rather than boundary effects in English. The cross-linguistic difference was interpreted as coming from different prosodic systems between Korean and English, presumably instantiated in dynamical terms of how the temporal and the spatial modulation gestures are phased with constriction gestures in relation to boundary marking versus prominence marking.
The current work examines native Korean speakers' perception and production of stop contrasts in their native language (L1, Korean) and second language (L2, English), focusing on three acoustic dimensions that are all used, albeit to different extents, in both languages: voice onset time (VOT), f0 at vowel onset, and closure duration. Participants used all three cues to distinguish the L1 Korean three-way stop distinction in both production and perception. Speakers' productions of the L2 English contrasts were reliably distinguished using both VOT and f0 (even though f0 is only a very weak cue to the English contrast), and, to a lesser extent, closure duration. In contrast to the relative homogeneity of the L2 productions, group patterns on a forced-choice perception task were less clear-cut, due to considerable individual differences in perceptual categorization strategies, with listeners using either primarily VOT duration, primarily f0, or both dimensions equally to distinguish the L2 English contrast. Differences in perception, which were stable across experimental sessions, were not predicted by individual variation in production patterns. This work suggests that reliance on multiple cues in representation of a phonetic contrast can form the basis for distinct individual cue-weighting strategies in phonetic categorization.
This paper discusses some current issues regarding how prosodic structure is manifested in fine-grained phonetic details, how prosodically-conditioned articulatory variation is explained in terms of speech dynamics, and how such phonetic manifestation of prosodic structure may be exploited in spoken word recognition. Prosodic structure is phonetically manifested in prosodically important landmark locations such as prosodic domain-final position, domain-initial position and stressed/accented syllables. It will be discussed how each of the prosodic landmarks engenders particular phonetic patterns, how articulatory variation in such locations are dynamically accounted for, and how prosodically-driven fine-grained phonetic detail is exploited by listeners in speech comprehension.
No abstract is provided for this article.
Speech perception relies on multiple acoustic cues whose relative weighting varies across languages. The present study examines how long-term language experience shapes cue weighting in second-language (L2) speech perception, refining an attentional-learning account of cross-linguistic transfer. Native speakers of English, Dutch, Spanish, Korean, and Mandarin completed a cue-weighting task targeting English lexical stress, in which vowel quality, pitch, and duration were orthogonally manipulated. Results revealed robust, dimension-specific differences across first-language (L1) groups that could not be explained solely by the presence or absence of lexical stress or lexical tone in the L1. Instead, cue weighting reflected how acoustic dimensions function within the L1 cue ecology, including their relative contribution to lexical distinctions and the stability and interpretability of these mappings across contexts. Cue redundancy constrained relative cue strength without eliminating attentional sensitivity to secondary dimensions. Machine-learning classification further showed that L1-linked attentional profiles were sufficiently structured to support prediction, even among L2 listeners with substantial English proficiency, demonstrating the persistence of L1-shaped attentional tuning. These findings support a view of cue weighting as reflecting durable, multidimensional attentional priors shaped by long-term experience and highlight the importance of L1 cue ecologies in understanding cross-linguistic transfer in speech perception.
This paper examines the effects of morpheme boundaries on intergestural timing, and demonstrates that low-level phonetic realization is influenced by morphological structure, i.e. compounding and affixation. It reports two experiments, one using electromagnetic midsagittal articulography (EMA) and one electropalatography (EPG), examining Korean data. The results of the EMA study show that intergestural timing is less variable for adjacent gestures across the word boundary inside a lexicalized compound than inside a nonlexicalized compound, and inside a monomorphemic word than across a morpheme boundary. The EPG study (which examined the timing in palatalization of a coronal) shows that both [ti] and [ni] have more variability in gestural timing when heteromorphemic than when tautomorphemic. Furthermore, the phonetic details of gestural overlap shed light on the asymmetry on palatalization between tautomorphemic and heteromorphemic gestural sequences (e.g. ni vs. n-i), presumably driven by paradigmatic contrast and preference of overlap. In short, what emerges from two experiments is that gestures are coordinated more stably within a single lexical item (a morpheme or a lexicalized compound) than across a boundary between lexical items. In accounting for the stability of intergestural timing within a lexical entry, several hypotheses were discussed including the Phase Window, Bonding Strength, Phonological Timing and Extended Phase Window model newly proposed here. The implication is that the morphological structure may be encoded in the phonetic realization, as is the case with other linguistic structure (e.g. prosodic structure).
This study compares prosodic structural effects on nasal (N) duration and coarticulatory vowel (V) nasalization in NV (Nasal-Vowel) and CVN (Consonant-Vowel-Nasal) sequences in Mandarin Chinese with those found in English and Korean. Focus-induced prominence effects show cross-linguistically applicable coarticulatory resistance that enhances the vowel's phonological features. Boundary effects on the initial NV reduced N's nasality without having a robust effect on V-nasalization, whose direction is comparable to that in English and Korean. Boundary effects on the final CVN showed language specificity of V-nasalization, which could be partly attributable to the ongoing sound change of coda nasal lenition in Mandarin.
No abstract is provided for this article.
No abstract is provided for this article.
No abstract is provided for this article.
Prosodic structure has been assumed to serve as a frame for articulation, so that phonetic shaping of abstract phonological representations is fine‐tuned as a function of the prosodic system of the language. The intricate relationship between phonetics and prosodic structure has been explored in the literature under the rubric of the phonetics–prosody interface. This paper reviews various aspects of the phonetics–prosody interface and discusses how prosodic structure modulates phonetic realization within and across languages. A particular attention is paid to boundary‐related prosodic strengthening (i.e., spatial and/or temporal expansion of articulation that arises in the vicinity of prosodic junctures), especially in association with domain‐initial positions (also known as domain‐initial strengthening, DIS, effects). Prosodic boundary strengthening is further discussed in terms of how it is language‐specifically fine‐tuned, how it is understood in dynamical terms, and how it relates to linguistic functions (as syntagmatic vs. paradigmatic contrast enhancement) that are all further conditioned by other factors of the linguistic sound system of individual languages such as the prominence system and the phonetic feature system.
This study examined preboundary lengthening and other kinematic characteristics of articulatory gestures in CV.CV and CV.CVC before prosodic boundaries in Korean. Preboundary lengthening was found to be extended to initial syllables in both&nbsp;CV.CVand CV.CVC, while its magnitude was largest on the final syllable. The preboundary lengthening effect was also reflected in the time-to-peak velocity (acceleration duration), but only on gestures of the final syllable. Preboundary lengthening was accompanied by substantial increase in both displacement and peak velocity, showing domain-final articulatory strengthening. This articulatory strengthening effect on preboundary gestures (at the right edge of prosodic constituent) was largely dovetailed with the notion of an edge-prominence language where boundary marking is assumed to be closely related with prominence lending. These results were compared in two different conditions driven by information structure (‘new’ vs. ‘given’) and were discussed to understand the observed kinematic pattern in dynamical terms in the theoretical framework of the π-gesture model.&nbsp;
• Fourteen studies show that fine phonetic detail is systematically regulated in shaping sound systems. • Fine phonetic detail links production, perception, and learning, sustaining contrasts and enabling sound change. • Sound systems emerge from controlled allocation of continuous phonetic parameters within prosodic structure. • Prosodic structure guides segmental and suprasegmental realization across languages and domains. • Phonetic grammar is a language-specific, system-internal control system shaped by motor, perceptual, and cognitive pressures. This special issue examines how fine phonetic detail participates in the shaping of sound systems. Across fourteen studies, the central theme is that subtle temporal, spectral, and articulatory patterns are not incidental by-products of articulation, but are systematically regulated aspects of speakers’ phonetic knowledge. They provide the means through which phonological contrasts and prosodic structure are realized, maintained, and sometimes reorganized. The contributions show how languages allocate continuous phonetic parameters—such as timing, coordination, voice quality, and nasality—within prosodic domains (e.g., phrases, words, and syllables) and under general biomechanical and communicative pressures. Studies of Irish, Hawaiian, Japanese, and Mandarin illustrate how prosodic structure guides segmental and suprasegmental realization. Work on English, German, Danish, and Cantonese demonstrates how fine phonetic detail underlies patterns of variation and creates potential pathways for change. Production connects naturally to perception and learning: findings from English accent adaptation and Samoan iterated learning reveal how listeners stabilize or reinterpret detail, linking individual processing to community-level patterning. A set of studies on Italian, Korean, English, and L2 German show how prominence reorganizes cues across articulation, interaction, and acquisition, shaping how speakers signal and listeners recover linguistic structure. These studies converge on a view in which fine phonetic detail arises from a central phonetic component (or the phonetic grammar) of linguistic structure—controlled by speakers, shaped by universal motor and perceptual constraints, and continually adjusted through perception and learning. In this perspective, sound systems emerge from the interplay of these regulated patterns, which sustain contrasts, support communication, and open principled routes for change.