This study was supported by the National Research Foundation of Korean Grant funded by the Korean Government (NRF2013S1A2A2035410) to Taehong Cho.
Recent studies on perceptual learning have indicated that listeners use intermediate units between the acoustic input and lexical representations of words. The same paradigm may also reveal the nature of these intermediate units based on patterns of generalization of learning. We here test whether learning generalizes to other units of the same underlying or surface representation. This was achieved by exposing listeners to tensified Korean stops (i.e., underlying plain stops produced as tense due to a phonological process) and testing the consequences for later presented underlying tense or plain stops. Our results show that learning generalizes to underlying tense stops, while generalization to underlying plain stops could not be found. This indicates that the difference in the underlying phonological representation as tense or plain do not hinder learning as long as there is phonetic similarity on the surface.
In this special collection entitled Marking 50 Years of Research on Voice Onset Time and the Voicing Contrast in the World's Languages, we have compiled eleven studies investigating the voicing contrast in 19 languages. The collection provides extensive data obtained from 270 speakers across those languages, examining VOT and other acoustic, aerodynamic and articulatory measures. The languages studied may be divided into four groups: 'aspirating' languages with a two-way contrast (English, three varieties of German); 'true voicing' languages with a two-way contrast (Russian, Turkish, Brazilian Portuguese, two Iranian languages Pashto and Wakhi); languages with a three-way contrast (Thai, Vietnamese, Khmer, Yerevan Armenia, three Indo-Aryan languages, Dawoodi, Punjabi and Shina, and Burushaki spoken in India); and Indo-Aryan languages with a more than three-way contrast (Jangli and Urdu with a four-way contrast, and Sindhi and Siraiki with a five-way contrast). We discuss the cross-linguistic data, focusing on how much VOT alone tell s us above the voicing contrast in these languages, and what other phonetic dimensions (such as consonant-induced F0 and voice quality) are needed for a complete understanding of laryngeal contrast in these languages. Implications for various issues emerge: universal phonetic feature systems, effects of language contact on linguistic levelling, and the relation between laryngeal contrast and supralaryngeal articulation. The cross-linguistic VOT data also lead us to discuss how the distribution of VOT as measured acoustically may allow us to infer the underlying articulation and how it might be approached in gestural phonologies. The discussion on these multiple issues sparks new questions to be resolved, and provide indications of where the field may be best directed in exploring laryngeal contrast in voicing in the world's languages.
This study investigates how second-language (L2) listeners from five first-language (L1) backgrounds—English, Dutch, Mandarin, Spanish, and Korean—perceive English lexical stress, focusing on their use of vowel quality, pitch, and duration cues. Participants completed a cue-weighting perception task (Tremblay et al., 2021) in which two acoustic dimensions were manipulated orthogonally while the third was neutralized. Data for Dutch listeners come from the original study. Predictions about cross-linguistic transfer were based on the functional weight of each cue in the L1. The following L1 effects were predicted: For vowel quality: English, Mandarin > Dutch > Spanish, Korean; for pitch: Mandarin > Korean > Dutch, Spanish > English; for duration: English, Mandarin > Dutch, Spanish > Korean. Bayesian mixed-effects models tested the effects of cues and L1 with L2 proficiency (Lemhöfer & Broersma, 2012) as a covariate. The results aligned broadly with our predictions: for vowel quality, English > Mandarin > Dutch > Korean > Spanish; for pitch: Mandarin > Korean, Dutch > Spanish > English; for duration: English, Mandarin, Dutch > Spanish > Korean. These findings support a cue-weighting typology shaped by L1-specific cue prominence, with implications for theories of transfer and perceptual learning in L2 acquisition.
Application of a phonological rule is often conditioned by prosodic structure, which may create a potential perceptual ambiguity, calling for phonological inferencing. Three eye-tracking experiments were conducted to examine how spoken word recognition may be modulated by the interaction between the prosodically-conditioned rule application and phonological inferencing. The rule examined was post-obstruent tensing (POT) in Korean, which changes a lax consonant into a tense after an obstruent only within a prosodic domain of Accentual Phrase (AP). Results of Experiments 1 and 2 revealed that, upon hearing a derived tense form, listeners indeed recovered its underlying (lax) form. The phonological inferencing effect, however, was observed only in the absence of its tense competitor which was acoustically matched with the auditory input. In Experiment 3, a prosodic cue to an AP boundary (which blocks POT) was created before the target using an F0 cue alone (i.e., without any temporal cues), and the phonological inferencing effect disappeared. This supports the view that phonological inferencing is modulated by listeners' online computation of prosodic structure (rather than through a low-level temporal normalization). Further analyses of the time course of eye movement suggested that the prosodic modulation effect occurred relatively later in the lexical processing. This implies that speech processing involves segmental processing in conjunction with prosodic structural analysis, and calls for further research on how prosodic information is processed along with segmental information in language-specific vs. universally applicable ways.
List of Figures Acknowledgements Preface Chapter 1. Introduction 1.1. Background 1.2. Limitations of Previous Studies 1.3. Research Questions and Theoretical Background 1.3.1. Questions regarding articulation at domain-edges 1.3.2. Questions regarding strengthening and coarticulation 1.3.3. Questions Regarding Effects of Prosody on Kinematic Variations and Dynamical Accounts 1.3.3.1. Task dynamic model and dynamical parameters 1.3.3.2. Some previous studies 1.3.3.3. Stiffness and the p-gesture 1.3.4. Prosodic Organization 1.5. Outline of The Dissertation Chapter 2. Methods 2.1. Speech Material 2.2. Speakers 2.3. Procedures 2.4. Measurements 2.4.1. Measurements in Chapter 3: Maximum positions of the tongue, jaw, and lips. 2.4.2. Measurements in Chapter 4: Tongue position differences 2.4.3. Measurements in Chapter 5: Lip kinematics 2.4.4. Measurements in Chapter 6: Tongue kinematics 2.5. Prosodic Transcription 2.6. Statistics Chapter 3. Effects of Accent and Prosodic Boundary on Articulatory Maxima: Tongue, Jaw, and Lip 3.1. Hypotheses 3.2. Measurements and Statistics 3.3. Result 1: Articulatory Maxima in Preboundary Domain-final) Vowels (V1#) 3.3.1 Effect of Accent on V1 3.3.1.1. Accent effect on V1 maximum tongue position 3.3.1.2. Accent effect on V1 lip and jaw opening maxima 3.3.2. Effect of Boundary Type on V1 3.3.2.1. Boundary effect on V1 maximum tongue position 3.3.2.2. Boundary effect on V1 lip and jaw opening maxima 3.3.3. Effects of the neighboring vowel (V2) on V1 3.4. Result 2: Articulatory Maxima in Postboundary (domain-initial) Vowels (#CV2) 3.4.1. Effect of Accent on C2V2 3.4.1.1. Accent effect on CV2 maximum tongue position 3.4.1.2. Accent effect on V2 lip and jaw opening maxima 3.4.2. Effect of Boundary Type on C2V2 3.4.2.1. Boundary effect on C2V2 maximum tongue position 3.4.2.2. Boundary effect on V2 lip and jaw maxima 3.4.3. Effects of the neighboring vowel (V1) on V2 3.5. Summary & Discussion 3.5.1. Characteristics of Accent and Featural Enhancement 3.5.2. Asymmetry between domain-initial and domain-final articulations 3.5.3. Boundary-induced vs. Accent-induced Articulatory strengthening Chapter 4. Prosodically Conditioned V-to-V Coarticulatory Resistance 4.1. Hypotheses 4.2. Measurements and Statistics 4.3. Results 4.3.1. Carryover effect 4.3.1.1. Postboundary (domain-initial) /a/ in /i#a/ 4.3.1.2. Postboundary (domain-initial) /#i/ in /a#i/ 4.3.1.3. Summary of carryover coarticulatory effect 4.3.2. Anticipatory Effect 4.3.2.1. Preboundary (domain-intial) /a/ in /a#i/ 4.3.2.2. Preboundary (domain-final) /i/ in /i#a/ 4.3.2.3. Summary of anticipatory coarticulatory effect 4.3.3. The Reciprocal coarticulatory effect 4.3.3.1. Effect of accent 4.3.3.2. Effect of Boundary Type 4.3.3.3. Effect of Vowel Sequence (/a#i/ vs. /i#a/) 4.3.3.4. Interactions 4.3.3.5. Summary of reciprocal coarticulatory effect 4.3.4. Temporal distance and coarticulatory resistance 4.4. General Discussion 4.4.1. Coarticulatory resistance as a function of accent 4.4.2. Coarticulatory resistance across a prosodic boundary 4.4.3. Accent- vs. boundary-induced coarticulatory resistance 4.4.4. Coarticulatory resistance vs. coarticulatory aggression 4.4.5. Inherent articulatory properties and coarticulatory resistance 4.4.6. Revised Window Model Chapter 5. Effects of Prosody on Lip Movement Kinematics 5.1. Hypotheses 5.2. Measurements and Statistics 5.3. Results 5.3.1. Domain-final C1-to-V1 lip opening gesture 5.3.1.1. Displacement 5.3.1.2. Peak Velocity 5.3.1.3. Total Movement Duration: C1ONS-TO-V1TARG 5.3.1.4. C1ONS-To-V1PKVEL (time-to-peak-velocity) and V1PKVEL-To-V1TARG (deceleration duration) 5.3.2. Discussion about dynamics underlying kinematic differences for C1-to-V1 lip opening gesture 5.3.2.1. Dynamical aspect of accent effects on C1-to-V1 (C1V1# ) lip opening gesture 5.3.2.2. Dynamical aspect of boundary effects on C1-to-V1 lip opening gesture 5.3.3. Domain-initial C2-to-V2 lip opening gesture 5.3.3.1. Displacement (#C2V2) 5.3.3.2. Peak Velocity (#C2V2) 5.3.3.3. Total Movement Duration: C2ONS-TO-V2TARG (#C2V2) 5.3.3.4. C2ONS-To-V2PKVEL and V2PKVEL-To-V2TARG (#C2V2) 5.3.4. Summary and discussion about dynamics underlying kinematic differences for C2-to-V2 lip opening gesture 5.3.4.1. Dynamical aspect of accent effects on C2-to-V2 (#C2V2) lip opening gesture 5.3.4.2. Dynamical aspect of boundary effects on C2-to-V2 (#C2V2) lip opening gesture 5.3.5. V1-to-C2 Closing Gestures 5.3.5.1. Displacement (V1#C2) 5.3.5.2. Peak Velocity (V1#C2) 5.3.5.3. Total Movement Duration: V1ONS-TO-C2TARG 5.3.5.4. V1ONS-To-C2PKVEL and C2PKVEL-To-C2TARG 5.3.6. Summary and discussion about dynamics underlying kinematic differences for V1-to-C2 lip closing gesture 5.3.6.1. Dynamical aspect of V1 Accent effects on V1-to-C2 lip closing gesture 5.3.6.2. Dynamical aspect of V2 Accent effects on V1-to-C2 lip closing gesture 5.3.6.3. Dynamical aspect of boundary effects on V1-to-C2 lip closing gesture 5.4. General Discussion 5.4.1. What are the accent-driven kinematic characteristics underlying lip opening and closing gestures? 5.4.2. Can accent-driven kinematic variations be modeled by a particular dynamical parameter setting? 5.4.4. What are the boundary-driven kinematic characteristics underlying lip opening and closing gestures? 5.4.5. Can boundary-driven kinematic variations modeled by a particular dynamical parameter setting? 5.4.6. Articulatory signature of prosodic structure Chapter 6. Effects of Prosody on V1-to-V2 Tongue Movement Kinematics 6.1. Hypotheses 6.2. Measurements and Statistics 6.3. Results 6.3.1. Effects of Accent 6.3.1.1. Effect of Accent on vertical (y) tongue movement 6.3.1.2. Effect of Accent on horizontal (x) tongue movement 6.3.2. Discussion of dynamics underlying ACC/UNACC differences in kinematics 6.3.2.1. Dynamical aspects of accent for the /i/-to-/a/ vocalic gesture in the y dimension. 6.3.2.2. Dynamical aspects of accent for the /a/-to-/i/ vocalic gesture in the x dimension. 6.3.3. Effects of Boundary Type 6.3.3.1. Effects of Boundary Type on the vertical (y) tongue movement 6.3.3.2. Effects of Boundary Type on the horizontal (x) movement 6.3.4. Discussion of dynamics underlying boundary-induced differences in kinematics 6.3.4.1. Dynamics of boundary-induced kinematic differences in the y dimension 6.3.4.2. Dynamics of boundary-induced kinematic differences in the x dimension 6.4. General Discussion 6.4.1. Accent-induced kinematic variation and task dynamics 6.4.1.1. Accent-induced kinematic variation and strengthening 6.4.1.2. Dynamical accounts 6.4.2. Boundary-induced kinematic variation and task dynamics 6.4.2.1. Boundary-induced kinematic variation and strengthening 6.4.2.2. Dynamical accounts 6.4.2.3. The p-gesture and pre- and postboundary lengthening 6.4.3. Accent- vs. boundary-induced kinematic patterns and articulatory strengthening Chapter 7. Conclusion 7.1. Articulatory strengthening in accented syllables 7.2. Articulatory Strengthening at Domain-edges 7.3. Accent- vs. Boundary-induced Articulatory Strengthening 7.4. Consonantal vs. vocalic domain-initial strengthening 7.5. Articulatory strengthening and its linguistic significance 7.6. Other implications 7.6. Closing remarks Appendix I: The Speech Corpus, Appendix II: R2 Values for Lip Movement Temporal Relationships References, Index
A recent study (Kim & Cho, 2013, The Journal of the Acoustical Society of America) reported that the perception of a prosodic boundary leads to a shift in a stop-identification function in English, so that stops with a relatively long VOT are accepted as voiced if occurring after a major prosodic boundary. Even Korean learners of English showed such a shift. This shift would seem to result from compensation for post-boundary lengthening effects (or domain-initial strengthening) and thereby help to overcome the invariance problem in speech perception. In two experiments, we ask how this effect comes about. The first experiment tested whether a simple adjustment to a change in overall speaking rate would be sufficient to account for the shift. Results showed that while the global speaking-rate change modulates phonetic categorization in a similar way as a change in the prosodic boundary strength, the speaking-rate effect is not sufficient to explain the boundary effect. That is, there was a more robust shift in a stop identification function with localized slowing down of the final syllable due to an intonational phrase (IP) boundary than with global slowing down of speaking rate. The second experiment therefore investigated the contribution of an F0 cue to the observed perceptual shift and found that the presence or absence of the F0 cue did not mediate the effect of prosodic boundaries on phonetic categorization. This suggests that a perception shift in phonetic categorization stems primarily from the listeners' adjustment to temporal variation, though its source is different from the speaking rate. The results are considered in terms of two possible accounts: one that takes both the boundary-induced and the speaking rate-induced effects as listeners' adjustments to low-level temporal variation, and the other that separates them by taking the boundary-induced effects to arise with computation of higher-level prosodic structure, given that the source of the localized slowing down effect is a prosodic boundary.
In two experiments we examine how listeners make reference to prosodic phrasing in their perception of temporally cued segmental contrasts. We test how the prosodic-structurally conditioned modulation of segmental cues (in domain-initial strengthening) translates into speech perception. We adopt the test case of stop contrasts in Seoul Korean (aspirated versus fortis), which are cued by vowel duration and voice onset time (VOT). The phrasing manipulation was carried out at the level of the Accentual Phrase (AP), a small phrase that is marked by intonational features. The AP was chosen because it was possible to create two prosodic phrasing contexts (AP-initial versus AP-medial) by manipulating only f0 before the target segment with the duration of contextual segments unchanged, controlling for temporal context effects. In Experiment 1, listeners shift their perception of a VOT continuum based on phrasing, in line with the domain-initial strengthening pattern of post-stop vowel lengthening, where AP-initial post-fortis vowels are lengthened. Experiment 2 shows that vowel duration is used as a cue to the contrast and that perceptual categorization of vowel duration itself is also mediated by contextual phrasing information. Results thus suggest that prosodic phrasing, signaled by intonation only, mediates perception of the segmental contrast, with temporal context controlled. We discuss these findings in terms of their implications for the role of phrasing in segmental perception and in processing.
How do Dutch and Korean listeners use acoustic–phonetic information when learning words in an artificial language? Dutch has a voiceless ‘unaspirated’ stop, produced with shortened Voice Onset Time (VOT) in prosodic strengthening environments (e.g., in domain-initial position and under prominence), enhancing the feature {−spread glottis}; Korean has a voiceless ‘aspirated’ stop produced with lengthened VOT in similar environments, enhancing the feature {+spread glottis}. Given this cross-linguistic difference, two competing hypotheses were tested. The phonological-superiority hypothesis predicts that Dutch and Korean listeners should utilize shortened and lengthened VOTs, respectively, as cues in artificial-language segmentation. The phonetic-superiority hypothesis predicts that both groups should take advantage of the phonetic richness of longer VOTs (i.e., their enhanced auditory–perceptual robustness). Dutch and Korean listeners learned the words of an artificial language better when word-initial stops had longer VOTs than when they had shorter VOTs. It appears that language-specific phonological knowledge can be overridden by phonetic richness in processing an unfamiliar language. Listeners nonetheless performed better when the stimuli were based on the speech of their native languages, suggesting that the use of richer phonetic information was modulated by listeners' familiarity with the stimuli.
This study investigated the role of phrase-level prosodic boundary information in word segmentation in Korean with two word-spotting experiments. In experiment 1, it was found that intonational cues alone helped listeners with lexical segmentation. Listeners paid more attention to local intonational cues (…H#L…) across the prosodic boundary than the intonational information within a prosodic phrase. The results imply that intonation patterns with high frequency are used, though not exclusively, in lexical segmentation. In experiment 2, final lengthening was added to see how multiple prosodic cues influence lexical segmentation. The results showed that listeners did not necessarily benefit from the presence of both intonational and final lengthening cues: Their performance was improved only when intonational information contained infrequent tonal patterns for boundary marking, showing only partially cumulative effects of prosodic cues. When the intonational information was optimal (frequent) for boundary marking, however, poorer performance was observed with final lengthening. This is arguably because the phrase-initial segmental allophonic cues for the accentual phrase were not matched with the prosodic cues for the intonational phrase. It is proposed that the asymmetrical use of multiple cues was due to interaction between prosodic and segmental information that are computed in parallel in lexical segmentation.
Categorical perception experiments were performed on an English /b-p/ voice onset time (VOT) continuum with native (American English) and non-native (Korean) listeners to examine whether and how phonetic categorization is modulated by prosodic boundary and language experience. Results demonstrated perceptual shifting according to prosodic boundary strength: A longer VOT was required to identify a sound as /p/ after an intonational phrase than a word boundary, regardless of the listeners' language experience. This suggests that segmental perception is modulated by the listeners' computation of an abstract prosodic structure reflected in phonetic cues of phrase-final lengthening and domain-initial strengthening, which are common across languages.