176 publications from this institution
This study investigates how the fine-grained phonetic realization of tonal cues impacts speech segmentation when the cues signal the same word boundary in the native and unfamiliar languages but do so differently. Korean listeners use the phrase-final high (H) tone and the phrase-initial low (L) tone to segment speech into words (Kim, Broersma, & Cho, 2012; Kim & Cho, 2009), but it is unclear how the alignment of the phrase-final H tone and the scaling of the phrase-initial L tone modulate their speech segmentation. Korean listeners completed three artificial-language (AL) tasks (within-subject): (a) one AL without tonal cues; (b) one AL with later-aligned phrase-final H cues (non-Korean-like); and (c) one AL with earlier-aligned phrase-final H cues (Korean-like). Three groups of Korean listeners heard (b) and (c) in three phrase-initial L scaling conditions (between-subject): high (non-Korean-like), mid (non-Korean-like), or low (Korean-like). Korean listeners’ segmentation improved as the L tone was lowered, and (b) enhanced segmentation more than (c) in the high- and mid-scaling conditions. We propose that Korean listeners tune in to low-level cues (the greater H-to-L slope in [b]) that conform to the Korean intonational grammar when the phrase-initial L tone is not canonical phonologically.
This study explores the relationship between prosodic strengthening and linguistic contrasts in English by examining temporal realization of nasals (N-duration) in CVN# and #NVC, and their coarticulatory influence on vowels (V-nasalization). Results show that different sources of prosodic strengthening bring about different types of linguistic contrasts. Prominence enhances the consonant׳s [nasality] as reflected in an elongation of N-duration, but it enhances the vowel׳s [orality] (rather than [nasality]) showing coarticulatory resistance to the nasal influence even when the nasal is phonologically focused (e.g., mob-bob; bomb-bob). Boundary strength induces different types of enhancement patterns as a function of prosodic position (initial vs. final). In the domain-initial position, boundary strength reduces the consonant׳s [nasality] as evident in a shortening of N-duration and a reduction of V-nasalization, thus enhancing CV contrast. The opposite is true with the domain-final nasal in which N-duration is lengthened accompanied by greater V-nasalization, showing coarticulatory vulnerability. The systematic coarticulatory variation as a function of prosodic factors indicates that V-nasalization as a coarticulatory process is indeed under speaker control, fine-tuned in a linguistically significant way. In dynamical terms, these results may be seen as coming from differential intergestural coupling relationships that may underlie the difference in V-nasalization in CVN# vs. #NVC. It is proposed that the timing initially determined by such coupling relationships must be fine-tuned by prosodic strengthening in a way that reflects the relationship between dynamical underpinnings of speech timing and linguistic contrasts.
No abstract is provided for this article.
This study tests whether potential differences in the perceptual robustness of speech sounds influence continuous-speech processes. Two phoneme-monitoring experiments examined place assimilation in Korean. In Experiment 1, Koreans monitored for targets which were either labials (/p,m/) or alveolars (/t,n/), and which were either unassimilated or assimilated to a following /k/ in two-word utterances. Listeners detected unaltered (unassimilated) labials faster and more accurately than assimilated labials; there was no such advantage for unaltered alveolars. In Experiment 2, labial–velar differences were tested using conditions in which /k/ and /p/ were illegally assimilated to a following /t/. Unassimilated sounds were detected faster than illegally assimilated sounds, but this difference tended to be larger for /k/ than for /p/. These place-dependent asymmetries suggest that differences in the perceptual robustness of segments play a role in shaping phonological patterns.
The data reported in this article contain eleven (6 female and 5 male) individual speaker's speech production patterns for the word-initial voiced and voiceless stops (/p,t/ and /b,d/) in American English. The production patterns are documented in the acoustic parameter: the Integrated Voicing Index (IVI) obtained from Voice Onset Time (VOT) and voicing duration in the stop closure (Voicing-in-Closure), in various prosodic contexts: lexically-stressed vs. unstressed; accented (focused) vs. unaccented (unfocused); phrase-initial vs. phrase-medial. The data also contain a CVS file with each speaker׳s mean values of the IVI, VOT and Voicing-in-Closure for each prosodic condition for the voiced and voiceless stops, along with the information about the speaker gender. For further discussion of the data, please refer to the full length article entitled "Prosodic-structural modulation of stop voicing contrast along the VOT continuum in trochaic and iambic words in American English" (Kim et al., 2018).
This study investigated the role of phrase-level prosodic boundary information in word segmentation in Korean with two word-spotting experiments. In experiment 1, it was found that intonational cues alone helped listeners with lexical segmentation. Listeners paid more attention to local intonational cues (…H#L…) across the prosodic boundary than the intonational information within a prosodic phrase. The results imply that intonation patterns with high frequency are used, though not exclusively, in lexical segmentation. In experiment 2, final lengthening was added to see how multiple prosodic cues influence lexical segmentation. The results showed that listeners did not necessarily benefit from the presence of both intonational and final lengthening cues: Their performance was improved only when intonational information contained infrequent tonal patterns for boundary marking, showing only partially cumulative effects of prosodic cues. When the intonational information was optimal (frequent) for boundary marking, however, poorer performance was observed with final lengthening. This is arguably because the phrase-initial segmental allophonic cues for the accentual phrase were not matched with the prosodic cues for the intonational phrase. It is proposed that the asymmetrical use of multiple cues was due to interaction between prosodic and segmental information that are computed in parallel in lexical segmentation.
Categorical perception experiments were performed on an English /b-p/ voice onset time (VOT) continuum with native (American English) and non-native (Korean) listeners to examine whether and how phonetic categorization is modulated by prosodic boundary and language experience. Results demonstrated perceptual shifting according to prosodic boundary strength: A longer VOT was required to identify a sound as /p/ after an intonational phrase than a word boundary, regardless of the listeners' language experience. This suggests that segmental perception is modulated by the listeners' computation of an abstract prosodic structure reflected in phonetic cues of phrase-final lengthening and domain-initial strengthening, which are common across languages.
This study investigates how prosodic strengthening is kinematically manifested in V-to-V lingual movement in English CV#CV context (where # is a prosodic boundary). Results showed that both boundary and accent gave rise to a kind of prosodic strengthening (showing spatial and temporal expansion), but exact kinematic patterns of prosodic strengthening were different as a function of the type of gesture (tongue lowering versus raising) associated with different vowels (/i/to-/ alpha/ vs. /alpha/-to-/i/) and the source of prosodic strengthening (boundary versus accentuation). This implies that speakers must know about prosodic structure and differentiate the two sources of prosodic strengthening in a systematic fine-grained fashion. From a theoretical point of view regarding a mass-spring gestural model, results suggested that kinematic patterns of prosodic strengthening could not be fully accounted for by any particular dynamical parameter, presenting a complex nature of prosodic strengthening. The results also implied that the theory of the pi-gesture (the prosodic boundary gesture) under the rubric of the mass-spring gestural model needs to be refined in terms of how the theory defines the exact scope of the pi-gesture's influence in the temporal dimension and how it differentiates boundary-induced articulation from an accent-induced one.
This study examines how young speakers of Seoul Korean produce tri-consonantal clusters /1kt/ and /1pt/ as in palk-ta ('to be bright') and palp-ta ('to step on'). Production data were collected from 20 speakers of Seoul Korean. The results of narrow transcription of the data showed that simplification is not obligatory as some speakers often preserve all three consonants. When simplified, there was a clear asymmetry between /1kt/ and /1pt/. Speakers showed no clear preference for either C1 preservation (C1=/1/) or C2 preservation (C2=/k/ in /1kt/ and /p/ in /1pt/) in production of /1kt/, but in production of /1pt/, strong preference was found for C1-preserved to C2-preserved variant. When compared with production data in Cho (1999), simplification patterns appear to have changed over the past 10 years, in a direction to preserve the first member of the cluster (/1/) more often, especially with /1kt/. There was no substantial between-item variation, indicating that simplification patterns are not lexically specified. Finally, the results suggest that the process of tri-consonantal simplification has not been fully phonologized in the grammar of the language as evident in substantial inter- and intra-speaker variation.
o/√ merger; and (4) at least for the urban Cheju speakers, the merger is best accounted for by the merger-by-transfer model, a unidirectional change in which one phonemic category becomes another (cf. Labov, 1994). Further, when our data are compared with other acoustic data available (including studies of the standard Korean in the 1960s and 1990s), it suggests that the directionality of the diachronic sound change is guided by both auditorily and articulatorily based principles such as contrast maximization and effort minimization principles.
Prosodic influences on phonetic realizations of four Dutch consonants (/t d s z/) were examined. Sentences were constructed containing these consonants in word-initial position; the factors lexical stress, phrasal accent and prosodic boundary were manipulated between sentences. Eleven Dutch speakers read these sentences aloud. The patterns found in acoustic measurements of these utterances (e.g., voice onset time (VOT), consonant duration, voicing during closure, spectral center of gravity, burst energy) indicate that the low-level phonetic implementation of all four consonants is modulated by prosodic structure. Boundary effects on domain-initial segments were observed in stressed and unstressed syllables, extending previous findings which have been on stressed syllables alone. Three aspects of the data are highlighted. First, shorter VOTs were found for /t/ in prosodically stronger locations (stressed, accented and domain-initial), as opposed to longer VOTs in these positions in English. This suggests that prosodically driven phonetic realization is bounded by language-specific constraints on how phonetic features are specified with phonetic content: Shortened VOT in Dutch reflects enhancement of the phonetic feature {−spread glottis}, while lengthened VOT in English reflects enhancement of {+spread glottis}. Prosodic strengthening therefore appears to operate primarily at the phonetic level, such that prosodically driven enhancement of phonological contrast is determined by phonetic implementation of these (language-specific) phonetic features. Second, an accent effect was observed in stressed and unstressed syllables, and was independent of prosodic boundary size. The domain of accentuation in Dutch is thus larger than the foot. Third, within a prosodic category consisting of those utterances with a boundary tone but no pause, tokens with syntactically defined Phonological Phrase boundaries could be differentiated from the other tokens. This syntactic influence on prosodic phrasing implies the existence of an intermediate-level phrase in the prosodic hierarchy of Dutch.
This study investigates how second-language (L2) listeners from five first-language (L1) backgrounds—English, Dutch, Mandarin, Spanish, and Korean—perceive English lexical stress, focusing on their use of vowel quality, pitch, and duration cues. Participants completed a cue-weighting perception task (Tremblay et al., 2021) in which two acoustic dimensions were manipulated orthogonally while the third was neutralized. Data for Dutch listeners come from the original study. Predictions about cross-lin-<br/>guistic transfer were based on the functional weight of each cue in the L1. The following L1 effects were predicted: For vowel quality: English, Mandarin > Dutch > Spanish, Korean; for pitch: Mandarin > Korean > Dutch, Spanish > English; for duration: English, Mandarin> Dutch, Spanish > Korean. Bayesian mixed-effects models tested the effects of cues and L1 with L2 proficiency (Lemh€ofer & Broersma, 2012) as a covariate. The results aligned broadly with our predictions: for vowel quality, English-> Mandarin > Dutch > Korean > Spanish; for pitch: Mandarin > Korean, Dutch > Spanish > English; for duration: English, Mandarin, Dutch > Spanish > Korean. These findings support a cue-weighting typology shaped by L1-specific cue prominence, with implications for theories of transfer and perceptual learning in L2 acquisition.
This study investigated how three different kinds of hyper-articulation, one communicatively driven (in clear speech), and two prosodically driven (with boundary and prominence/focus), are acoustic-phonetically realized in Korean. Several important points emerged from the results obtained from an acoustic study with eight speakers of Seoul Korean. First, clear speech gave rise to global modification of the temporal and prosodic structures over the course of the utterance, showing slowing down of the utterance and more prosodic phrases. Second, although the three kinds of hyper-articulation were similar in some aspects, they also differed in many aspects, suggesting that different sources of hyper-articulation are encoded separately in speech production. Third, the three kinds of hyper-articulation interacted with each other; the communicatively driven hyper-articulation was prosodically modulated, such that in a clear speech mode not every segment was hyper-articulated to the same degree, but prosodically important landmarks (e.g., in IP-initial and/or focused conditions) were weighted more. Finally, Korean, a language without lexical stress and pitch accent, showed different hyper-articulation patterns compared to other, Indo-European languages such as English—i.e., it showed more robust domain-initial strengthening effects (extended beyond the first initial segment), focus effects (extended to V1 and V2 of the entire bisyllabic test word) and no use of global F0 features in clear speech. Overall, the present study suggests that the communicatively driven and the prosodically driven hyper-articulations are intricately intertwined in ways that reflect not only interactions of principles of gestural economy and contrast enhancement, but also language-specific prosodic systems, which further modulate how the three kinds of hyper-articulations are phonetically expressed.
The present study investigates effects of Boundary and Prominence (focus) on the /a/-to-/i/ tongue movement in Korean in two contexts: V#V and V#/m/V. Results show that the tongue movement at an IP boundary is larger, longer, and faster. Prominence effects show a relatively weaker but comparable pattern to the boundary effect, showing a larger, longer, and faster movement. The observed boundary-induced strengthening pattern in Korean is clearly different from that in English, implying that Korean, a language without constraints from the lexical stress system, has more freedom to strengthen articulation at prosodic junctures, creating strengthening patterns which are often encountered with prominence marking in English. Results also reveal that the presence of a consonant influences transboundary vocalic movement, and that the consonantal influence is further modulated by boundary strength. These results taken together are further discussed in terms of language-specificity of prosodic strengthening and its implications for the pi-gesture model.
Voice onset time (VOT) is known to vary with place of articulation. For any given place of articulation there are differences from one language to another. Using data from multiple speakers of 18 languages, all of which were recorded and analyzed in the same way, we show that most, but not all, of the within language place of articulation variation can be described by universally applicable phonetic rules (although the physiological bases for these rules are not entirely clear). The between language variation is also largely (but not entirely) predictable by assuming that languages choose one of the three possibilities for the degree of aspiration of voiceless stops. Some languages, however, have VOTs that are markedly different from the generally observed values. The phonetic output of a grammar has to contain language specific components to account for these results.
Recent studies have indicated that vowels in prosodically strong positions (e.g., in stressed syllables and at edges of prosodic boundaries) are not only strongly articulated, but also resistant to coarticulation with neighboring vowels. This paper further examines vowel-to-vowel coarticulation in English by analyzing extensive articulatory data from six American English speakers, using the Electromagnetic Articulograph (EMA). It is hypothesized that vowels in prosodically strong positions are more resistant to coarticulation with their neighbors, and at the same time encroach more on their neighbors. To test this, sentences were designed so that they included /V1♯bV2/ where V1 and V2 were manipulated, resulting in /a-a/, /i-i/ (control condition) and /a-i/, /i-a/ (test condition). Vowels also varied in sentence stress (accented versus unaccented) and in the intervening boundaries (♯ = Word, ip, IP). The vertical and horizontal positions of three tongue points and jaw are examined at five different points (onset, first quarter, middle, three quarters and end) in the vowel, to assess how much of the vowel articulation is anticipated or carried over at different points of the vowel. This shows variation in degree of V-to-V coarticulation under various prosodic conditions. [Work supported by NSF doctoral research grant.]
This study investigated how acoustic characteristics (i.e., duration, F1, F2) of English high front vowels /i, ɪ/ are modulated by boundary- and prominence-induced strengthening in native vs. non-native (Korean) speech production. The study also examined how the durational difference in vowels due to the voicing of a following consonant (i.e., voiced vs. voiceless) is modified by prosodic strengthening in two different (native vs. non-native) speaker groups. Five native speakers of Canadian English and eight Korean learners of English (intermediate-advanced level) produced 8 minimal pairs with the CVC sequence (e.g., 'beat'-'bit') in varying prosodic contexts. Native speakers distinguished the two vowels in terms of duration, F1, and F2, whereas non-native speakers only showed durational differences. The two groups were similar in that they maximally distinguished the two vowels when the vowels were accented (F2, duration), while neither group showed boundary-induced strengthening in any of the three measurements. The durational differences due to the voicing of the following consonant were also maximized when accented. The results are discussed further in terms of phonetics-prosody interface in L2 production.