This study explores the relationship between prosodic strengthening and linguistic contrasts in English by examining temporal realization of nasals (N-duration) in CVN# and #NVC, and their coarticulatory influence on vowels (V-nasalization). Results show that different sources of prosodic strengthening bring about different types of linguistic contrasts. Prominence enhances the consonant׳s [nasality] as reflected in an elongation of N-duration, but it enhances the vowel׳s [orality] (rather than [nasality]) showing coarticulatory resistance to the nasal influence even when the nasal is phonologically focused (e.g., mob-bob; bomb-bob). Boundary strength induces different types of enhancement patterns as a function of prosodic position (initial vs. final). In the domain-initial position, boundary strength reduces the consonant׳s [nasality] as evident in a shortening of N-duration and a reduction of V-nasalization, thus enhancing CV contrast. The opposite is true with the domain-final nasal in which N-duration is lengthened accompanied by greater V-nasalization, showing coarticulatory vulnerability. The systematic coarticulatory variation as a function of prosodic factors indicates that V-nasalization as a coarticulatory process is indeed under speaker control, fine-tuned in a linguistically significant way. In dynamical terms, these results may be seen as coming from differential intergestural coupling relationships that may underlie the difference in V-nasalization in CVN# vs. #NVC. It is proposed that the timing initially determined by such coupling relationships must be fine-tuned by prosodic strengthening in a way that reflects the relationship between dynamical underpinnings of speech timing and linguistic contrasts.
This study tests whether potential differences in the perceptual robustness of speech sounds influence continuous-speech processes. Two phoneme-monitoring experiments examined place assimilation in Korean. In Experiment 1, Koreans monitored for targets which were either labials (/p,m/) or alveolars (/t,n/), and which were either unassimilated or assimilated to a following /k/ in two-word utterances. Listeners detected unaltered (unassimilated) labials faster and more accurately than assimilated labials; there was no such advantage for unaltered alveolars. In Experiment 2, labial–velar differences were tested using conditions in which /k/ and /p/ were illegally assimilated to a following /t/. Unassimilated sounds were detected faster than illegally assimilated sounds, but this difference tended to be larger for /k/ than for /p/. These place-dependent asymmetries suggest that differences in the perceptual robustness of segments play a role in shaping phonological patterns.
No abstract is provided for this article.
The data reported in this article contain eleven (6 female and 5 male) individual speaker's speech production patterns for the word-initial voiced and voiceless stops (/p,t/ and /b,d/) in American English. The production patterns are documented in the acoustic parameter: the Integrated Voicing Index (IVI) obtained from Voice Onset Time (VOT) and voicing duration in the stop closure (Voicing-in-Closure), in various prosodic contexts: lexically-stressed vs. unstressed; accented (focused) vs. unaccented (unfocused); phrase-initial vs. phrase-medial. The data also contain a CVS file with each speaker׳s mean values of the IVI, VOT and Voicing-in-Closure for each prosodic condition for the voiced and voiceless stops, along with the information about the speaker gender. For further discussion of the data, please refer to the full length article entitled "Prosodic-structural modulation of stop voicing contrast along the VOT continuum in trochaic and iambic words in American English" (Kim et al., 2018).
This study investigates whether listeners’ cue weighting predicts their real-time processing of asynchronous acoustic information as the speech signal unfolds over time. It does so by testing the time course of acoustic cue integration in the processing of Seoul Korean stop contrasts by native listeners. The current study further tests whether listeners’ cue weighting is associated with cue integration at the individual level. Seoul Korean listeners’ (n = 62) perceptual weightings of voice onset time (VOT, available earlier in time) and onset fundamental frequency of the following vowel (F0, available later in time) to perceive Korean stop contrasts were measured with a speech perception task (Experiment 1), and the timing of VOT integration in lexical access was examined with a visual-world eye-tracking task (Experiment 2). The group results revealed that the timing of VOT integration is predicted by listeners’ reliance on F0, with delayed integration of VOT in target-competitor pairs where F0 is a primary cue to process the stop contrast. At the individual level, listeners who relied more on F0 than on VOT showed later integration of VOT, further elucidating the relationship between cue weighting and the time course of cue integration. These results suggest that listeners’ real-time processing of asynchronous acoustic information in lexical activation is modulated by the informativeness of perceptual cues. As such, this study provides a nuanced perspective for a better understanding of listeners’ moment-by-moment processing of acoustic information in spoken word recognition.
This study explores processing characteristics of a glottal stop in Maltese which occurs both as a phoneme and as an epenthetic stop for vowel-initial words. Experiment 1 shows that its hyperarticulation is not necessarily mapped onto an underlying form, although listeners may interpret it as underlying at a later processing stage. Experiment 2 shows that listeners’ experience with a particular speaker’s use of a glottal stop exclusively as a phoneme does not modulate competition patterns accordingly. Not only are vowel-initial words activated by [ʔ]-initial forms, but /ʔ/-initial words are also activated by vowel-initial forms, suggesting that lexical access is not constrained by an initial acoustic mismatch that involves a glottal stop. Experiment 3 reveals that the observed pattern is not generalizable to an oral stop /t/. We propose that glottal stops have a special status in lexical processing: it is prosodic in nature to be licensed by the prosodic structure.
No abstract is provided for this article.
The present study investigates effects of Boundary and Prominence (focus) on the /a/-to-/i/ tongue movement in Korean in two contexts: V#V and V#/m/V. Results show that the tongue movement at an IP boundary is larger, longer, and faster. Prominence effects show a relatively weaker but comparable pattern to the boundary effect, showing a larger, longer, and faster movement. The observed boundary-induced strengthening pattern in Korean is clearly different from that in English, implying that Korean, a language without constraints from the lexical stress system, has more freedom to strengthen articulation at prosodic junctures, creating strengthening patterns which are often encountered with prominence marking in English. Results also reveal that the presence of a consonant influences transboundary vocalic movement, and that the consonantal influence is further modulated by boundary strength. These results taken together are further discussed in terms of language-specificity of prosodic strengthening and its implications for the pi-gesture model.
The present study examined how listeners of Seoul Korean would recover deleted phonemes in consonant cluster simplification. In a phoneme monitoring experiment, listeners had to monitor for C2 (/k/ or /p/) in C1C2C3 when C2 was deleted (C1 was preserved) or preserved (C1 was deleted). The target consonant (C2) was either /k/ or /p/ (e.g., ilk-t?lato vs. palp-t?lato), and there were two listener groups, one group tested in 2002 and the other in 2009. Some points have emerged from the results. First, listeners were able to detect deleted phonemes as accurately and rapidly as preserved phonemes, showing that the physical presence of the acoustic information did not improve the listeners' performance. This suggests that listeners must have relied on language-specific phonological knowledge about the consonant cluster simplification, rather than relying on the low-level acoustic-phonetic information. Second, listener groups (participants in 2002 vs. 2009), differed in processing /p/ versus /k/: listeners in 2009 failed to detect /p/ more frequently than those in 2002, suggesting that the way the consonant cluster sequence is produced and perceived has changed over time. This result was interpreted as coming from statistical patterns of speech production in contemporary Seoul Korean as reported in a recent study by Cho & Kim (2009): /p/ is deleted far more often than /p/ is preserved, which is likely reflected in the way listeners process simplified variants. Finally, listeners processed /k/ more efficiently than /p/, especially when the target was physically present (in C-preserved condition), indicating that listeners benefited more from the presence of /k/ than of /p/. This was interpreted as supporting the view that velars are perceptually more robust than labials, which constrains shaping phonological patterns of the language. These results were then discussed in terms of their implications for theories of spoken word recognition.
The Korean three-way stop contrast (lenis, aspirated, fortis) is currently undergoing a sound change, such that the primary cue distinguishing lenis and aspirated stops is shifting from voice onset time (VOT) to F0. Despite recent discussions of this shift, research on voice quality, traditionally considered an additional cue signaling the contrast, remains sparse. This study investigated the extent to which the associated voice quality [as reflected in the acoustic measurements of H1*–H2*, H1*–A1*, and cepstral peak prominence (CPP)] contributes to the three-way stop contrast, and how the realization is conditioned by prominence-vs. boundary-induced prosodic strengthening amid the ongoing sound change. Results for 12 native Korean speakers indicate that there was a substantial distinction in voice quality among the three stop categories with the breathiness of the vowel being the greatest after the lenis, intermediate after the aspirated, and least after the fortis stops, indicating the role of voice quality in the maintenance of the three-way stop contrast. Furthermore, prosodic strengthening has different effects on the contrast and contributes to the enhancement of the phonological contrast contingent on whether it is induced by prominence or boundary.
This paper discusses some current issues regarding how prosodic structure is manifested in fine-grained phonetic details, how prosodically-conditioned articulatory variation is explained in terms of speech dynamics, and how such phonetic manifestation of prosodic structure may be exploited in spoken word recognition. Prosodic structure is phonetically manifested in prosodically important landmark locations such as prosodic domain-final position, domain-initial position and stressed/accented syllables. It will be discussed how each of the prosodic landmarks engenders particular phonetic patterns, how articulatory variation in such locations are dynamically accounted for, and how prosodically-driven fine-grained phonetic detail is exploited by listeners in speech comprehension.
Prosodic structure in English speech is signalled, in part, by stronger articulation of consonants at the onset of intonational phrases (IPs) than of consonants that are IP-medial. In two cross-modal priming experiments, American English listeners heard sentences and decided whether visual letter strings, presented during the sentences, were real words. We manipulated sentence type (either no IP boundary or an IP boundary in a critical two-word sequence), splicing (whether the onset of the sequence’s second word was spliced from another token of that sentence or cross-spliced from a matched sentence with or without an IP boundary), and relatedness (whether the visual target was the first word in the spoken sequence). There was a relatedness effect on target responses for sentences with no IP boundary only when they were cross-spliced, that is, where splicing provided evidence of domain-initial strengthening. Listeners thus use this evidence when segmenting continuous speech.