This study investigates how second-language (L2) listeners from five first-language (L1) backgrounds—English, Dutch, Mandarin, Spanish, and Korean—perceive English lexical stress, focusing on their use of vowel quality, pitch, and duration cues. Participants completed a cue-weighting perception task (Tremblay et al., 2021) in which two acoustic dimensions were manipulated orthogonally while the third was neutralized. Data for Dutch listeners come from the original study. Predictions about cross-lin-<br/>guistic transfer were based on the functional weight of each cue in the L1. The following L1 effects were predicted: For vowel quality: English, Mandarin > Dutch > Spanish, Korean; for pitch: Mandarin > Korean > Dutch, Spanish > English; for duration: English, Mandarin> Dutch, Spanish > Korean. Bayesian mixed-effects models tested the effects of cues and L1 with L2 proficiency (Lemh€ofer & Broersma, 2012) as a covariate. The results aligned broadly with our predictions: for vowel quality, English-> Mandarin > Dutch > Korean > Spanish; for pitch: Mandarin > Korean, Dutch > Spanish > English; for duration: English, Mandarin, Dutch > Spanish > Korean. These findings support a cue-weighting typology shaped by L1-specific cue prominence, with implications for theories of transfer and perceptual learning in L2 acquisition.
This study investigated how coda voicing contrast in English would be phonetically encoded in the temporal vs. spectral dimension of the preceding vowel (in vowel duration vs. F1/F2) by Korean L2 speakers of English, and how their L2 phonetic encoding pattern would be compared to that of native English speakers. Crucially, these questions were explored by taking into account the phonetics-prosody interface, testing effects of prominence by comparing target segments in three focus conditions (phonological focus, lexical focus, and no focus) that stem from information structure. Results showed that Korean speakers utilized the temporal dimension (vowel duration) to encode coda voicing contrast, but failed to use the spectral dimension (F1/F2), reflecting their native language experience—i.e., with a more sparsely populated vowel space in Korean, they are less sensitive to small changes in the spectral dimension, and hence fine-grained spectral cues in English are not readily accessible. Results also showed that along the temporal dimension, both the L1 and L2 speakers hyperarticulated coda voicing contrast under prominence (when phonologically or lexically focused), but hypoarticulated it in the non-prominent condition. This indicates that low-level phonetic realization and high-order information structure interact in a communicatively efficient way, regardless of the speakers' native language background. The Korean speakers, however, used the temporal phonetic space differently from the way the native speakers did, especially showing less reduction in the no focus condition. This was also attributable to their native language experience—i.e., the Korean speakers' use of temporal dimension is constrained in a way that is not detrimental to the preservation of coda voicing contrast, given that they fail to add additional cues along the spectral dimension. The results imply that the L2 phonetic system can be more fully illuminated through an investigation of the phonetics-prosody interface that is further modulated by higher-order linguistic structure such as information structure as well as by the L2 speakers' native language experience.
We investigated how listeners of two unrelated languages, Dutch and Korean, process phonotactically legitimate and illegitimate sounds spoken in Dutch and American English. To Dutch listeners, unreleased word-final stops are phonotactically illegal because word-final stops in Dutch are generally released in isolation, but to Korean listeners, released final stops are illegal because word-final stops are never released in Korean. Two phoneme monitoring experiments showed a phonotactic effect: Dutch listeners detected released stops more rapidly than unreleased stops whereas the reverse was true for Korean listeners. Korean listeners with English stimuli detected released stops more accurately than unreleased stops, however, suggesting that acoustic-phonetic cues associated with released stops improve detection accuracy. We propose that in non-native speech perception, phonotactic legitimacy in the native language speeds up phoneme recognition, the richness of acousticphonetic cues improves listening accuracy, and familiarity with the non-native language modulates the relative influence of these two factors.
Prosodic focus marking in Seoul Korean is known to be achieved primarily through prosodic phrasing, different from the use of prosody for this purpose in many other languages. This study investigates how children use prosodic phrasing for focus-marking purposes in Seoul Korean, compared to adults.
This article presents articulatory–kinematic data on preboundary lengthening (Intonational Phrase-final lengthening) from the productions of ten native speakers of American English—a relatively rare class of phonetic data compared with the more widely available acoustic data. The dataset includes three trisyllabic nonce words (bábaba, babába, bababá), each designed to manipulate the location of lexical stress. These were produced under prosodic conditions that varied in boundary position and focus-induced phrasal prominence, enabling analysis of how preboundary lengthening is distributed across words with different lexical stress locations and how it interacts with prosodic prominence. Articulatory data were collected using electromagnetic articulography (EMA, Carstens AG200), providing kinematic measurements such as movement duration, peak velocity, and displacement of articulatory gestures. The accompanying files allow examination of individual speaker variation in these measures as modulated by prosodic structure, including boundary and prominence effects. While theoretical findings have been reported in a previous study, the full dataset, including detailed descriptions of individual speaker patterns, is made available here. By making these less commonly available articulatory data publicly available, we aim to promote broad reuse and support further research in prosody, articulatory phonetics, and speech production.
This study investigates whether listeners’ cue weighting predicts their real‐time use of asynchronous acoustic information in spoken word recognition at both group and individual levels. By focusing on the time course of cue integration, we seek to distinguish between two theoretical views: the associated view (cue weighting is linked to cue integration strategy) and the independent view (no such relationship). The current study examines Seoul Korean listeners’ ( n = 62) weighting of voice onset time (VOT, available earlier in time) and onset fundamental frequency of the following vowel (F0, available later in time) when perceiving Korean stop contrasts (Experiment 1: cue‐weighting perception task) and the timing of VOT integration when recognizing Korean words that begin with a stop (Experiment 2: visual‐world eye‐tracking task). The group‐level results reveal that the timing of the early cue (VOT) integration is delayed when the later cue (F0) serves as the primary cue to process the stop contrast, supporting a relationship between cue weighting and the timing of cue integration (the associated view). At the individual level, listeners with greater reliance on F0 than VOT exhibited a further delayed integration of VOT. These findings suggest that the real‐time processing of asynchronously occurring acoustic cues for lexical activation is modulated by the weight that listeners assign to those cues, providing evidence for the associated view of cue integration. This study offers insights into the mechanisms of cue integration and spoken word recognition, and they shed light on variability in cue integration strategies among listeners.
The Korean three-way stop contrast (lenis, aspirated, fortis) is currently undergoing a sound change, such that the primary cue distinguishing lenis and aspirated stops is shifting from voice onset time (VOT) to F0. Despite recent discussions of this shift, research on voice quality, traditionally considered an additional cue signaling the contrast, remains sparse. This study investigated the extent to which the associated voice quality [as reflected in the acoustic measurements of H1*–H2*, H1*–A1*, and cepstral peak prominence (CPP)] contributes to the three-way stop contrast, and how the realization is conditioned by prominence-vs. boundary-induced prosodic strengthening amid the ongoing sound change. Results for 12 native Korean speakers indicate that there was a substantial distinction in voice quality among the three stop categories with the breathiness of the vowel being the greatest after the lenis, intermediate after the aspirated, and least after the fortis stops, indicating the role of voice quality in the maintenance of the three-way stop contrast. Furthermore, prosodic strengthening has different effects on the contrast and contributes to the enhancement of the phonological contrast contingent on whether it is induced by prominence or boundary.
This study explores processing characteristics of a glottal stop in Maltese which occurs both as a phoneme and as an epenthetic stop for vowel-initial words. Experiment 1 shows that its hyperarticulation is not necessarily mapped onto an underlying form, although listeners may interpret it as underlying at a later processing stage. Experiment 2 shows that listeners’ experience with a particular speaker’s use of a glottal stop exclusively as a phoneme does not modulate competition patterns accordingly. Not only are vowel-initial words activated by [ʔ]-initial forms, but /ʔ/-initial words are also activated by vowel-initial forms, suggesting that lexical access is not constrained by an initial acoustic mismatch that involves a glottal stop. Experiment 3 reveals that the observed pattern is not generalizable to an oral stop /t/. We propose that glottal stops have a special status in lexical processing: it is prosodic in nature to be licensed by the prosodic structure.
This study investigates whether listeners’ cue weighting predicts their real-time processing of asynchronous acoustic information as the speech signal unfolds over time. It does so by testing the time course of acoustic cue integration in the processing of Seoul Korean stop contrasts by native listeners. The current study further tests whether listeners’ cue weighting is associated with cue integration at the individual level. Seoul Korean listeners’ (n = 62) perceptual weightings of voice onset time (VOT, available earlier in time) and onset fundamental frequency of the following vowel (F0, available later in time) to perceive Korean stop contrasts were measured with a speech perception task (Experiment 1), and the timing of VOT integration in lexical access was examined with a visual-world eye-tracking task (Experiment 2). The group results revealed that the timing of VOT integration is predicted by listeners’ reliance on F0, with delayed integration of VOT in target-competitor pairs where F0 is a primary cue to process the stop contrast. At the individual level, listeners who relied more on F0 than on VOT showed later integration of VOT, further elucidating the relationship between cue weighting and the time course of cue integration. These results suggest that listeners’ real-time processing of asynchronous acoustic information in lexical activation is modulated by the informativeness of perceptual cues. As such, this study provides a nuanced perspective for a better understanding of listeners’ moment-by-moment processing of acoustic information in spoken word recognition.
The present study examined how listeners of Seoul Korean would recover deleted phonemes in consonant cluster simplification. In a phoneme monitoring experiment, listeners had to monitor for C2 (/k/ or /p/) in C1C2C3 when C2 was deleted (C1 was preserved) or preserved (C1 was deleted). The target consonant (C2) was either /k/ or /p/ (e.g., ilk-t?lato vs. palp-t?lato), and there were two listener groups, one group tested in 2002 and the other in 2009. Some points have emerged from the results. First, listeners were able to detect deleted phonemes as accurately and rapidly as preserved phonemes, showing that the physical presence of the acoustic information did not improve the listeners' performance. This suggests that listeners must have relied on language-specific phonological knowledge about the consonant cluster simplification, rather than relying on the low-level acoustic-phonetic information. Second, listener groups (participants in 2002 vs. 2009), differed in processing /p/ versus /k/: listeners in 2009 failed to detect /p/ more frequently than those in 2002, suggesting that the way the consonant cluster sequence is produced and perceived has changed over time. This result was interpreted as coming from statistical patterns of speech production in contemporary Seoul Korean as reported in a recent study by Cho & Kim (2009): /p/ is deleted far more often than /p/ is preserved, which is likely reflected in the way listeners process simplified variants. Finally, listeners processed /k/ more efficiently than /p/, especially when the target was physically present (in C-preserved condition), indicating that listeners benefited more from the presence of /k/ than of /p/. This was interpreted as supporting the view that velars are perceptually more robust than labials, which constrains shaping phonological patterns of the language. These results were then discussed in terms of their implications for theories of spoken word recognition.
No abstract is provided for this article.
This data article provides acoustic data for individual speakers' production of coda voicing contrast between stops in English, which are based on laboratory speech recorded by twelve native speakers of American English and twenty-four Korean learners of English. There were four pairs of English monosyllabic target words with voicing contrast in the coda position (bet-bed, pet-ped, bat-bad, pat-pad). The words were produced in carrier sentences in which they were placed in two different prosodic boundary conditions (Intonational Phrase initial and Intonation Phrase medial), two pitch accent conditions (nuclear-pitch accented and unaccented), and three focus conditions (lexical focus, phonological focus and no focus). The raw acoustic measurement values that are included in a CSV-formated file are F0, F1, F2 and duration of each vowel preceding a coda consonant; and Voice Onset Time of word-initial stops. This article also provides figures that exemplify individual speaker variation of vowel duration, F0, F1 and F2 as a function of focus conditions. The data can thus be potentially reused to observe individual variations in phonetic encoding of coda voicing contrast as a function of the aforementioned prosodically-conditioned factors (i.e., prosodic boundary, pitch accent, focus) in native vs. non-native English. Some theoretical aspects of the data are discussed in the full-length article entitled "Phonetic encoding of coda voicing contrast under different focus conditions in L1 vs. L2 English" [1].