• Title/Summary/Keyword: Prosodic Information

Search Result 90, Processing Time 0.03 seconds

A cross-modal naming study: Effects of prosodic boundaries on the comprehension of relative clauses in Japanese

  • Kang, Soyoung;Kashiwagi, Akiko;Nakayama, Mineharu;Speer, Shari R.
    • Cross-Cultural Studies
    • /
    • v.24
    • /
    • pp.157-169
    • /
    • 2011
  • Compared to studies on prosodic effects on the comprehension of syntactic ambiguity in English, there are relatively few that investigated prosodic effects in East-Asian languages. This study examined the role of prosodic information in processing syntactically ambiguous sentences in Japanese. For syntactically ambiguous sentences containing relative clauses, this paper investigated whether prosodic information is immediately available during the process of these ambiguous sentences. Results from an auditory comprehension experiment with an on-line, cross-modal naming task seemingly suggest that contrary to the findings from the off-line study that examined the same constructions, prosodic information may not be immediately available to Japanese listeners. A possible account for failure to obtain effects of prosodic information is provided.

The Effect of Prosodic Position and Word Type on the Production of Korean Plosives

  • Jang, Mi
    • Phonetics and Speech Sciences
    • /
    • v.3 no.4
    • /
    • pp.71-81
    • /
    • 2011
  • This paper investigated how prosodic position and word type affect the phonetic structure of Korean coronal stops. Initial segments of prosodic domains were known to be more strongly articulated and longer relative to prosodic domain-medial segments. However, there are few studies examining whether the properties of prosodic domain-initial segments are affected by the information content of words (real vs. nonsense words). In addition, since the scope of domain-initial effect was known to be local to the initial consonant and the effects on the following vowel have been found to be limited, it is thus worth examining whether the prosodic domain-initial effect extends into the vowel after the initial consonant in a systematic way across different prosodic domains. The acoustic properties of Korean coronal stops (lenis /t/, aspirated /$t^h$/, and tense /t'/) were compared across Intonational Phrase, Phonological Phrase and Word-initial positions both in real and nonsense words. The durational intervals such as VOT and CV duration were cumulatively lengthened for /t/ and /$t^h$/ in the higher prosodic domain-initial positions. However, tense stop /t'/ did not show any variation as a function of prosodic position and word type. The domain-initial lenis stop showed significantly longer duration in nonsense words than in real words. But the prosodic domain-initial effect was not found in the properties of F0 and [H1-H2] of the vowel after initial stops. The present study provided evidence that speakers tend to enhance speech clarity when there is less contextual information as in prosodic domain-initial position and in nonsense words.

  • PDF

Automatic Detection of Korean Prosodic Boundaries U sing Acoustic and Grammatical Information (음성정보와 문법정보를 이용한 한국어 운율 경계의 자동 추정)

  • Kim, Sun-Hee;Jeon, Je-Hun;Hong, Hye-Jin;Chung, Min-Hwa
    • MALSORI
    • /
    • no.66
    • /
    • pp.117-130
    • /
    • 2008
  • This paper presents a method for automatically detecting Korean prosodic boundaries using both acoustic and grammatical information for the performance improvement of speech information processing systems. While most of previous works are solely based on grammatical information, our method utilizes not only grammatical information constructed by a Maximum-Entropy-based grammar model using 10 grammatical features, but also acoustical information constructed by a GMM-based acoustic model using 14 acoustic features. Given that Korean prosodic structure has two intonationally defined prosodic units, intonation phrase (IP) and accentual phrase (AP), experimental results show that the detection rate of AP boundaries is 82.6%, which is higher than the labeler agreement rate in hand transcribing, and that the detection rate of IP boundaries is 88.7%, which is slightly lower than the labeler agreement rate.

  • PDF

Prosodic Contour Generation for Korean Text-To-Speech System Using Artificial Neural Networks

  • Lim, Un-Cheon
    • The Journal of the Acoustical Society of Korea
    • /
    • v.28 no.2E
    • /
    • pp.43-50
    • /
    • 2009
  • To get more natural synthetic speech generated by a Korean TTS (Text-To-Speech) system, we have to know all the possible prosodic rules in Korean spoken language. We should find out these rules from linguistic, phonetic information or from real speech. In general, all of these rules should be integrated into a prosody-generation algorithm in a TTS system. But this algorithm cannot cover up all the possible prosodic rules in a language and it is not perfect, so the naturalness of synthesized speech cannot be as good as we expect. ANNs (Artificial Neural Networks) can be trained to learn the prosodic rules in Korean spoken language. To train and test ANNs, we need to prepare the prosodic patterns of all the phonemic segments in a prosodic corpus. A prosodic corpus will include meaningful sentences to represent all the possible prosodic rules. Sentences in the corpus were made by picking up a series of words from the list of PB (phonetically Balanced) isolated words. These sentences in the corpus were read by speakers, recorded, and collected as a speech database. By analyzing recorded real speech, we can extract prosodic pattern about each phoneme, and assign them as target and test patterns for ANNs. ANNs can learn the prosody from natural speech and generate prosodic patterns of the central phonemic segment in phoneme strings as output response of ANNs when phoneme strings of a sentence are given to ANNs as input stimuli.

An Experimental Study on Prosodic Patterns of Subjective Particles (주어자리조사의 운율패턴에 관한 실험음성학적 연구)

  • Seong Cheol-Jae;Song Yun-Gyeong
    • MALSORI
    • /
    • no.33_34
    • /
    • pp.23-42
    • /
    • 1997
  • This study has two main purposes. One is to explore the relationship between syntactic aspects and prosodic aspects in Standard Korean. The other is to provide speech synthesis with the information about such relationship. This study will focus on the prosodic behavior of subjective particles'-i/-ga', '-eun/-neun'. The prosodic features of subjective particles are described respectively. How do the elements such as the position of particles in a sentence, the sentence constituents, the length of the sentence and the rhythmic boundaries influence on the prosodic behavior are also investigated.

  • PDF

Pronunciation Variation Modeling for Korean Point-of-Interest Data Using Prosodic Information (운율 정보를 이용한 한국어 위치 정보 데이타의 발음 모델링)

  • Kim, Sun-He;Park, Jeon-Gue;Na, Min-Soo;Jeon, Je-Hun;Chung, Min-Wha
    • Journal of KIISE:Software and Applications
    • /
    • v.34 no.2
    • /
    • pp.104-111
    • /
    • 2007
  • This paper examines how the performance of an automatic speech recognizer was improved for Korean Point-of-Interest (POI) data by modeling pronunciation variation using structural prosodic information such as prosodic words and syllable length. First, multiple pronunciation variants are generated using prosodic words given that each POI word can be broken down into prosodic words. And the cross-prosodic-word variations were modeled considering the syllable length of word. A total of 81 experiments were conducted using 9 test sets (3 baseline and 6 proposed) on 9 trained sets (3 baseline, 6 proposed). The results show: (i) the performance was improved when the pronunciation lexica were generated using prosodic words; (ii) the best performance was achieved when the maximum number of variants was constrained to 3 based on the syllable length; and (iii) compared to the baseline word error rate (WER) of 4.63%, a maximum of 8.4% in WER reduction was achieved when both prosodic words and syllable length were considered.

The Continuous Speech Recognition with Prosodic Phrase Unit (운율구 단위의 연속음 인식)

  • 강지영;엄기완;김진영;최승호
    • The Journal of the Acoustical Society of Korea
    • /
    • v.18 no.8
    • /
    • pp.9-16
    • /
    • 1999
  • Generally, a speaker structures utterances very clearly by grouping words into phrases. This facilitates the listener's recovery of the meaning of the utterance and the speaker's intention. To this purpose, a speaker uses, among other things, prosodic information such as intonation pause, duration, intensity, etc. The research described here is concerned with the relationship between the strength of prosodic boundaries in spoken utterances as perceived by untrained listeners(Perceptual boundary strength, PBS)-In this paper, the preceptual boundary strength is used as the same meaning of the prosodic boundary strength-and prosodic information. We made a rule determinating the prosodic boundaries and verified the usefulness of the prosodic phrase as a recognition unit. Experiments results showed that the performance of speech recognition(SR) is improved in aspect of recognition rate and time compared with that using sentences as recognition unit. In the future we will suggest the methods that estimate more appropriate boundaries and study more various methods of prosody assisted SR.

  • PDF

Rich Transcription Generation Using Automatic Insertion of Punctuation Marks (자동 구두점 삽입을 이용한 Rich Transcription 생성)

  • Kim, Ji-Hwan
    • MALSORI
    • /
    • no.61
    • /
    • pp.87-100
    • /
    • 2007
  • A punctuation generation system which combines prosodic information with acoustic and language model information is presented. Experiments have been conducted first for the reference text transcriptions. In these experiments, prosodic information was shown to be more useful than language model information. When these information sources are combined, an F-measure of up to 0.7830 was obtained for adding punctuation to a reference transcription. This method of punctuation generation can also be applied to the 1-best output of a speech recogniser. The 1-best output is first time aligned. Based on the time alignment information, prosodic features are generated. As in the approach applied in the punctuation generation for reference transcriptions, the best sequence of punctuation marks for this 1-best output is found using the prosodic feature model and an language model trained on texts which contain punctuation marks.

  • PDF

Working memory and sensitivity to prosody in spoken language processing (언어 처리에서 운율 제약 활용과 작업 기억의 관계)

  • Lee, Eun-Kyung
    • Korean Journal of Cognitive Science
    • /
    • v.23 no.2
    • /
    • pp.249-267
    • /
    • 2012
  • Individual differences in working memory predict qualitative differences in language processing. High span comprehenders are better able to integrate probabilistic information such as plausibility and animacy, the use of which requires the computation of real world knowledge in syntactic parsing (e.g.,[1]). However, it is unclear whether similar individual differences exist in the use of informative prosodic cues. This study examines whether working memory modulates the use of prosodic boundary information in attachment ambiguity resolution. Prosodic boundaries were manipulated in globally ambiguous relative clause sentences. The results show that high span listeners are more likely to be sensitive to the distinction between different types of prosodic boundaries than low span listeners. The findings suggest that like high-level constraints, the use of low-level prosodic information is resource demanding.

  • PDF

Effects of phonological and phonetic information of vowels on perception of prosodic prominence in English

  • Suyeon Im
    • Phonetics and Speech Sciences
    • /
    • v.15 no.3
    • /
    • pp.1-7
    • /
    • 2023
  • This study investigates how the phonological and phonetic information of vowels influences prosodic prominence among linguistically untrained listeners using public speech in American English. We first examined the speech material's phonetic realization of vowels (i.e., maximum F0, F0 range, phone rate [as a measure of duration considering the speech rate of the utterance], and mean intensity). Results showed that the high vowels /i/ and /u/ likely had the highest max F0, while the low vowels /æ/ and /ɑ/ tended to have the highest mean intensity. Both high and low vowels had similarly high phone rates. Next, we examined the effects of the vowels' phonological and phonetic information on listeners' perceptions of prosodic prominence. The results showed that vowels significantly affected the likelihood of perceived prominence independent of acoustic cues. The high and low vowels affected probability of perceived prominence less than the mid vowels /ɛ/ and /ʌ/, although the former two were more likely to be phonetically enhanced in the speech than the latter. Overall, these results suggest that perceptions of prosodic prominence in English are not directly influenced by signal-driven factors (i.e., vowels' acoustic information) but are mediated by expectation-driven factors (e.g., vowels' phonological information).