• 제목/요약/키워드: speech dimensions

검색결과 28건 처리시간 0.022초

Articulatory Manifestation of Prosodic Strengthening in English /i/ and /I/

  • Kim, Sa-Hyang;Cho, Tae-Hong
    • 말소리와 음성과학
    • /
    • 제3권4호
    • /
    • pp.13-21
    • /
    • 2011
  • The present study investigated the effects of two different sources of prosodic strengthening, i.e., boundary and accent, in the articulation of English high front vowels, /i/ and /I/. The vowels were investigated in vowel-initial ('eat' vs. 'it'), /h/-initial ('heat' vs. 'hit') and /p/-initial words ('Pete' vs. 'pit'), which were placed in varying prosodic conditions. Using Electromagnetic Articulograph (EMA), the tongue dorsum positions in the x and y dimensions, the lip opening and the jaw opening (lowering) were measured. With respect to the boundary-induced strengthening, results showed that /i/ and /I/ in vowel-initial words ('eat' - 'it') are produced with a higher tongue position in the domain-intial than domain-medial positions. The fact that the vowels only in the vowel-initial condition showed the domain-intial strengthening (DIS) effect suggests that the DIS effect is localized mainly to the initial position (the locality account). As for the accent-induced strengthening, vowels were produced with a more fronted tongue position and larger lip opening in accented than unaccented positions. This suggests that the presence of accent increases overall sonority of the vowels in various prosodic contexts, and enhances primarily the frontedness of the front high vowels. Taken together, the results indicate that the two types of prosodic strengthening are articulatorily realized differently, supporting the view that they are encoded separately in the speech planning process. The present study also showed the distinction between the two high front vowels in the tongue position (in both the frontedness and the height dimensions), while the jaw did not seem to contribute to the distinction robustly, suggesting that the tongue contributes more in distinguishing the two vowels than the jaw does.

  • PDF

자연스러운 정서 반응의 범주 및 차원 분류에 적합한 음성 파라미터 (Acoustic parameters for induced emotion categorizing and dimensional approach)

  • 박지은;박정식;손진훈
    • 감성과학
    • /
    • 제16권1호
    • /
    • pp.117-124
    • /
    • 2013
  • 본 연구는 음성 인식기에서 일반적으로 사용되는 음향적 특징인 MFCC, LPC, 에너지, 피치 관련 파라미터들을 이용하여 자연스러운 음성의 정서를 범주 및 차원으로 얼마나 잘 인식할 수 있는지 살펴보았다. 자연스러운 정서 반응 데이터를 얻기 위해 선행 연구에서 이미 타당도와 효과성이 밝혀진 정서 유발 자극을 사용하였고, 110명의 대학생들에게 7가지 정서 유발 자극을 제시한 후 유발된 음성 반응을 녹음하여 분석에 사용하였다. 각 음성 데이터에서 추출한 파라미터들을 독립변인으로 하여 선형 판별 분석(LDA)으로 7가지 정서 범주를 분류하였고, 범주 분류의 한계를 극복하기 위해 단계별 다중회귀(stepwise multiple regression) 모형을 도출하여 4가지 정서 차원(valence, arousal, intensity, potency)을 가장 잘 예측하는 음성 특징 파라미터를 산출하였다. 7가지 정서 범주 판별율은 평균 62.7%이었고, 4 차원 예측 회귀모형들도 p<.001수준에서 통계적으로 유의하였다. 결론적으로, 본 연구 결과는 자연스러운 감정의 음성 반응을 분류하는데 유용한 파라미터들을 선정하여 정서의 범주와 차원적 접근으로 정서 분류 가능성을 보였으며 논의에 본 연구의 개선방향에 대해 기술하였다.

  • PDF

PCA 퍼지 혼합 모델을 이용한 화자 식별 (Speaker Identification Using PCA Fuzzy Mixture Model)

  • 이기용
    • 음성과학
    • /
    • 제10권4호
    • /
    • pp.149-157
    • /
    • 2003
  • In this paper, we proposed the principal component analysis (PCA) fuzzy mixture model for speaker identification. A PCA fuzzy mixture model is derived from the combination of the PCA and the fuzzy version of mixture model with diagonal covariance matrices. In this method, the feature vectors are first transformed by each speaker's PCA transformation matrix to reduce the correlation among the elements. Then, the fuzzy mixture model for speaker is obtained from these transformed feature vectors with reduced dimensions. The orthogonal Gaussian Mixture Model (GMM) can be derived as a special case of PCA fuzzy mixture model. In our experiments, with having the number of mixtures equal, the proposed method requires less training time and less storage as well as shows better speaker identification rate compared to the conventional GMM. Also, the proposed one shows equal or better identification performance than the orthogonal GMM does.

  • PDF

Voice onset time in English and Korean stops with respect to a sound change

  • Kim, Mi-Ryoung
    • 말소리와 음성과학
    • /
    • 제13권2호
    • /
    • pp.9-17
    • /
    • 2021
  • Voice onset time (VOT) is known to be a primary acoustic cue that differentiates voiced from voiceless stops in the world's languages. While much attention has been given to the sound change of Korean stops, little attention has been given to that of English stops. This study examines VOT of stop consonants as produced by English speakers in comparison to Korean speakers to see whether there is any VOT change for English stops and how the effects of stop, place, gender, and individual on VOT differ cross-linguistically. A total of 24 native speakers (11 Americans and 13 Koreans) participated in this experiment. The results showed that, for Korean, the VOT merger of lax and aspirated stops was replicated, and, for English, voiced stops became initially devoiced and voiceless stops became heavily aspirated. English voiceless stops became longer in VOT than Korean counterparts. The results suggest that, similar to Korean stops, English stops may also undergo a sound change. Since it is the first study to be revealed, more convincing evidence is necessary.

K-L 전개를 이용한 연속 숫자음 인식에 관한 연구 (A Study on Connected Digits Recognition Using the K-L Expansion)

  • 김주곤;오세진;황철준;김범국;정현열
    • 융합신호처리학회논문지
    • /
    • 제2권3호
    • /
    • pp.24-31
    • /
    • 2001
  • K-L 전개 방법은 특징의 차원을 효과적으로 압축하므로 인식 처리에서 계산량을 줄일 수 있는 방법으로 잘 알려져 있다. 본 논문에서는 한국어 인식 시스템의 인식 정도를 개선하기 위해, 음성의 특징 파라미터에 대하여 효과적으로 K-L전개를 적용하는 방법(K-L 계수)을 제안한다. 그리고 제안한 방법으로 얻어진 새로운 음성 특징 파라미터를 이용하여 화자 독립 연속 숫자음 인식실험을 수행하고, 기존의 Mel-cepstrum과 회귀계수의 인식 결과와 비 교, 분석하였다. 인식 실험 결과, 제안한 K-L 계수를 이용한 방법이 기존의 방법보다 높은 인식률을 얻어 제안한 방법의 유효성을 확인할 수 있었다.

  • PDF

A Study on the Syllable Recognition Using Neural Network Predictive HMM

  • Kim, Soo-Hoon;Kim, Sang-Berm;Koh, Si-Young;Hur, Kang-In
    • The Journal of the Acoustical Society of Korea
    • /
    • 제17권2E호
    • /
    • pp.26-30
    • /
    • 1998
  • In this paper, we compose neural network predictive HMM(NNPHMM) to provide the dynamic feature of the speech pattern for the HMM. The NNPHMM is the hybrid network of neura network and the HMM. The NNPHMM trained to predict the future vector, varies each time. It is used instead of the mean vector in the HMM. In the experiment, we compared the recognition abilities of the one hundred Korean syllables according to the variation of hidden layer, state number and prediction orders of the NNPHMM. The hidden layer of NNPHMM increased from 10 dimensions to 30 dimensions, the state number increased from 4 to 6 and the prediction orders increased from 10 dimensions to 30 dimension, the state number increased from 4 to 6 and the prediction orders increased from the second oder to the fourth order. The NNPHMM in the experiment is composed of multi-layer perceptron with one hidden layer and CMHMM. As a result of the experiment, the case of prediction order is the second, the average recognition rate increased 3.5% when the state number is changed from 4 to 5. The case of prediction order is the third, the recognition rate increased 4.0%, and the case of prediction order is fourth, the recognition rate increased 3.2%. But the recognition rate decreased when the state number is changed from 5 to 6.

  • PDF

설소대의 크기와 운동이 발음에 미치는 영향 (THE EFFECT OF THE LENGTH OF THE LINGUAL FRENUM AND THE TONGUE MOTION ON SPEECH)

  • 박성희;손우성;김용덕;신상훈;김욱규;정인교;권순복
    • Journal of the Korean Association of Oral and Maxillofacial Surgeons
    • /
    • 제27권6호
    • /
    • pp.526-534
    • /
    • 2001
  • Purpose : The objective of this study is to ascertain whether the positive exists among the frenum length, the tongue movement and the speech and to present the normal range of tongue movement and guidelines for the choice of surgery, observation if necessary. Materials and Methods : 180 patients were evaluated. We divided 180 patients into 6 group by age. Each group was separated as follows; the age of 2.5-4, 5-6, 7-9, 10-12, 16-18. We measured the frenal length, the range of tongue motion and evaluated the speech so that we really questioned about the positive relationship between the tongue-tie and speech. We let the patient exercise the protrusive both(right, left) laterotrusive superior movement of the tongue. During these movements, we measured the distance between the vermilion border and the tongue tip. We also measured the distance from the tongue tip to the point contacting the upper lip with dorsum of the tongue during the maximal protrusive movement of the tongue. Three linear measurement of the anterior, inferior segment of the tongue including the lingual frenum, are made. These measurements are as follows: 1. Distance A. Free anterior portion of the tongue from the point of frenular insertion to the tongue tip. 2. Distance B. The distance from the initiating point of the lingual frenum to the point connecting the two sublingual carundcles to the lingual frenum perpendicularly. 3. Distance C. The distance from the point contacting the line crossing the sublingual caruncles with the lingual frenum to the terminating point of the lingual frenum. We transform three linear measures into a statistical ratio, A/(A-B+C), representing the length of the free portion of the tongue compared with the total sublingual dimensions. In addition, we assessed the speech through Picture Consonant Articulation Test(PCAT) and tried to find out the relationship between the length of the lingual frenum and speech. Conclusion : As people are born, they have small and restricted tongue. As people grow old, tongue motions are more liberate, and unrestricted and they can speak so freely. Therefore we suggest that until age 5, oral and maxillofacial surgeons postpone the surgery if not urgent, evaluate the maximal lingual motions and PCAT according to this article and observe their changes.

  • PDF

L2 Proficiency Effect on the Acoustic Cue-Weighting Pattern by Korean L2 Learners of English: Production and Perception of English Stops

  • Kong, Eun Jong;Yoon, In Hee
    • 말소리와 음성과학
    • /
    • 제5권4호
    • /
    • pp.81-90
    • /
    • 2013
  • This study explored how Korean L2 learners of English utilize multiple acoustic cues (VOT and F0) in perceiving and producing the English alveolar stop with a voicing contrast. Thirty-four 18-year-old high-school students participated in the study. Their English proficiency level was classified as either 'high' (HEP) or 'low' (LEP) according to high-school English level standardization. Thirty different synthesized syllables were presented in audio stimuli by combining a 6-step VOTs and a 5-step F0s. The listeners judged how close the audio stimulus was to /t/ or /d/ in L2 using a visual analogue scale. The L2 /d/ and /t/ productions collected from the 22 learners (12 HEP, 10 LEP) were acoustically analyzed by measuring VOT and F0 at the vowel onset. Results showed that LEP listeners attended to the F0 in the stimuli more sensitively than HEP listeners, suggesting that HEP listeners could inhibit less important acoustic dimensions better than LEP listeners in their L2 perception. The L2 production patterns also exhibited a group-difference between HEP and LEP in that HEP speakers utilized their VOT dimension (primary cue in L2) more effectively than LEP speakers. Taken together, the study showed that the relative cue-weighting strategies in L2 perception and production are closely related to the learner's L2 proficiency level in that more proficient learners had a better control of inhibiting and enhancing the relevant acoustic parameters.

모음-자음-모음 연결에서 자음의 조음특성과 모음-모음 동시조음 (Consonantal Production and V-to-V Coarticulation in Korean VCV Sequences)

  • 신지영
    • 음성과학
    • /
    • 제1권
    • /
    • pp.55-81
    • /
    • 1997
  • In the present paper, V-to-V coarticulation in Korean VCV sequences is discussed, focusing on links between consonantal production and degree of V-to-V coarticulation. Temporal and spatial differences between three types of Korean alveolar stops (lax /t/. aspirated /$t^h$/ and thense /t'/) are examined from VCV sequences involving all possible combinations of three Korean unrounded vowels /a, i,/ based on spectrographic and electrographic data(two male speakers and one female speaker and one female speaker respectively). Closure duration and voice onset time (VOT) were measured from acoustic data. 'Total duration', which is defined as the sum of the closure duration and the VOT, was also calculated in order to see the temporal distance between two vowels in a VCV sequence. Differences in lingual-palatal contact pattern at the maximum contact (MC) point between the three types of stop were observed from EPG data. V-to-V coarticulation was investigated by measuring the offset or onset of the second formant (F2) of the target vowels from spectrograms. Two different dimensions of articulation, temporal and spatial, seem to playa role in determining the degree of V-to-V coarticulation. The degree of V-to-V anticipatory coarticulation is influenced by the spatial characteristics of the intervening consonant while the degree of carryover coarticulation is influenced by the temporal characteristics of the consonant.

  • PDF

Distinct Segmental Implementations in English and Spanish Prosody

  • Lee, Joo-Kyeong
    • 음성과학
    • /
    • 제11권4호
    • /
    • pp.199-206
    • /
    • 2004
  • This paper attempts to provide a substantial explanation of different prosodic implementations on segments in English and Spanish, arguing that the phonetic modification invoked by prosody may effectively reflect phonological structure. In English, a high front vowel in accented syllables is acoustically realized as higher F1 and F2 frequencies than in unaccented syllables, due to its more peripheral and sonorous articulation (Harrington et al. 1999). In this paper, an acoustic experiment was conducted to see if such a manner of segmental modification invoked by prosody in English extends to other languages such as Spanish. Results show that relatively more prominent syllables entail higher F1 values as a result of their more sonorous articulation in Spanish, but either front or back vowel does not show a higher F2 or a lower F2 frequency. This is interpreted as an indication that a prosodically prominent syllable entails its vocalic enhancement in both horizontal and vertical dimensions of articulation in English. In Spanish, however, only the vertical dimensional articulation is maximized, resulting in a higher F1. I suggest that this difference may be attributed to the different phonological structures of vowels in English and Spanish, and that sonority expansion alone would be sufficient in the articulation of prosodic prominence as long as the phonological distinction of vowels is well retained.

  • PDF