• Title/Summary/Keyword: Speech function

Search Result 696, Processing Time 0.023 seconds

음성 강화를 위한 a priori SNR 추정기반 적응 바람소리 저감 방법 (An Adaptive Wind Noise Reduction Method Based on a priori SNR Estimation for Speech Eenhancement)

  • 서지훈;이석필
    • 전기학회논문지
    • /
    • 제64권12호
    • /
    • pp.1756-1760
    • /
    • 2015
  • This paper focuses on a priori signal to noise ratio (SNR) estimation method for the speech enhancement. There are many researches for speech enhancement with several ambient noise cancellation methods. The method based on spectral subtraction (SS) which is widely used in noise reduction has a trade-off between the performance and the distortion of the signals. So the need of adaptive method like an estimated a priori SNR being able to making a high performance and low distortion is increasing. The decision directed (DD) approach is used to determine a priori SNR in noisy speech signals. A priori SNR is estimated by using only the magnitude components and consequently follows a posteriori SNR with one frame delay. We propose a modified a priori SNR estimator and the weighted rational transfer function for speech enhancement with wind noises. The experimental result shows the performance of our proposed estimator is better Perceptual Evaluation of Speech Quality scores (PESQ, ITU-T P.862) compare to the conventional DD approach-based systems and different noise reduction methods.

A Study on Pitch Period Detection Algorithm Based on Rotation Transform of AMDF and Threshold

  • 서현수;김남호
    • 융합신호처리학회논문지
    • /
    • 제7권4호
    • /
    • pp.178-183
    • /
    • 2006
  • As a lot of researches on the speech signal processing are performed due to the recent rapid development of the information-communication technology. the pitch period is used as an important element to various speech signal application fields such as the speech recognition. speaker identification. speech analysis. or speech synthesis. A variety of algorithms for the time and the frequency domains related with such pitch period detection have been suggested. One of the pitch detection algorithms for the time domain. AMDF (average magnitude difference function) uses distance between two valley points as the calculated pitch period. However, it has a problem that the algorithm becomes complex in selecting the valley points for the pitch period detection. Therefore, in this paper we proposed the modified AMDF(M-AMDF) algorithm which recognizes the entire minimum valley points as the pitch period of the speech signal by using the rotation transform of AMDF. In addition, a threshold is set to the beginning portion of speech so that it can be used as the selection criteria for the pitch period. Moreover the proposed algorithm is compared with the conventional ones by means of the simulation, and presents better properties than others.

  • PDF

Differentiation of Aphasic Patients from the Normal Control Via a Computational Analysis of Korean Utterances

  • Kim, HyangHee;Choi, Ji-Myoung;Kim, Hansaem;Baek, Ginju;Kim, Bo Seon;Seo, Sang Kyu
    • International Journal of Contents
    • /
    • 제15권1호
    • /
    • pp.39-51
    • /
    • 2019
  • Spontaneous speech provides rich information defining the linguistic characteristics of individuals. As such, computational analysis of speech would enhance the efficiency involved in evaluating patients' speech. This study aims to provide a method to differentiate the persons with and without aphasia based on language usage. Ten aphasic patients and their counterpart normal controls participated, and they were all tasked to describe a set of given words. Their utterances were linguistically processed and compared to each other. Computational analyses from PCA (Principle Component Analysis) to machine learning were conducted to select the relevant linguistic features, and consequently to classify the two groups based on the features selected. It was found that functional words, not content words, were the main differentiator of the two groups. The most viable discriminators were demonstratives, function words, sentence final endings, and postpositions. The machine learning classification model was found to be quite accurate (90%), and to impressively be stable. This study is noteworthy as it is the first attempt that uses computational analysis to characterize the word usage patterns in Korean aphasic patients, thereby discriminating from the normal group.

Fillers in the Hong Kong Corpus of Spoken English (HKCSE)

  • Seto, Andy
    • 아시아태평양코퍼스연구
    • /
    • 제2권1호
    • /
    • pp.13-22
    • /
    • 2021
  • The present study employed an analytical framework that is characterised by a synthesis of quantitative and qualitative analyses with a specially designed computer software SpeechActConc to examine speech acts in business communication. The naturally occurring data from the audio recordings and the prosodic transcriptions of the business sub-corpora of the HKCSE (prosodic) are manually annotated with a speech act taxonomy for finding out the frequency of fillers, the co-occurring patterns of fillers with other speech acts, and the linguistic realisations of fillers. The discoursal function of fillers to sustain the discourse or to hold the floor has diverse linguistic realisations, ranging from a sound (e.g. 'uhuh') and a word (e.g. 'well') to sounds (e.g. 'um er') and words, namely phrase ('sort of') and clause (e.g. 'you know'). Some are even combinations of sound(s) and word(s) (e.g. 'and um', 'yes er um', 'sort of erm'). Among the top five frequent linguistic realisations of fillers, 'er' and 'um' are the most common ones found in all the six genres with relatively higher percentages of occurrence. The remaining more frequent realisations consist of clause ('you know'), word ('yeah') and sound ('erm'). These common forms are syntactically simpler than the less frequent realisations found in the genres. The co-occurring patterns of fillers and other speech acts are diverse. The more common co-occurring speech acts with fillers include informing and answering. The findings show that fillers are not only frequently used by speakers in spontaneous conversation but also mostly represented in sounds or non-linguistic realisations.

편도암 절제술후 전완유리피판술을 이용한 연구개 결손부 재건의 기능적 결과 (Functional Results of Soft Palate Defect Reconstruction using Radial Forearm Free Flap after Tonsil Cancer Surgery)

  • 김민식;선동일;박해섭;조승호;제현순
    • 대한기관식도과학회지
    • /
    • 제5권2호
    • /
    • pp.191-197
    • /
    • 1999
  • Background and Objective : Soft palate plays a great role in function of speech and swallowing. Ablation of tonsil cancer results in multi-demensional defect including soft palate in most cases and restoration of the postoperative oral cavity function is a continuing surgical challenge. Although a variety of techniques are available, radial forearm free flap has been known as an effective method for these defect, which offers a thin, pliable, and relatively hairless skin, and a long vascular pedicle. The aim of the present study is to report the speech and swallowing function test results of our 5 consecutive radial forearm free flaps used for tonsil cancers. Materials and Methods : We reviewed the medical records of 5 patients who were offered intraoral reconstruction with a radial forearm free flap after ablative surgery for tonsil cancers, from Dec. 1997 to Oct. 1998, and analyzed the surgical methods, complications, and speech and swallowing function test results. We have examined with modified barium swallow to evaluate postoperative wallowing function and articulation and resonance test for speech. Results : The tumor sizes by TNM stage(AJCC, 1997) were T1(1), T2(2), and T4(3). The paddles of flaps were tailored in multilobed designs from oval shape to pentalobed design and in variable size from 24$cm^2$ to 108$cm^2$(average size = 78.4$cm^2$), according to the defect after ablation. This procedures resulted in satisfactory flap success and functional results all but 1 case of flap contracture in 2 postoperative week, achieved early oral diet until 16-57 postoperative day(average, 28 days) and social speech. The oropharyngeal defect including soft palate reconstruction with radial forearm free flap might be an excellent method for the maximal functional results, after ablative surgery of tonsil cancer that results in multidimensional defect.

  • PDF

구개열(口蓋裂) 환자(患者)에 있어서 구개(口蓋) 성형술후(成形術後) 비인강(鼻咽腔) 폐쇄(閉鎖)에 관(關)한 임상적(臨床的) 연구(硏究) (CLINICAL STUDY OF VELOPHARYNGEAL CLOSURE AFTER THE PRIMARY PALATORRHAPHY IN CLEFT PALATE PATIENTS)

  • 고광희;신효근
    • Maxillofacial Plastic and Reconstructive Surgery
    • /
    • 제14권1_2호
    • /
    • pp.1-21
    • /
    • 1992
  • In order to find the causes of velopharyngeal incompetency after primary palatorrhaphy in cleft patients, we analyzed the form and function of the velopharyngeal space of fifteen operated cleft palate patients and five normal subjects. The velopharyngeal function was evaluated by lateral cephalometric radiography, velopharyngography and hypernasality cul-de-sac test. The obtained results were as follows. 1. The rate of velopharyngeal incompetency was twenty percent, three of the fifteen operated patients. Two of them were complete cleft palate and the other was incomplete one. 2. The length of soft palate and levator eminence were longer in normal group than those of good speech group and complete cleft palate group during phonation of /i/ (P<0.05). The lengthening rate of soft palate was smaller in good and poor speech group than that of normal group(P<0.05), and, reduced in order, normal group, complete cleft palate group and incomplete palate group(P<0.05). 3. The nasopharyngeal distance had no significant difference between all groups at rest, but, smaller in normal group than that of both cleft palate group(P<0.05), good speech group and poor speech group(P<0.05) during phonation of /i/ The difference in nasopharyngeal distance between rest and /i/ phonation was greater in normal group than that of both cleft palate group, good speech group and poor speech group. 4. The moving distance of sop palate reduced in order, normal group, incomplete cleft palate group, complete cleft palate group(P<0.05). 5. The distance between lateral pharyngeal wall had no significant difference between all groups in rest, but, smaller than that of complete cleft palate group in normal group(P<0.01) and increased in order normal group, good speech group, poor speech group(P<0.01) during phonation of /a/. The mobility of lateral wall was reduced in order, normal group, good speech group poor speech group(P<0. 01). 6. There was low corelationship between the mobility of lateral pharyngeal wall and soft palate. Therfore, it suggest that the movements of lateral pharyngeal wall and soft palate occurs independently.

  • PDF

효과적인 복소 스펙트럼 기반 음성 향상을 위한 시간과 주파수 영역 손실함수 조합에 관한 연구 (A study on loss combination in time and frequency for effective speech enhancement based on complex-valued spectrum)

  • 정재희;김우일
    • 한국음향학회지
    • /
    • 제41권1호
    • /
    • pp.38-44
    • /
    • 2022
  • 잡음에 오염된 음성의 명료도와 음질을 향상시키고자 음성 향상을 수행한다. 본 연구에서는 복소값 스펙트럼을 이용한 마스크기반 음성 향상에서 시간 영역 손실함수와 주파수 영역 손실함수에 따른 학습 결과를 비교하였다. 시간 영역의 음성 파형과 주파수 영역의 스펙트럼의 세부정보를 고려해 두 영역의 장점을 활용할 수 있도록 손실함수 조합에 관해 연구를 진행하였다. 시간 영역 손실함수는 Scale Invariant-Source to Noise Ratio(SI-SNR)을 이용해 계산하고, 주파수 영역 손실함수는 복소값 스펙트럼과 크기 스펙트럼을 Mean Squared Error(MSE)로 계산하여 사용하였고, sin 함수를 이용해 위상에 대한 손실함수를 계산하였다. 손실함수 조합은 시간 영역 손실함수인 SI-SNR과 각 주파수 영역 손실함수를 조합하였다. 또한 크기 값과 위상 값을 모두 고려할 수 있도록 SI-SNR과 크기 스펙트럼, 위상에 관련된 손실함수들도 조합하여 실험을 진행하였다. 음성 향상 결과는 Source-to-Distortion Ratio(SDR), Perceptual Evaluation of Speech Quality(PESQ), Short-Time Objective Intelligibility(STOI)를이용해 성능 비교 평가를 진행하였다. 음성 향상 결과를 확인해보기 위해 스펙트럼 상에서 비교를 진행하였다. TIMIT 데이터베이스를 이용한 실험 결과, 시간 영역 또는 주파수 영역 손실함수보다 SI-SNR과 크기 스펙트럼을 조합한 손실함수를 사용하여 음성 향상을 학습했을 때 가장 높은 성능을 보였다.

강인 음성 인식을 위한 가중화된 음원 분산 및 잡음 의존성을 활용한 보조함수 독립 벡터 분석 기반 음성 추출 (Speech extraction based on AuxIVA with weighted source variance and noise dependence for robust speech recognition)

  • 신의협;박형민
    • 한국음향학회지
    • /
    • 제41권3호
    • /
    • pp.326-334
    • /
    • 2022
  • 이 논문에서는 배경 잡음이 포함되는 환경에서 강인한 음성 인식을 하기 위한 전처리 단계로서 쓰이는 목표 음성 향상 방법을 제안한다. 보조 함수 기반의 독립 벡터 분석(Auxiliary-function-based Independent Vector Analysis, AuxIVA) 기법을 기반으로 가중 공분산 행렬에서 시간에 따라 변하는 분산에 의해서 가중치가 결정된다. 목표 음성에 대한 시간-주파수별 기여도를 나타내는 마스크를 통해 분산의 크기를 조절한다. 이러한 마스크는 음성 향상을 위해서 학습된 신경망 혹은 목표 화자로부터의 직선 성분의 기여도를 찾기 위한 확산성으로부터 추정할 수 있다. 이에 더하여 둘러싼 잡음에 대한 출력들은 서로 다차원 독립 성분 분석을 도입하여 의존성을 주어 안정적으로 노이즈 성분을 추출할 수 있다. 이 AuxIVA 기반의 목표 음성 추출 알고리즘은 또한 노이즈에 대해서 비음수 행렬 분해(Non-negative Matrix Factorization, NMF)를 비음수 텐서 분해(Non-negative Tensor Factorization, NTF)로 확장하여 독립 단순 행렬 분석(Independent Low-Rank Matrix Analysis, ILRMA)의 틀에서도 수행될 수 있다. 이러한 확장을 통해서 여전히 잡음 출력 채널에서의 채널간 의존성을 유지할 수 있다. CHiME-4데이터셋에 대한 실험 결과는 소개된 알고리즘에 대한 효과를 보여준다.

성악인의 발성능력 향상에 Vocal Function Exercise가 미치는 영향 (The Study on the Effects of Vocal Function Exercise for Trained Singers)

  • 권영경;심현섭;진성민;정성민
    • 음성과학
    • /
    • 제10권2호
    • /
    • pp.169-189
    • /
    • 2003
  • Trained singers, one group of professional voice users, have much more interest on the voice than common people, and on its management, too. They train for singing beautiful songs, and, at the same time, try for efficient voice production. The present study was performed with three tenors and three baritones, undergraduate students majored in classical singing, to investigate the degree of improvement of their voice production efficiency through vocal function exercise, by measuring the three dependent variables, maximum phonation time, speed quotient of glottal contact, and the number of semi tones. For the baseline establishment, dependent variables were measured 3$\sim$6 times for two weeks. Then, the subjects exercised vocal function exercise for seven weeks, and after the termination of training, evaluation was performed four times for two weeks, to find the maintenance of the training effect. Vocal function exercise is composed of four successive steps: warm-up, stretching exercise, contracting exercise, power exercise. As results, all of six subjects showed improvement in the aspect of maximum phonation time, speed quotient if glottal contact, and the number of semitones.

  • PDF

노년층의 담화 산출 특성: 노화, 성별, 교육정도에 따른 차이 (Discourse Characteristics in Healthy Elderly: Effects of Aging, Gender and Educational Level)

  • 최현주
    • 말소리와 음성과학
    • /
    • 제4권2호
    • /
    • pp.135-143
    • /
    • 2012
  • Discourse is regarded as an important component of communication assessment, but studies about the discourse characteristics of the elderly are scant. The purpose of this study was to confirm the effects of aging, gender, and educational level on discourse in elderly people with normal cognitive function. Forty normal elderly and forty young people participated in this study. A picture description task (Boston Cookie-Theft picture) was used to examine discourse function. The description task was analyzed for both productivity (total number of sentences, total number of syllables, and syllables per sentence) and semantics (CIU ratio). The results were as follows: 1) Only CIU ratio differed significantly according to age. 2) In the total number of syllables and syllables per sentence, females demonstrate a higher number than males. 3) The CIU ratio differed significantly according to educational level. These results suggest that impairment of communicative function is an aspect of cognitive impairment that can be related to aging. Also, discourse performance in the elderly is associated with their gender and educational level.