• 제목/요약/키워드: vocal process

검색결과 75건 처리시간 0.021초

음성 변환을 사용한 감정 변화에 강인한 음성 인식 (Emotion Robust Speech Recognition using Speech Transformation)

  • 김원구
    • 한국지능시스템학회논문지
    • /
    • 제20권5호
    • /
    • pp.683-687
    • /
    • 2010
  • 본 논문에서는 인간의 감정 변화에 강인한 음성 인식 시스템을 구현하기 위하여 음성 변환 방법 중의 한가지인 주파수 와핑 방법을 사용한 연구를 수행하였다. 이러한 목표를 위하여 다양한 감정이 포함된 음성 데이터베이스를 사용하여 감정의 변화에 따라 음성의 스펙트럼이 변화한다는 것과 이러한 변화는 음성 인식 시스템의 성능을 저하시키는 원인 중의 하나임을 관찰하였다. 본 논문에서는 이러한 음성의 변화를 감소시키는 방법으로 주파수 와핑을 학습 과정에 사용하는 방법을 제안하여 감정 변화에 강인한 음성 인식 시스템을 구현하였고 성도 길이 정규화 방법을 사용한 방법과 성능을 비교하였다. HMM을 사용한 단독음 인식 실험에서 제안된 학습 방법은 사용하면 감정이 포함된 데이터에 대한 인식 오차가 기존 방법보다 감소되었다.

Speech Recognition by Neural Net Pattern Recognition Equations with Self-organization

  • Kim, Sung-Ill;Chung, Hyun-Yeol
    • The Journal of the Acoustical Society of Korea
    • /
    • 제22권2E호
    • /
    • pp.49-55
    • /
    • 2003
  • The modified neural net pattern recognition equations were attempted to apply to speech recognition. The proposed method has a dynamic process of self-organization that has been proved to be successful in recognizing a depth perception in stereoscopic vision. This study has shown that the process has also been useful in recognizing human speech. In the processing, input vocal signals are first compared with standard models to measure similarities that are then given to a process of self-organization in neural net equations. The competitive and cooperative processes are conducted among neighboring input similarities, so that only one winner neuron is finally detected. In a comparative study, it showed that the proposed neural networks outperformed the conventional HMM speech recognizer under the same conditions.

삽관성 육아종과 접촉성 육아종에 대한 치료 결과 분석 (Treatment Results of Vocal Process Granuloma: Intubation Versus Contact Granuloma)

  • 정진욱;오재환;김슬;김동영;우주현
    • 대한후두음성언어의학회지
    • /
    • 제32권3호
    • /
    • pp.135-141
    • /
    • 2021
  • Background and Objectives Vocal process granulomas (VPGs) are benign lesions of the larynx, typically contact granulomas (CG) and intubation granulomas (IG). The two diseases are known to have different clinical manifestations despite having the same pathological features. The purpose of this study was to analyze the treatment results for CG and IG and to obtain clinical information. Materials and Method We retrospectively reviewed the medical records of patients diagnosed with VPG between January 2015 and December 2018. The patient's age, sex, medical history, lesion size, lesion type, reflux finding score, response to treatment, duration of treatment, and follow-up period were compared. Results In total, 32 patients were included in the study, of which 18 were CG and 14 were IG. In the CG group, males were dominant (n=15, 83.3%), whereas in the IG group, females were dominant (n=11, 78.6%) (p=0.0009). The response to medical treatment using proton pump inhibitor and steroid inhaler was better in the IG group (11/14, 78.6%) than in the CG group (7/18, 38.9%) (p=0.036). Of the 14 patients who did not respond to medical treatment, 5 received botulium toxin injections, and all 5 had complete remission. The duration of medical treatment was significantly longer in the IG group (p=0.0029). Conclusion IG was more common in female, and CG was more dominant in male. IG had better response to medical treatment using proton pump inhibitor and steroid inhaler than CG.

위상 보상된 고조파 스케일링에 의한 음성합성용 피치변경법 (On a Pitch Alteration Method using Scaling the Harmonics Compensated with the Phase for Speech Synthesis)

  • 배명진
    • 한국음향학회지
    • /
    • 제13권6호
    • /
    • pp.91-97
    • /
    • 1994
  • 신호처리에서, 파형부화법은 음성신호의 잉여성분을 감소시킴으로써 파형을 유지하는 부호화 방법이다. 음성 합성의 경우, 고음질의 파형부호화법은 주로 분석에 의한 합성법에 이용된다. 그러나, 파형부호화법은 여기 파라미터와 성도 파라미터로 분리하지 않고 처리하기 때문에 규칙에 의한 합성에 적용되기 어렵다. 따라서 파형부호화법을 규칙에 의한 합성에 이용하기 위해서는 피치변경이 필요하다. 본 논문에서, 우리는 파형부호화법에서 음성신호를 성도 파라미터와 여기 파라미터로 분리함으로써 피치 주기를 바꿀 수 있는 새로운 피치변경법을 제안한다. 이 방법은 시-주파수 혼성영억 방법으로 시간영역에서 파형의 위상성분과 주파수영역에서 파형의 진폭성분을 보존한다. 따라서 파형부호화법은 음성처리에 있어 규칙에 의한 합성을 할 수 있다. 본 논문에서 제안한 알고리즘을 이용한 경우, 단지 $2.94\%의$ 스펙트럼 왜곡만이 일어났다. 즉, 스펙트럼 왜곡이 시간영역에서의 피치변경법보다 $5.06\%$ 이상 감소되었다.

  • PDF

스테레오 음악 신호에서의 보컬 음원 분리를 위한 통합 알고리즘 (A Unified Method for Vocal Source Separation From Stereophonic Music Signals)

  • 김민제;장인선;강경옥
    • 대한전자공학회논문지SP
    • /
    • 제47권5호
    • /
    • pp.89-99
    • /
    • 2010
  • 본 논문에서는 스테레오 형식의 음악 신호에서 가창 신호와 같은 음원을 분리하기 위한 통합 알고리즘을 제시한다. 스테레오 형식의 음악 신호에서 특정한 악기 음원을 분리하는 문제는, 획득한 음악 신호가 다양한 악기들이 동시에 연주되는 혼합신호라는 점을 고려하고, 각각의 악기를 음원이라고 가정할 때, 획득한 혼합 신호의 개수가 음원의 개수보다 적은 비결정(underdetermined) 환경에서의 음원 분리 문제가 된다. 비결정 환경에서는 신호가 혼합되는 공간에 대한 가정을 기반하는 전통적 음원 분리 방식을 적용하기 힘들며, 목표 음원의 특정한 특성을 활용하여 추출하게 된다. 본 논문에서 제안하는 통합 알고리즘은 이종의 특성을 활용하는 음악 음원 분리 알고리즘들을 유기적으로 통합하는 구조이며, 구체적으로는 가창 신호와 같은 특정한 음원 추출을 위해 주로 사용되어 왔던 스테레오 채널 정보를 활용하는 방식과, 모노 혼합 신호에서 두드러지는 음원의 음정을 이용하여 음원을 추출하는 두 가지 방식을 통합하는 것을 목표로 한다. 본 논문에서 제안하는 구조는 각각의 음악 음원 분리 알고리즘이 가지고 있는 고유의 약점을 해소함으로써, 목표 음원의 복원 신호가 통합 과정에 의해 향상될 수 있다는 강점이 있으며, 그것을 실제 상업 음악 콘텐츠를 대상으로 한 실험을 통해 검증한다.

Electromyographic evidence for a gestural-overlap analysis of vowel devoicing in Korean

  • Jun, Sun-A;Beckman, M.;Niimi, Seiji;Tiede, Mark
    • 음성과학
    • /
    • 제1권
    • /
    • pp.153-200
    • /
    • 1997
  • In languages such as Japanese, it is very common to observe that short peripheral vowel are completely voiceless when surrounded by voiceless consonants. This phenomenon has been known as Montreal French, Shanghai Chinese, Greek, and Korean. Traditionally this phenomenon has been described as a phonological rule that either categorically deletes the vowel or changes the [+voice] feature of the vowel to [-voice]. This analysis was supported by Sawashima (1971) and Hirose (1971)'s observation that there are two distinct EMG patterns for voiced and devoiced vowel in Japanese. Close examination of the phonetic evidence based on acoustic data, however, shows that these phonological characterizations are not tenable (Jun & Beckman 1993, 1994). In this paper, we examined the vowel devoicing phenomenon in Korean using data from ENG fiberscopic and acoustic recorders of 100 sentences produced by one Korean speaker. The results show that there is variability in the 'degree of devoicing' in both acoustic and EMG signals, and in the patterns of glottal closing and opening across different devoiced tokens. There seems to be no categorical difference between devoiced and voiced tokens, for either EMG activity events or glottal patterns. All of these observations support the notion that vowel devoicing in Korean can not be described as the result of the application of a phonological rule. Rather, devoicing seems to be a highly variable 'phonetic' process, a more or less subtle variation in the specification of such phonetic metrics as degree and timing of glottal opening, or of associated subglottal pressure or intra-oral airflow associated with concurrent tone and stricture specifications. Some of token-pair comparisons are amenable to an explanation in terms of gestural overlap and undershoot. However, the effect of gestural timing on vocal fold state seems to be a highly nonlinear function of the interaction among specifications for the relative timing of glottal adduction and abduction gestures, of the amplitudes of the overlapped gestures, of aerodynamic conditions created by concurrent oral tonal gestures, and so on. In summary, to understand devoicing, it will be necessary to examine its effect on phonetic representation of events in many parts of the vocal tracts, and at many stages of the speech chain between the motor intent and the acoustic signal that reaches the hearer's ear.

  • PDF

비회귀성 후두 신경; 수술 전 경부 CT를 통한 신경 손상의 예방 (Nonrecurrent Laryngeal Nerve; Prevention of Neural Injury by Preoperative Neck CT)

  • 김진성;소상수;최동일;양윤수;홍기환
    • 대한후두음성언어의학회지
    • /
    • 제18권1호
    • /
    • pp.67-70
    • /
    • 2007
  • Background and Objectives: The nonrecurrent laryngeal nerve(NRLN) is exceedingly rare nerve anomaly that is associated with developmentally aberrant subclavian artery. The presence of NRLN is associated with an increased risk of vocal cord palsy in thyroid surgery. The purpose of this study is to investigate its prevalence, associated vascular anomaly and necessity of recognizing its possibility for prevention of intraoperative nerve damage. Materials and Methods: Between January 2004 and December 2006, 583 thyroidectomy were performed at our hospital. Of these cases, 529 cases(90.7%) were checked preoperative neck CT. Results: Patients with preopreative neck CT, 6 cases show the retroesophageal abberant right subclavian artery that arising directly form the aortic arch. 5 cases of these 6 cases(5/6, 83.3%) and of 583 patients(5/583, 0.8%) performed thyroid surgery were identified NRLN per-operatively. All of them are identified on the right side. There were 4 women and 1 man. In all cases, there were no clinical symptoms. I case was performed only left hemithyroidectomy, so we cannot identified NRLN. No vocal cord palsy was observed. Conclusion: It is possible to predict NRLN from preoperative neck CT. When NRLN is suspected, careful, complete dissection of the nerve is always advocated. These process can reduce the operative morbidity.

  • PDF

기능적 실성증에 대한 음성치료의 효과 분석: 기초 연구 (The Effect of Voice Therapy for the Treatment of Functional Aphonia: A Preliminary Study)

  • 김노을;김준석;오재환;김동영;우주현
    • 대한후두음성언어의학회지
    • /
    • 제32권2호
    • /
    • pp.75-80
    • /
    • 2021
  • Background and Objectives Functional aphonia refers to in which by presenting whispering voice and almost producing very high-pitched tensed voices are produced. Voice therapy is the most effective treatment, but there is a lack of consensus for application of voice therapy. The purpose of this study was to examine the vocal characteristics of functional aphonia and the effect of voice therapy applied accordingly. Materials and Method From October 2019 to December 2020, 11 patients with functional aphonia were treated using voice therapy which was processing three stages such as vocal hygiene, trial therapy, and behavioral therapy. Of these, 7 patients who completed the voice evaluation before and after voice therapy was enrolled in this study. By retrospective chart review, clinical information such as sex, age, symptoms, duration, social and medical history, process of voice therapy, subjective and objective findings were analyzed. Voice parameters before and after voice therapy were compared. Results In GRBAS study, grade, rough, and asthenic, and in Consensus Auditory-Perceptual Evaluation of Voice, overall severity, roughness, pitch, and loudness were significantly improved after voice therapy. In Voice handicap index, all of the scores of total and sub-categories were significantly decreased. In objective voice analysis, jitter, cepstral peak prominence, and maximum phonation time were significantly improved. Conclusion The voice therapy was effective for the treatment of functional aphonia by restoring patient's vocalization and improving voice quality, pitch and loudness.

후두 접촉성 육아종의 치료 (Management of Laryngeal Contact Granuloma)

  • 고문희;손영익;장전엽;소윤경;정만기
    • 대한후두음성언어의학회지
    • /
    • 제19권2호
    • /
    • pp.128-132
    • /
    • 2008
  • Background: Laryngeal contact granuloma is an inflammatory hypertrophic granulation tissue arising at around the vocal process of arytenoid cartilage. Various approaches are currently used for the treatment, but a solid guideline has not been established. Objectives: We aimed to compare the each treatment modality in the hope of suggesting a guideline for the successful management of laryngeal contact granuloma. Method: Eighty-seven treatment cases of 56 patients were analyzed. Cases having recent intubation history were excluded from the study. All patients received vocal hygiene education. Proton pump inhibitors (PPI, N = 33) or H2 receptor antagonists ($H_{2}RA$, N =26) were used as a first-line treatment. Among the non-responders to $H_{2}RA$, 11 cases received PPI as a second-line therapy. Eight cases received botulinum toxin injection and 9 cases had laryngomicrosurgical removal. Results: As an initial therapy, response rate to PPI and $H_{2}RA$ was 60.6% and 38.5% respectively, which was not statistically different (p=0.091). Response rate of PPI as the second-line therapy was 36.3% (p=0.162 when compared to that of first-line PPI therapy). Response rate of Botulinum toxin injection was 75%. All cases of surgical removal recurred in a relatively short period (mean 1.9months). Conclusion: In patients having laryngeal contact granuloma, combined therapy with vocal hygiene education and PPI medication would provide more than 60% of therapeutic response. Botulinum toxin injection is highly effective even in non-responders to antireflux therapy. The only indications of surgery are to resolve diagnostic doubt or to treat acute airway compromise.

  • PDF

후두 삽관육아종 16례에 대한 임상적 고찰 (A clinical study on the 16 cases of intubation granuloma)

  • 김용신;김정은;차형근;장백암
    • 대한기관식도과학회:학술대회논문집
    • /
    • 대한기관식도과학회 1993년도 제27차 학술대회 초록집
    • /
    • pp.76-76
    • /
    • 1993
  • 기관내 삽관은 전신마취 및 기도 확보를 위해 시행되어 왔으나, 이비인후과 영역의 합병증으로는 육아종 등이 유발될 수 있는 문제점을 안고 있다. 이에 저자들은 1982년부터 1992년까지 만 10년 동안 본원 이비인후과에서 경험한 16례에서 아래와 같은 결과를 얻었다. 1. 연령 분포는 20세에서 49세까지가 84% 로써 가장 많았으며 남녀 비는 1 : 7 로써 여성에서 호발 하였다. 2. 임상증상으로 애성 12례(75 %), 후두 이물감 3례(18 %), 호흡곤란 1례(6 %)이었다. 3. 발생 부위로는 양측성 6례(37 %), 일측성 10례(63 %)중 우측 7례(70 %), 좌측 3례(30 %)이었고 발생 장소는 피열연골 성대돌기 8례(50 %), 성대 후방1/3 부위 6례(37 %), 성대 중앙 부위 2례(12 %)이었다. 4. 과거력상 삽관후 임상 증상 발현 기간은 1 개월이내 7례(44 %)로 가장 많았으며 4 개월 이상은 없었다. 5. 과거력상 수술 종류 및 빈도수는 제왕절개술이 6례(37 %)로써 가장 많았다. 6. 평균삽관 시간은 2시간 5분 이었다. 7. 튜브재질은 모두 rubber tube 이었다. 8. 수술후 재발은 1례(6 %)이었다.

  • PDF