Search | Korea Science

Phoneme-Boundary-Detection and Phoneme Recognition Research using Neural Network (음소경계검출과 신경망을 이용한 음소인식 연구)

임유두;강민구;최영호
- Proceedings of the Korean Institute of Information and Commucation Sciences Conference
- /
- 1999.11a
- /
- pp.224-229
- /
- 1999
In the field of speech recognition, the research area can be classified into the following two categories: one which is concerned with the development of phoneme-level recognition system, the other with the efficiency of word-level recognition system. The resonable phoneme-level recognition system should detect the phonemic boundaries appropriately and have the improved recognition abilities all the more. The traditional LPC methods detect the phoneme boundaries using Itakura-Saito method which measures the distance between LPC of the standard phoneme data and that of the target speech frame. The MFCC methods which treat spectral transitions as the phonemic boundaries show the lack of adaptability. In this paper, we present new speech recognition system which uses auto-correlation method in the phonemic boundary detection process and the multi-layered Feed-Forward neural network in the recognition process respectively. The proposed system outperforms the traditional methods in the sense of adaptability and another advantage of the proposed system is that feature-extraction part is independent of the recognition process. The results show that frame-unit phonemic recognition system should be possibly implemented.
PDF

Cross-language Transfer of Phonological Awareness and Its Relations with Reading and Writing in Korean and English (음운인식의 언어 간 전이와 한글 및 영어의 읽기 쓰기와의 관계)

Kim, Sangmi;Cho, Jeung-Ryeul;Kim, Ji-Youn
- Korean Journal of Cognitive Science
- /
- v.26 no.2
- /
- pp.125-146
- /
- 2015
This study investigated the contribution of Korean phonological awareness to English phonological awareness and the relations of phonological awareness with reading and writing in Korean Hangul and English among Korean 5th graders. With age and vocabulary knowledge statistically controlled, Korean phonological awareness was transferred to English phonological awareness. Specifically, syllable and phoneme awareness in Korean transferred to syllable awareness in English, and Korean phoneme awareness transferred to English phoneme awareness. In addition, English phoneme awareness independently explained significant variance of reading and writing in Korean and English after controlling for age and vocabulary. Syllable awareness in Korean and English explained Hangul reading and writing, respectively. The results suggest cross-language transfer of phonological awareness that is a metalinguistic skill. Phoneme awareness is important in reading and writing in English whereas both of syllable and phoneme awareness are important in literacy of Korean.
PDF KSCI

Real-time Phoneme Recognition System Using Max Flow Matching (최대 흐름 정합을 이용한 실시간 음소인식 시스템 구현)

Lee, Sang-Yeob;Park, Seong-Won
- Journal of Korea Game Society
- /
- v.12 no.1
- /
- pp.123-132
- /
- 2012
There are many of games using smart devices. Voice recognition is can be useful way for input. In the game, voice have to be quickly recognized, at the same time it have to be manipulated promptly as well. In this study, we developed the optimized real-time phoneme recognition using max flow matching that it can be efficiently used in the game field. Firstly, voice wavelength is transformed to FFT, secondly, transformed value is made by a graph in Z plane, thirdly, data is extracted in specific area, and then data is saved in database. After all the value is recognized using weighted bipartite max flow matching. This way would be useful method in game or robot field when researchers hope to recognize the fast voice recognition.
https://doi.org/10.7583/JKGS.2012.12.1.123 인용 PDF KSCI

Typical Frame Etraction for Korean Phoneme Recognition (한국어 음소인식을 위한 기준 프레임 추출)

김범국
- Proceedings of the Acoustical Society of Korea Conference
- /
- 1994.06c
- /
- pp.121-124
- /
- 1994
음소를 인식의 기본으로 하는 한국어 음성인식 시스템을 구현하기 위한 기초 연구의 일환으로서 각 음소의 특징 가장 잘 표현하는 기준프레임 추출을 위한 연구를 수행하였다. 이를 위하여 먼저 선행 실험과 분산비 분석을 통해서 인식에 필요로한 시간 패턴의 길이를 추출한 후 이를 바탕으로 통계적 인식방법인 베이즈 결정법칙을 이용하여 시단 프레임으로부터 3프레임씩 시점을 1프레임씩 옮기면서 인식 실험을 해？여, 각 음소별 특징이 가장 풍부한 기준 프레임을 추출하였다. 그리고 이 기준 프레임을 중심으로 각 음소군별 인식 실험을 수행하여 그 결과를 시단을 기준으로 한 경우와 비교 검토하고 한국어 전 음소별로 확장하여 인식 실험을 실시하였다. 이 실험 결과 모음의 경우 시단으로부터 5프레임, 파열음은 시단에서부터 5프레임사이, 마찰음은 3프레임에서부터 10프레임까지, 파찰음은 5프레임까지, 비음과 유음의 경우 초성은 시단 프레임에서 6프레임, 종성은 종단으로부터 전 4프레임 구간이 인식률이 높게 나타나 이 부분의 특징이 인식에 가장 유효함을 알 수 있었다.
PDF

Phoneme-based Recognition of Korean Speech Using HMM(Hidden Markov Model) and Genetic Algorithm (HMM과 GA를 이용한 한국어 음성의 음소단위 인식)

박준하;조성원
- Proceedings of the Korean Institute of Intelligent Systems Conference
- /
- 1997.10a
- /
- pp.291-295
- /
- 1997
현재에 주로 개발되어 상용화가 시작되고 있는 음성인식 시스템의 대부분은 단어인식을 기분으로 하는 시스템으로 적용 단어수를 늘려줌으로서 인식범위를 늘일 수 있으나, 그에 따라 검색해야하는 단어수가 늘어남으로서 전체적인 시스템의 속도 및 성능이 저하되는 경향이 있다. 이러한 단점의 극복을 위하여 본 논문에서는 HMM(Hidden Markov Model)과 GA(Genetic Algorithm)를 이용한 한국어 음성의 음소단위 인식 시스템을 구현하였다. 음성 특징으로는 LPC Cepstrum 계수를 사용하였으며, 인식시는 인식대상이 되는 단어에 대하여 GA(Genetic Algorithm)을 통하여 각 음소를 분리하고, 음소단위로 학습된 HMM 파라미터를 적용하여 인식함으로써 각각의 음소별 가능하도록 하는 방법을 제안하였다.
PDF

A Study on the Phoneme Recognition using RBFN (RBFN을 이용한 음소인식에 관한 연구)

안종영
- Proceedings of the Acoustical Society of Korea Conference
- /
- 1995.06a
- /
- pp.88-91
- /
- 1995
개층형 신경망은 교사신호들의 학습으로 원하는 입출력간의 매핑을 할 수 있으므로 패턴분류를 위해 사용되어왔다. 본 논문은 계층형 신경망의 일종인 RBFN 중 GPFN 과 PNN으로 한국어 음소인식을 수행하였다. RBFN 의 구조는 계층형 신경망과 유사하나 차이점으로는 은닉층에서 시그모이드 함수, 참조벡터 및 학습알고리듬의 선택이 다르다. 특히 PNN 의 시그모이드 함수는 지수를 포함한 함수들로 대체되며 학습없이 패턴을 분류하므로 계산시간이 빠르게 수행된다. 본 실험에서는 한국어 단음절에서 모음과 자음을 추출하여 음소인식을 수행하였다. 실험 결과 학습과 평가데이타에 의한 인식률은 계층형 신경망과 비교하여 향상 되었으며, Hybrid 구성에 의한 실험에서도 항상된 인식률을 얻을 수 있었다.
PDF

Korean Phonological Viseme for Lip Synch Based on Phoneme Recognition (음소인식 기반의 립싱크 구현을 위한 한국어 음운학적 Viseme의 제안)

Joo Heeyeol;Kang Sunmee;Ko Hanseok
- Proceedings of the Acoustical Society of Korea Conference
- /
- spring
- /
- pp.70-73
- /
- 1999
본 논문에서는 한국어에 대한 실시간 음소 인식을 통한 Lip Synch 구현에 필수요소인 Viseme(Visual Phoneme)을 한국어의 음운학적 접근 방법을 통해 제시하고, Lip Synch에서 입술의 모양에 결정적인 영향을 미치는 모음에 대한 모음 인식 실험 및 결과 분석을 한다.모음인식 실험에서는 한국어 음소 51개 각각에 대해 3개의 State로 이루어진 CHMM (Continilous Hidden Makov Model)으로 모델링하고, 각각의 음소가 병렬로 연결되어진 음소네트워크를 사용한다. 입력된 음성은 12차 MFCC로 특징을 추출하고, Viterbi 알고리즘을 인식 알고리즘으로 사용했으며, 인식과정에서 Bigrim 문법과 유사한 구조의 음소배열 규칙을 사용해서 인식률과 인식 속도를 향상시켰다.
PDF

A Study on the Analysis and Recognition of Korean Speech Signal using the Phoneme (음소를 이용한 한국어 음성 신호의 분석과 인식에 관한 연구)

Kim Y. I.;Hwang Y. S.;Youn D. H.;Cha I. W.
- The Journal of the Acoustical Society of Korea
- /
- v.8 no.5
- /
- pp.70-77
- /
- 1989
In this paper, Korean language recognition using the phoneme is studied. The experiment is carried out by dividing 545 isolated words into phonemes. Using linear prediction coefficients the recognition rate of consonants, vowels, and end-consonants are $87.3(\%), 91.0(\%), 91.7(\%)$, respectively. Recognition rate of isolated words combined with the phonemes is $71.4(\%)$. Itakura-saito distortion measure is used to phoneme segmentation and phoneme recognition.
PDF

A Parallel Speech Recognition System based on Hidden Markov Model (은닉 마코프 모델 기반 병렬음성인식 시스템)

Jeong, Sang-Hwa;Park, Min-Uk
- Journal of KIISE:Computer Systems and Theory
- /
- v.27 no.12
- /
- pp.951-959
- /
- 2000
본 논문의 병렬음성인식 모델은 연속 은닉 마코프 모델(HMM; hidden Markov model)에 기반한 병렬 음소인식모듈과 계층구조의 지식베이스에 기반한 병렬 문장인식모듈로 구성된다. 병렬 음소인식 모듈은 수천개의 HMM을 병렬 프로세서에 분산시킨 수, 할당된 HMM에 대한 출력확률 계산과 Viterbi 알고리즘을 담당한다. 지식베이스 기반 병렬 문장인식모듈은 음소모듈에서 공급되는 음소열과 지안하는 병렬 음성인식 알고리즘은 분산메모리 MIMD 구조의 다중 트랜스퓨터와 Parsytec CC 상에 구현되었다. 실험결과, 병렬 음소인식모듈을 통한 실행시간 향상과 병렬 문장인식모듈을 통한 인식률 향상을 얻을 수 있었으며 병렬 음성인식 시스템의 실시간 구현 가능성을 확인하였다.
PDF

A Study on PLU (Phone-Likely Unit) for Korean Continuous Speech Recognition (강건한 한국어 연속음성인식을 위한 유사음소단일에 대한 연구)

Seo Jun-Bae;Kim Joo-Gon;Kim Min-Jung;Jung Ho-Youl;Chung Hyun-Yeol
- Proceedings of the Acoustical Society of Korea Conference
- /
- spring
- /
- pp.37-40
- /
- 2004
본 논문은 한국어 연속음성인식에 효율적인 문맥의존 음향모델 수에 대한 연구로써 유사음소단위 수에 따른 인식 성능을 비교, 평가하였다. 기존에 본연구실에서는 48음소를 기본인식단위로 이용하고 있으나 연속음성인식의 경우 문맥종속모델이 사용되고 문맥종속모델은 변이 음을 고려한 음소가 이미 포함되어 있어 이를 고려하면 기본 음소를 줄이므로서 계산량의 감소와 인식 성능 향상을 기대할 수 있을 것으로 생각된다. 따라서 , 본 논문에서는 기존의 48음소와 이를 39음소로 줄여 인식실험에 사용하여 그 성능을 비교 평가하기로 하였다. 이를 위하여 다양한 태스크의 데이터베이스를 통합하여 부족한 문맥요소들을 확장한 후 인식실험을 수행하였다. 실험결과 변이음의 개수를 줄이면서도 인식 성능저하가 없음을 확인할 수 있었으며 연속 음성의 경우 39음소를 이용한 경우가 $10\%$정도의 향상된 인식성능을 얻을 수 있음을 확인할 수 있었다.
PDF

Search Result 302, Processing Time 0.025 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)