Search | Korea Science

An Isolated Word Recognition Using the Mellin Transform (Mellin 변환을 이용한 격리 단어 인식)

김진만;이상욱;고세문
- Journal of the Korean Institute of Telematics and Electronics
- /
- v.24 no.5
- /
- pp.905-913
- /
- 1987
This paper presents a speaker dependent isolated digit recognition algorithm using the Mellin transform. Since the Mellin transform converts a scale information into a phase information, attempts have been made to utilize this scale invariance property of the Mellin transform in order to alleviate a time-normalization procedure required for a speech recognition. It has been found that good results can be obtained by taking the Mellin transform to the features such as a ZCR, log energy, normalized autocorrelation coefficients, first predictor coefficient and normalized prediction error. We employed a difference function for evaluating a similarity between two patterns. When the proposed algorithm was tested on Korean digit words, a recognition rate of 83.3% was obtained. The recognition accuracy is not compatible with the other technique such as LPC distance however, it is believed that the Mellin transform can effectively perform the time-normalization processing for the speech recognition.
PDF

An improved automatic segmentation algorithm (자동 음성 분할 시스템의 성능 향상)

Kim Mu Jung;Kwon Chul Hong
- Proceedings of the Acoustical Society of Korea Conference
- /
- spring
- /
- pp.45-48
- /
- 2002
본 논문에서는 한국어 음성 합성기 데이터베이스 구축을 위하여 HMM을 이용하여 자동으로 음소경계를 추출하고, 음성 파라미터를 이용하여 그 결과를 보정하는 반자동 음성분할 시스템을 구현하였다. 개발된 시스템은 16KHz로 샘플링된 음성을 대상으로 삼았고, 레이블링 단위인 음소는 39개를 선정하였고, 음운현상을 고려한 확장 모노폰도 선정하였다. 그리고 언어학적 입력방식으로는 음소표기와 철자표기를 사용하였으며, 패턴 매칭 방법으로는 HMM을 이용하였다. 유성음/무성음/묵음 구간 분류에는 ZCR, Log Energy, 주파수 대역별 에너지 분포 등의 파라미터를 사용하였다. 개발된 시스템의 훈련된 음성은 정치, 경제, 사회, 문화, 날씨 등의 코퍼스를 사용하였으며, 성능평가를 위해 훈련에 사용되지 않은 문장 데이터베이스에 대해서 자동 음성 분할 실험을 수행하였다. 실험 결과, 수작업에 의해서 분할된 음소경계 위치와의 오차가 10ms 이내가 $87\%$, 30ms 이내가 $91\%$가 포함되었다.
PDF

A study on real-time implementation of speech recognition and speech control system using dSPACE board (dSPACE 보드를 이용한 음성인식 명령처리시스템 실시간 구현에 관한 연구)

김재웅;정원용
- Proceedings of the Korea Institute of Convergence Signal Processing
- /
- 2000.12a
- /
- pp.173-176
- /
- 2000
음성은 인간이 가진 가장 편리한 제어전송수단으로 이를 통한 제어는 인간에게 많은 편리함을 제공할 것이다. 본 논문에서는 다층구조 신경망(Multi-Layer Perceptron)을 이용하여 간단한 음성인식 명령처리시스템을 Matlab 상에서 구성해 보았다. 음성인식을 통한 제어의 목적을 위해 화자종속, 고립단어인식기를 목표로 설정하여 연구를 수행하였다. 음성의 시작점과 끝점을 검출하기 위해 단구간 에너지와 영교차율(ZCR)을 이용하였고 인식기의 특징파라미터로는 12차 LPC켑스트럼 계수를 사용하였다. 그리고 신경망의 출력값을 기동, 정지시에 활성화되도록 3개의 계층으로 하였고, 신경망의 뉴런의 개수를 각각 12, 12, 2으로 설정하였다. 먼저 기준음성패턴으로 학습시킨 후에 Matlab 환경하에 동작하는 dSPACE 실시간처리보드에 변환된 C프로그램을 다운로드하고, 음성을 입력하여 인식 후 dSPACE보드의 D/A컨버터의 출력단에 연결된 DC모터를 기동, 정지제어를 수행하였다. 실시간 음성인식 명령처리 시스템 구현을 통하여 원격제어와 같은 음성명령을 통한 제어가 가능함을 확인할 수 있었다.
PDF

Generating Speech feature vectors for Effective Emotional Recognition (효과적인 감정인식을 위한 음성 특징 벡터 생성)

Sim, In-woo;Han, Eui Hwan;Cha, Hyung Tai
- Proceedings of the Korea Information Processing Society Conference
- /
- 2019.05a
- /
- pp.687-690
- /
- 2019
본 논문에서는 효과적인 감정인식을 위한 효과적인 특징 벡터를 생성한다. 이를 위해서 음성 데이터 셋 RAVDESS를 이용하였으며, 그 중 neutral, calm, happy, sad 총 4가지 감정을 나타내는 음성 신호를 사용하였다. 본 논문에서는 기존에 감정인식에 사용되는 MFCC1~13차 계수와 pitch, ZCR, peakenergy 중에서 효과적인 특징을 추출하기 위해 클래스 간, 클래스 내 분산의 비를 이용하였다. 실험결과 감정인식에 사용되는 특징 벡터들 중 peakenergy, pitch, MFCC2, MFCC3, MFCC4, MFCC12, MFCC13이 효과적임을 확인하였다.
https://doi.org/10.3745/PKIPS.y2019m05a.687 인용 PDF

Variable Quad Rate ADPCM for Efficient Speech Transmission and Real Time Implementation on DSP (효율적인 음성신호의 전송을 위한 4배속 가변 변환율 ADPCM기법 및 DSP를 이용한 실시간 구현)

한경호
- Journal of the Korean Institute of Illuminating and Electrical Installation Engineers
- /
- v.18 no.1
- /
- pp.129-136
- /
- 2004
In this paper, we proposed quad variable rates ADPCM coding method for efficient speech transmission and real time porcessing is implemented on TMS320C6711-DSP. The modified ADPCM with four variable coding rates, 16[kbps], 24[kbps], 32[kbps] and 40[kbps] are used for speech window samples for good quality speech transmission at a small data bits and real time encoding and decoding is implemented using DSP. ZCR is used to identify the influence of the noise on the speech signal and to decide the rate change threshold. For noise superior signals, low coding rates are applied to minimize data bit and for noise inferior signals, high coding rates are applied to enhance the speech quality. In most speech telecommunications, silent period takes more than half of the signals, speech quality close to 40[kbps] can be obtained at comparabley low data bits and this is shown by simulation and experiments. TMS320C6711-DSK board has 128K flash memory and performance of 1333MIPS and has meets the requirements for real time implementation of proposed coding algorithm.
https://doi.org/10.5207/JIEIE.2004.18.1.129 인용 PDF KSCI

A Study on the Automatic Recognition of Korean Basic Spoken Digit Using Energy of Special Bandwidth (특정 대역 에너지를 이용한 한국어 기본 수자 음성의 백동 인식에 관한 연구)

Han, Hee;Kim, Soon-Hyob;Park, Kyu-Tae
- Journal of the Korean Institute of Telematics and Electronics
- /
- v.19 no.3
- /
- pp.5-12
- /
- 1982
Through the use of energy ratio of special bandwidths of basic vowels, recognition of Korean basic spoken digit is performed in logical combination with a zero-crossing rate and an energy parameter. In the experiments for recognition of the digits, the speech signal of spoken digits is filtered by a lowpass filter of which the cutoff frequency is 10KHz, and then sampled at 20KHz of sampling rate, In the speech signal processing, we used four FIR digital filters, and the order of filter lengths is 61, 120, 25, 25respectively. The filters are designed by using Remetz exchange algorithm.[13],[14] As a result, the recognition rate of 92% for the three speakers is obstained.
PDF

Estimation of Concrete Strength Based on Artificial Intelligence Techniques (인공지능 기법에 의한 콘크리트 강도 추정)

김세동;신동환;이영석;노승용;김성환
- The Journal of the Acoustical Society of Korea
- /
- v.18 no.7
- /
- pp.101-111
- /
- 1999
This paper presents concrete pattern recognition method to identify the strength of concrete by evidence accumulation with multiple parameters based on artificial intelligence techniques. At first, variance(VAR), zero-crossing(ZCR), mean frequency(MEANF), and autoregressive model coefficient(ARC) and linear cepstrum coefficient(LCC) are extracted as feature parameters from ultrasonic signal of concrete. Pattern recognition is carried out through the evidence accumulation procedure using distance measured with reference parameters. A fuzzy mapping function is designed to transform the distances for the application of the evidence accumulation method. Results(92% successful pattern recognition rate) are presented to support the feasibility of the suggested approach for concrete pattern recognition.
PDF

The Comparison of Sensitivity of Numerical Parameters for Quantification of Electromyographic (EMG) Signal (근전도의 정량적 분석시 사용되는 수리적 파라미터의 민감도 비교)

Kim, Jung-Yong;Jung, Myung-Chul
- Journal of Korean Institute of Industrial Engineers
- /
- v.25 no.3
- /
- pp.330-335
- /
- 1999
The goal of the study is to determine the most sensitive parameter to represent the degree of muscle force and fatigue. Various numerical parameters such as the first coefficient of Autoregressive (AR) Model, Root Mean Square (RMS), Zero Crossing Rate (ZCR), Mean Power Frequency (MPF), Median Frequency (MF) were tested in this study. Ten healthy male subjects participated in the experiment. They were asked to extend their trunk by using the right and left erector spinae muscles during a sustained isometric contraction for twenty seconds. The force levels were 15%, 30%, 45%, 60%, and 75% of Maximal Voluntary Contraction (MVC), and the order of trials was randomized. The results showed that RMS was the best parameter to measure the force level of the muscle, and that the first coefficient of AR model was relatively sensitive parameter for the fatigue measurement at less than 60% MVC condition. At the 75% MVC, however, both MPF and the first coefficient of AR Model showed the best performance in quantification of muscle fatigue. Therefore, the sensitivity of measurement can be improved by properly selecting the parameter based upon the level of force during a sustained isometric condition.
PDF

A Digital Audio Watermark Using Wavelet Transform and Masking Effect (웨이브릿과 마스킹 효과를 이용한 디지털 오디오 워터마킹)

Hwang, Won-Young;Kang, Hwan-Il;Han, Seung-Soo;Kim, Kab-Il;Kang, Hwan-Soo
- Proceedings of the IEEK Conference
- /
- 2003.11b
- /
- pp.243-246
- /
- 2003
In this paper, we propose a new digital audio watermarking technique with the wavelet transform. The watermark is embedded by eliminating unnecessary information of audio signal based on human auditory system (HAS). This algorithm is an audio watermarking method, which does not require any original audio information in watermark extraction process. In this paper, the masking effect is used for audio watermarking, that is, post-tempera] masking effect. We construct the window with the synchronization signal and we extract the best frame in the window by using the zero-crossing rate (ZCR) and the energy of the audio signal. The watermark may be extracted by using the correlation of the watermark signal and the portion of the frame. Experimental results show good robustness against MPEG1-layer3 compression and other common signal processing manipulations. All the attacks are made after the D/A/D conversion.
PDF

Classification of Korean Traditional Musical Instruments Using Feature Functions and k-nearest Neighbor Algorithm (특성함수 및 k-최근접이웃 알고리즘을 이용한 국악기 분류)

Kim Seok-Ho;Kwak Kyung-Sup;Kim Jae-Chun
- Journal of Korea Multimedia Society
- /
- v.9 no.3
- /
- pp.279-286
- /
- 2006
Classification method used in this paper is applied for the first time to Korean traditional music. Among the frequency distribution vectors, average peak value is suggested and proved effective comparing to previous classification success rate. Mean, variance, spectral centroid, average peak value and ZCR are used to classify Korean traditional musical instruments. To achieve Korean traditional instruments automatic classification, Spectral analysis is used. For the spectral domain, Various functions are introduced to extract features from the data files. k-NN classification algorithm is applied to experiments. Taegum, gayagum and violin are classified in accuracy of 94.44% which is higher than previous success rate 87%.
PDF

Search Result 59, Processing Time 0.029 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)