통합 검색 | Korea Science

분산음성인식 환경에서 서버에서의 스케일러블 고품질 음성복원 (Scalable High-quality Speech Reconstruction in Distributed Speech Recognition Environments)

윤재삼;김홍국;강병옥
- 대한전자공학회:학술대회논문집
- /
- 대한전자공학회 2007년도 하계종합학술대회 논문집
- /
- pp.423-424
- /
- 2007
In this paper, we propose a scalable high-quality speech reconstruction method for distributed speech recognition (DSR). It is difficult to reconstruct speech of high quality with MFCCs at the DSR server. Depending on the bit-rate available by the DSR system, we can send additional information associated with speech coding to the DSR sorrel, where the bit-rate is variable from 4.8 kbit/s to 11.4 kbit/s. The experimental results show that the speech quality reproduced by the proposed method when the bit-rate is 11.4 kbit/s is comparable with that of ITU-T G.729 under both ideal channel and frame error channel conditions while the performance of DSR is maintained to that of wireline speech recognition.
PDF

외국어 학습용 어플리케이션의 음성 인식 기술 활용 현황 - 영어와 프랑스어 말하기 학습을 중심으로 - (A Study on the Utilization of Speech Recognition Technology in Foreign Language Learning Applications - Focusing on English and French Speech -)

김선희;정현훈
- 디지털콘텐츠학회 논문지
- /
- 제19권4호
- /
- pp.621-630
- /
- 2018
본 연구는 외국어 학습 어플리케이션에서의 음성 인식 기술의 활용 현황에 관한 연구로서, 외국어 말하기 교육에 적용된 음성인식 기술의 현황과 그 한계를 파악하는 것을 그 목적으로 한다. 연구 대상으로 선정된 영어와 프랑스어 학습 어플리케이션에 대하여 말하기 학습을 중심으로 살펴 본 결과, 음성 인식 기술의 활용은 학습자의 말하기 연습 환경을 만들고 말하기 평가를 기반으로 한 피드백을 줄 수 있다는 장점이 있으나, 학습자들에게 오류를 스스로 교정할 수 있는 적절한 교정 피드백을 제공하지 않는 한계를 보이고 있음을 알 수 있었다.
https://doi.org/10.9728/dcs.2018.19.4.621 인용 PDF KSCI

마이크로폰의 종류에 따른 음성인식성능의 검토 (The Validation of Speech Recognition Performance according to Microphones)

김연화;이광현;정영조;김봉완;이용주
- 대한음성학회:학술대회논문집
- /
- 대한음성학회 2003년도 5월 학술대회지
- /
- pp.183-186
- /
- 2003
Speech recognition performance depends on various factors. One of the factors is the characteristic of a microphone which is used when speech data is collected. Thus, in the present experiment speech databases for tests are created through varying types of microphones. Then, acoustic models are built based on these databases, and each of the acoustic models is assessed by the data to determine recognition performance depending on various microphones.
PDF

앤트로피 거절을 활용한 음성인식 시스템의 성능 향상 (Improvement of Speech Recognition System using Entropy Rejection)

송점동
- 정보학연구
- /
- 제2권2호
- /
- pp.139-144
- /
- 1999
본 논문은 음성인식 시스템에서 정확도를 높이기 위해 후처리 단계에서 후보 단어들의 엔트로피 정보를 이용하였다. 기존의 우도비 검출방법은 음성 데이터에 따라 음성인식 시스템의 성능이 변하고 N개의 후보단어들의 우도값이 비슷하여 오인식 발생확률이 높았다. 그러나 본 눈문에서는 각 후보 단어들의 엔트로피 값보다 인식대상 단어 외의 단어들의 엔트로피 값이 상대적으로 낮은 후보를 거절하는 후처리 방법을 사용하여 음성 데이터에 독립적이면서도 변별력을 높인 정확한 음성인식 시스템을 얻을 수 있었다. 실험 결과 본 논문에서 제안하는 엔트로피에 의한 후처리 방법은 우도비에 의한 방법보다 인식 시스템의 성능을 false alarm이 20%일 때 최대 3.6% 향상시킬 수 있었다.
PDF

음성인식 기반의 자동 프롬프터 시스템 (Auto-Scrolling Prompter System using Speech Recognition Technology)

김길연;김진우
- 대한음성학회:학술대회논문집
- /
- 대한음성학회 2006년도 춘계 학술대회 발표논문집
- /
- pp.95-98
- /
- 2006
A prompter software is used, behind the camera, to scroll the script for a TV narrator. So far it has been manually operated by an assistant, who scrolls the caption following narrator's speech. Automating this procedure using a speech recognition technology has been investigated in this project. The developed auto-scrolling software was tested in offline and online, which shows performance good enough to replace an existing prompter software. This paper describes the whole development process and concerns to be cared.
PDF

한국인을 위한 외국어 발음 교정 시스템의 개발 및 성능 평가 (Performance Evaluation of English Word Pronunciation Correction System)

김무중;김효숙;김선주;김병기;하진영;권철홍
- 대한음성학회지:말소리
- /
- 제46호
- /
- pp.87-102
- /
- 2003
In this paper, we present an English pronunciation correction system for Korean speakers and show some of experimental results on it. The aim of the system is to detect mispronounced phonemes in spoken words and to give appropriate correction comments to users. There are several English pronunciation correction systems adopting speech recognition technology, however, most of them use conventional speech recognition engines. From this reason, they could not give phoneme based correction comments to users. In our system, we build two kinds of phoneme models: standard native speaker models and Korean's error models. We also design recognition network based on phonemes to detect Koreans' common mispronunciations. We get 90% detection rate in insertion/deletion/replacement of phonemes, but we cannot get high detection rate in diphthong split and accents.
PDF

내용기반 비디오 색인 및 검색을 위한 음성인식기술 이용에 관한 연구 (A Study on the Use of Speech Recognition Technology for Content-based Video Indexing and Retrieval)

손종목;배건성;강경옥;김재곤
- 한국음향학회지
- /
- 제20권2호
- /
- pp.16-20
- /
- 2001
비디오 프로그램 색인 및 검색에 있어서 비디오 프로그램을 의미 있는 부분으로 분할하는 것, 즉 내용기반 비디오 프로그램 분할은 중요하다. 본 논문에서는 내용기반 비디오 프로그램 분할을 위해 음성인식기술을 이용하는 새로운 방법을 제안한다. 제안한 방법은 음성신호와 캡션 (Closed Caption)의 정확한 동기를 위해 음성인식 기법을 사용한다. 실험을 통하여 내용기반 비디오 프로그램 분할을 위해 제안한 방법의 가능성을 확인하였다.
PDF

훈련데이터 기반의 temporal filter를 적용한 4연숫자 전화음성 인식 (Recognition of Korean Connected Digit Telephone Speech Using the Training Data Based Temporal Filter)

정성윤;배건성
- 대한음성학회지:말소리
- /
- 제53호
- /
- pp.93-102
- /
- 2005
The performance of a speech recognition system is generally degraded in telephone environment because of distortions caused by background noise and various channel characteristics. In this paper, data-driven temporal filters are investigated to improve the performance of a specific recognition task such as telephone speech. Three different temporal filtering methods are presented with recognition results for Korean connected-digit telephone speech. Filter coefficients are derived from the cepstral domain feature vectors using the principal component analysis. According to experimental results, the proposed temporal filtering method has shown slightly better performance than the previous ones.
PDF

음성특성 학습 모델을 이용한 음성인식 시스템의 성능 향상 (Improvement of Speech Recognition System Using the Trained Model of Speech Feature)

송점동
- 정보학연구
- /
- 제3권4호
- /
- pp.1-12
- /
- 2000
음성은 특성에 따라 고음성분이 강한 음성과 저음성분이 강한 음성으로 구분할 수 있다. 그러나 이제까지 음성인식의 연구에 있어서는 이러한 특성을 고려하지 않고, 인식기를 구성함으로써 상대적으로 낮은 인식률과 인식모델을 구성할 때 많은 데이터를 필요로 하고 있다. 본 논문에서는 화자의 이러한 특성을 포만트 주파수를 이용하여 구분할 수 있는 방법을 제안하고, 화자음성의 고음과 저음특성을 반영하여 인식모델을 구성한 후 인식하는 방법을 제안한다. 한국어에서 가능한 47개의 모노폰을 이용하여 인식모델을 구성하였으며, 여성과 남성 각각 20명의 음성을 이용하여 인식모델을 학습시켰다. 포만트 주파수를 추출하여 구성한 포만트 주파수 테이불과 피치 정보값을 이용하여 음성의 특성을 구분한 후, 음성특성에 따라 학습된 인식모델을 이용하여 인식을 수행하였다. 본 논문에서 제안한 시스템을 이용하여 실험한 결과 기존의 방법보다 인식률이 향상됨을 보였다.
PDF

자동 음성 인식기를 위한 단채널 음질 향상 알고리즘의 성능 분석 (Performance Analysis of a Class of Single Channel Speech Enhancement Algorithms for Automatic Speech Recognition)

송명석;이창헌;이석필;강홍구
- The Journal of the Acoustical Society of Korea
- /
- 제29권2E호
- /
- pp.86-99
- /
- 2010
This paper analyzes the performance of various single channel speech enhancement algorithms when they are applied to automatic speech recognition (ASR) systems as a preprocessor. The functional modules of speech enhancement systems are first divided into four major modules such as a gain estimator, a noise power spectrum estimator, a priori signal to noise ratio (SNR) estimator, and a speech absence probability (SAP) estimator. We investigate the relationship between speech recognition accuracy and the roles of each module. Simulation results show that the Wiener filter outperforms other gain functions such as minimum mean square error-short time spectral amplitude (MMSE-STSA) and minimum mean square error-log spectral amplitude (MMSE-LSA) estimators when a perfect noise estimator is applied. When the performance of the noise estimator degrades, however, MMSE methods including the decision directed module to estimate a priori SNR and the SAP estimation module helps to improve the performance of the enhancement algorithm for speech recognition systems.
PDF KSCI

검색결과 523건 처리시간 0.027초

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

자세히 찾기

이미지 검색 (β)