Search | Korea Science

Intra-and Inter-frame Features for Automatic Speech Recognition

Lee, Sung Joo;Kang, Byung Ok;Chung, Hoon;Lee, Yunkeun
- ETRI Journal
- /
- v.36 no.3
- /
- pp.514-517
- /
- 2014
In this paper, alternative dynamic features for speech recognition are proposed. The goal of this work is to improve speech recognition accuracy by deriving the representation of distinctive dynamic characteristics from a speech spectrum. This work was inspired by two temporal dynamics of a speech signal. One is the highly non-stationary nature of speech, and the other is the inter-frame change of a speech spectrum. We adopt the use of a sub-frame spectrum analyzer to capture very rapid spectral changes within a speech analysis frame. In addition, we attempt to measure spectral fluctuations of a more complex manner as opposed to traditional dynamic features such as delta or double-delta. To evaluate the proposed features, speech recognition tests over smartphone environments were conducted. The experimental results show that the feature streams simply combined with the proposed features are effective for an improvement in the recognition accuracy of a hidden Markov model-based speech recognizer.
https://doi.org/10.4218/etrij.14.0213.0181 인용 PDF KSCI KPUBS

Performance Improvement of Speaker Recognition System Using Genetic Algorithm (유전자 알고리즘을 이용한 화자인식 시스템 성능 향상)

문인섭;김종교
- The Journal of the Acoustical Society of Korea
- /
- v.19 no.8
- /
- pp.63-67
- /
- 2000
This paper deals with text-prompt speaker recognition based on dynamic time warping (DTW). The Genetic Algorithm was applied to the creation of reference patterns for suitable reflection of the speaker characteristics, one of the most important determinants in the fields of speaker recognition. In order to overcome the weakness of text-dependent and text-independent speaker recognition, the text-prompt type was suggested. Performed speaker identification and verification in close and open set respectively, hence the Genetic algorithm-based reference patterns had been proven to have better performance in both recognition rate and speed than that of conventional reference patterns.
PDF

Improvement of Recognition Performance for Limabeam Algorithm by using MLLR Adaptation

Nguyen, Dinh Cuong;Choi, Suk-Nam;Chung, Hyun-Yeol
- IEMEK Journal of Embedded Systems and Applications
- /
- v.8 no.4
- /
- pp.219-225
- /
- 2013
This paper presents a method using Maximum-Likelihood Linear Regression (MLLR) adaptation to improve recognition performance of Limabeam algorithm for speech recognition using microphone array. From our investigation on Limabeam algorithm, we can see that the performance of filtering optimization depends strongly on the supporting optimal state sequence and this sequence is created by using Viterbi algorithm trained with HMM model. So we propose an approach using MLLR adaptation for the recognition of speech uttered in a new environment to obtain better optimal state sequence that support for the filtering parameters' optimal step. Experimental results show that the system embedded with MLLR adaptation presents the word correct recognition rate 2% higher than that of original calibrate Limabeam and also present 7% higher than that of Delay and Sum algorithm. The best recognition accuracy of 89.4% is obtained when we use 4 microphones with 5 utterances for adaptation.
https://doi.org/10.14372/IEMEK.2013.8.4.219 인용 PDF KSCI

On a Study of the Improvement of Speaker Recognition with Characteristics of High Order Reflection Coefficients (고차 반사계수 특성을 이용한 화자인식의 성능 향상에 관한 연구)

이윤주;오세영;함명규;배명진
- Proceedings of the IEEK Conference
- /
- 1999.06a
- /
- pp.667-670
- /
- 1999
As the number of reference patterns increase in the text dependant speaker recognition, the recognition performance of the system degrades. So, if reference patterns were decreased the high recognition rate can be obtained. It’s because the speaker recognition can obtain the high discrimination. In this paper, to decrease the number of reference patterns, we choose candidate reference patterns to perform pattern matching with test pattern by high order component of the reflection coefficients of the uttered speech signal Consequently the total recognition rate of the proposed method is about 2% higher than that of the conventional method.
PDF

The Performance Improvement of Speech Recognition System based on Stochastic Distance Measure

Jeon, B.S.;Lee, D.J.;Song, C.K.;Lee, S.H.;Ryu, J.W.
- International Journal of Fuzzy Logic and Intelligent Systems
- /
- v.4 no.2
- /
- pp.254-258
- /
- 2004
In this paper, we propose a robust speech recognition system under noisy environments. Since the presence of noise severely degrades the performance of speech recognition system, it is important to design the robust speech recognition method against noise. The proposed method adopts a new distance measure technique based on stochastic probability instead of conventional method using minimum error. For evaluating the performance of the proposed method, we compared it with conventional distance measure for the 10-isolated Korean digits with car noise. Here, the proposed method showed better recognition rate than conventional distance measure for the various car noisy environments.
https://doi.org/10.5391/IJFIS.2004.4.2.254 인용 PDF KSCI

Improvement of Bit Recognition Rate for Color QR Codes By Multiplexing Color and Pattern Information (색 및 패턴 정보 다중화를 이용한 칼라 QR코드의 비트 인식률 개선)

Kim, Jin-Soo
- Journal of Korea Multimedia Society
- /
- v.24 no.8
- /
- pp.1012-1019
- /
- 2021
Currently, since the black-white QR (Quick Response) codes have limited storage capacity, color QR codes have been actively being studied. By multiplexing 3 colors, the color QR codes can allow the code capacity to be increased by three times, however, the color multiplexing brings about the possibility of crosstalk and noises in the acquisition process of the final image, incurring the decrease of bit-recognition rate. In order to improve the bit recognition rate, while keeping the storage capacity high, this paper proposes a new type of color QR code which uses the pattern information as well as the color information, and then analyzes how to increase the bit recognition rate. For this aim, the paper presents an efficient system which extracts embedded information from color QR code and then, through practical experiments, it is shown that the proposed color QR codes improves the bit recognition rate and are useful for commercial applications, compared to the conventional color codes.
https://doi.org/10.9717/kmms.2021.24.8.1012 인용 PDF KSCI HTML

A Study on the Performance Analysis of Entity Name Recognition Techniques Using Korean Patent Literature

Gim, Jangwon
- Journal of Advanced Information Technology and Convergence
- /
- v.10 no.2
- /
- pp.139-151
- /
- 2020
Entity name recognition is a part of information extraction that extracts entity names from documents and classifies the types of extracted entity names. Entity name recognition technologies are widely used in natural language processing, such as information retrieval, machine translation, and query response systems. Various deep learning-based models exist to improve entity name recognition performance, but studies that compared and analyzed these models on Korean data are insufficient. In this paper, we compare and analyze the performance of CRF, LSTM-CRF, BiLSTM-CRF, and BERT, which are actively used to identify entity names using Korean data. Also, we compare and evaluate whether embedding models, which are variously used in recent natural language processing tasks, can affect the entity name recognition model's performance improvement. As a result of experiments on patent data and Korean corpus, it was confirmed that the BiLSTM-CRF using FastText method showed the highest performance.
https://doi.org/10.14801/JAITC.2020.10.2.139 인용

KMSAV: Korean multi-speaker spontaneous audiovisual dataset

Kiyoung Park;Changhan Oh;Sunghee Dong
- ETRI Journal
- /
- v.46 no.1
- /
- pp.71-81
- /
- 2024
Recent advances in deep learning for speech and visual recognition have accelerated the development of multimodal speech recognition, yielding many innovative results. We introduce a Korean audiovisual speech recognition corpus. This dataset comprises approximately 150 h of manually transcribed and annotated audiovisual data supplemented with additional 2000 h of untranscribed videos collected from YouTube under the Creative Commons License. The dataset is intended to be freely accessible for unrestricted research purposes. Along with the corpus, we propose an open-source framework for automatic speech recognition (ASR) and audiovisual speech recognition (AVSR). We validate the effectiveness of the corpus with evaluations using state-of-the-art ASR and AVSR techniques, capitalizing on both pretrained models and fine-tuning processes. After fine-tuning, ASR and AVSR achieve character error rates of 11.1% and 18.9%, respectively. This error difference highlights the need for improvement in AVSR techniques. We expect that our corpus will be an instrumental resource to support improvements in AVSR.
https://doi.org/10.4218/etrij.2023-0352 인용 PDF

A Study on Convergence Development Direction of Gesture Recognition Game (동작 인식 게임의 융합 발전 방향)

Lee, MyounJae
- Journal of the Korea Convergence Society
- /
- v.5 no.4
- /
- pp.1-7
- /
- 2014
Gesture recognition provides the ease and immediacy to users in the processing technique for recognizing the gesture. Because of these benefits, gesture recognition technology has been applied and fused in many areas, such as the military, health care, education. In particular, the gesture recognition in game field since it can provide users to play games similar to the actual gesture, it being fused with many areas such as medical, military, and education. This paper is to discuss the future convergence direction of motion recognition games based on this background. In this paper, it looks at technology status and the game of gesture recognition, describe the problem and improvement of gesture recognition game. This paper can help improving the competitiveness of domestic convergence on gesture recognition game.
https://doi.org/10.15207/JKCS.2014.5.4.001 인용 PDF KSCI

The Vocabulary Recognition Optimize using Acoustic and Lexical Search (음향학적 및 언어적 탐색을 이용한 어휘 인식 최적화)

Ahn, Chan-Shik;Oh, Sang-Yeob
- Journal of Korea Multimedia Society
- /
- v.13 no.4
- /
- pp.496-503
- /
- 2010
Speech recognition system is developed of standalone, In case of a mobile terminal using that low recognition rate represent because of limitation of memory size and audio compression. This study suggest vocabulary recognition highest performance improvement system for separate acoustic search and lexical search. Acoustic search is carry out in mobile terminal, lexical search is carry out in server processing system. feature vector of speech signal extract using GMM a phoneme execution, recognition a phoneme list transmission server using Lexical Tree Search algorithm lexical search recognition execution. System performance as a result of represent vocabulary dependence recognition rate of 98.01%, vocabulary independence recognition rate of 97.71%, represent recognition speed of 1.58 second.
PDF KSCI

Search Result 1,491, Processing Time 0.025 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)