통합 검색 | Korea Science

Fast offline transformer-based end-to-end automatic speech recognition for real-world applications

Oh, Yoo Rhee;Park, Kiyoung;Park, Jeon Gue
- ETRI Journal
- /
- 제44권3호
- /
- pp.476-490
- /
- 2022
With the recent advances in technology, automatic speech recognition (ASR) has been widely used in real-world applications. The efficiency of converting large amounts of speech into text accurately with limited resources has become more vital than ever. In this study, we propose a method to rapidly recognize a large speech database via a transformer-based end-to-end model. Transformers have improved the state-of-the-art performance in many fields. However, they are not easy to use for long sequences. In this study, various techniques to accelerate the recognition of real-world speeches are proposed and tested, including decoding via multiple-utterance-batched beam search, detecting end of speech based on a connectionist temporal classification (CTC), restricting the CTC-prefix score, and splitting long speeches into short segments. Experiments are conducted with the Librispeech dataset and the real-world Korean ASR tasks to verify the proposed methods. From the experiments, the proposed system can convert 8 h of speeches spoken at real-world meetings into text in less than 3 min with a 10.73% character error rate, which is 27.1% relatively lower than that of conventional systems.
https://doi.org/10.4218/etrij.2021-0106 인용 PDF KSCI

FUZZY LOGIC KNOWLEDGE SYSTEMS AND ARTIFICIAL NEURAL NETWORKS IN MEDICINE AND BIOLOGY

Sanchez, Elie
- 한국지능시스템학회논문지
- /
- 제1권1호
- /
- pp.9-25
- /
- 1991
This tutorial paper has been written for biologists, physicians or beginners in fuzzy sets theory and applications. This field is introduced in the framework of medical diagnosis problems. The paper describes and illustrates with practical examples, a general methodology of special interest in the processing of borderline cases, that allows a graded assignment of diagnoses to patients. A pattern of medical knowledge consists of a tableau with linguistic entries or of fuzzy propositions. Relationships between symptoms and diagnoses are interpreted as labels of fuzzy sets. It is shown how possibility measures (soft matching) can be used and combined to derive diagnoses after measurements on collected data. The concepts and methods are illustrated in a biomedical application on inflammatory protein variations. In the case of poor diagnostic classifications, it is introduced appropriate ponderations, acting on the characterizations of proteins, in order to decrease their relative influence. As a consequence, when pattern matching is achieved, the final ranking of inflammatory syndromes assigned to a given patient might change to better fit the actual classification. Defuzzification of results (i.e. diagnostic groups assigned to patients) is performed as a non fuzzy sets partition issued from a "separating power", and not as the center of gravity method commonly employed in fuzzy control. It is then introduced a model of fuzzy connectionist expert system, in which an artificial neural network is designed to build the knowledge base of an expert system, from training examples (this model can also be used for specifications of rules in fuzzy logic control). Two types of weights are associated with the connections: primary linguistic weights, interpreted as labels of fuzzy sets, and secondary numerical weights. Cell activation is computed through MIN-MAX fuzzy equations of the weights. Learning consists in finding the (numerical) weights and the network topology. This feed forward network is described and illustrated in the same biomedical domain as in the first part.
PDF

CTC를 적용한 CRNN 기반 한국어 음소인식 모델 연구 (CRNN-Based Korean Phoneme Recognition Model with CTC Algorithm)

홍윤석;기경서;권가진
- 정보처리학회논문지:소프트웨어 및 데이터공학
- /
- 제8권3호
- /
- pp.115-122
- /
- 2019
지금까지의 한국어 음소 인식에는 은닉 마르코프-가우시안 믹스쳐 모델(HMM-GMM)이나 인공신경망-HMM을 결합한 하이브리드 시스템이 주로 사용되어 왔다. 하지만 이 방법은 성능 개선 여지가 적으며, 전문가에 의해 제작된 강제정렬(force-alignment) 코퍼스 없이는 학습이 불가능하다는 단점이 있다. 이 모델의 문제로 인해 타 언어를 대상으로 한 음소 인식 연구에서는 이 단점을 보완하기 위해 순환 신경망(RNN) 계열 구조와 Connectionist Temporal Classification(CTC) 알고리즘을 결합한 신경망 기반 음소 인식 모델이 연구된 바 있다. 그러나 RNN 계열 모델을 학습시키기 위해 많은 음성 말뭉치가 필요하고 구조가 복잡해질 경우 학습이 까다로워, 정제된 말뭉치가 부족하고 기반 연구가 비교적 부족한 한국어의 경우 사용에 제약이 있었다. 이에 본 연구는 강제정렬이 불필요한 CTC 알고리즘을 도입하되, RNN에 비해 더 학습 속도가 빠르고 더 적은 말뭉치로도 학습이 가능한 합성곱 신경망(CNN)을 기반으로 한국어 음소 인식 모델을 구축하여 보고자 시도하였다. 총 2가지의 비교 실험을 통해 본 연구에서는 한국어에 존재하는 49가지의 음소를 판별하는 음소 인식기 모델을 제작하였으며, 실험 결과 최종적으로 선정된 음소 인식 모델은 CNN과 3층의 Bidirectional LSTM을 결합한 구조로, 이 모델의 최종 PER(Phoneme Error Rate)은 3.26으로 나타났다. 이는 한국어 음소 인식 분야에서 보고된 기존 선행 연구들의 PER인 10~12와 비교하면 상당한 성능 향상이라고 할 수 있다.
https://doi.org/10.3745/KTSDE.2019.8.3.115 인용 PDF KSCI HTML

국부 유사사상의 퍼지통합에 기반한 비선형사상의 식별 (Identification of Nonlinear Mapping based on Fuzzy Integration of Local Affine Mappings)

최진영;최종호
- 전자공학회논문지B
- /
- 제32B권5호
- /
- pp.812-820
- /
- 1995
This paper proposes an approach of identifying nonlinear mappings from input/output data. The approach is based on the universal approximation by the fuzzy integration of local affine mappings. A connectionist model realizing the universal approximator is suggested by using a processing unit based on both the radial basis function and the weighted sum scheme. In addition, a learning method with self-organizing capability is proposed for the identifying of nonlinear mapping relationships with the given input/output data. To show the effectiveness of our approach, the proposed model is applied to the function approximation and the prediction of Mackey-Glass chaotic time series, and the performances are compared with other approaches.
PDF

회귀신경예측 모델을 이용한 음성인식 (Speech Recognition Using Recurrent Neural Prediction Models)

류제관;나경민;임재열;성경모;안성길
- 전자공학회논문지B
- /
- 제32B권11호
- /
- pp.1489-1495
- /
- 1995
In this paper, we propose recurrent neural prediction models (RNPM), recurrent neural networks trained as a nonlinear predictor of speech, as a new connectionist model for speech recognition. RNPM modulates its mapping effectively by internal representation, and it requires no time alignment algorithm. Therefore, computational load at the recognition stage is reduced substantially compared with the well known predictive neural networks (PNN), and the size of the required memory is much smaller. And, RNPM does not suffer from the problem of deciding the time varying target function. In the speaker dependent and independent speech recognition experiments under the various conditions, the proposed model was comparable in recognition performance to the PNN, while retaining the above merits that PNN doesn't have.
PDF

지능형 교육 시스템을 위한 적응적 지식베이스 객체 모형 개발 (Development of a Adaptive Knowledge Base Object Model for Intelligent Tutoring System)

김용범;김영식
- 정보처리학회논문지B
- /
- 제13B권4호
- /
- pp.421-428
- /
- 2006
Intelligent Tutoring System(ITS)이 다양한 학습자 변인을 고려한 개별화된 학습 환경을 제공하여 영역 전문가를 대신할 효율적인 대안으로 인식되어짐에 따라, Learning Companion System(LCS)에 대한 연구도 긍정적으로 검토되어지고 있다. 하지만 LCS에서의 원활한 상호작용을 위해서는 동일한 역할을 하는 복수 LC의 결합이 필요하고, 이는 개별적 지식베이스의 확보를 선행 조건으로 요구한다. 따라서 본 연구에서는 인지구조의 연결주의적 관점을 근거로, 지식베이스 자체의 자기 학습(self learning)이 가능하고, 지식베이스 객체의 소유자에 의해 적응적으로 성장 가능한 지식베이스 객체 모형을 설계하고, 이를 검증하였다. 이 지식베이스 객체 모형은 개별적 지식베이스의 구축을 가능하게 하여, 지식베이스 객체를 이용한 적응적 ITS 개발의 기회를 제공한다.
https://doi.org/10.3745/KIPSTB.2006.13B.4.421 인용 PDF KSCI

통합적 인지 모형의 가능성 (Toward a Possibility of the Unified Model of Cognition)

이영의
- 과학기술학연구
- /
- 제1권2호
- /
- pp.399-422
- /
- 2001
인지과학에서 최근 논의되고 있는 인지 이론들은 인지에 대한 적절한 모형을 제공하지 못하고 있다. 전통적인 인공지능 이론은 추리나 문제 해결과 같은 과제에는 적절한 것처럼 보이지만 문자와 음성 인식과 같은 패턴 인식 분야에서는 여전히 비효율적이다. 연결주의는 전통적인 인공지능 이론과는 정반대의 양상을 보이고 있다. 연결주의 체계는 패턴 인식에는 강하지만 추리에는 약하다. 한편 최근에 제시된 상황화 된 행동 이론은 전통적인 인공지능과 연결주의에서 기본적으로 전제되고 있는 표상의 개념을 부정하고 실제 세계에서 직접 유래되는 지각에 바탕을 둔 모형을 제시하지만 인간의 인지를 효과적으로 설명하고 있지 못하다. 인지 모형들이 갖고 있는 이러한 한계점들을 강조하여 나는 이 글에서 인공지능, 연결주의, 상황화된 행동 이론을 각각 좌뇌 모형, 우뇌 모형, 로봇 모형이라고 부르고 그러한 한계 상황을 벗어날 수 있는 방법으로서 모형들간의 양립가능성을 이용한 통합적 인지 모형의 구축을 모색한다.
PDF

Deep CNN 기반의 한국어 음소 인식 모델 연구 (Korean Phoneme Recognition Model with Deep CNN)

홍윤석;기경서;권가진
- 한국정보처리학회:학술대회논문집
- /
- 한국정보처리학회 2018년도 춘계학술발표대회
- /
- pp.398-401
- /
- 2018
본 연구에서는 심충 합성곱 신경망(Deep CNN)과 Connectionist Temporal Classification (CTC) 알고리즘을 사용하여 강제정렬 (force-alignment)이 이루어진 코퍼스 없이도 학습이 가능한 음소 인식 모델을 제안한다. 최근 해외에서는 순환 신경망(RNN)과 CTC 알고리즘을 사용한 딥 러닝 기반의 음소 인식 모델이 활발히 연구되고 있다. 하지만 한국어 음소 인식에는 HMM-GMM 이나 인공 신경망과 HMM 을 결합한 하이브리드 시스템이 주로 사용되어 왔으며, 이 방법 은 최근의 해외 연구 사례들보다 성능 개선의 여지가 적고 전문가가 제작한 강제정렬 코퍼스 없이는 학습이 불가능하다는 단점이 있다. 또한 RNN 은 학습 데이터가 많이 필요하고 학습이 까다롭다는 단점이 있어, 코퍼스가 부족하고 기반 연구가 활발하게 이루어지지 않은 한국어의 경우 사용에 제약이 있다. 이에 본 연구에서는 강제정렬 코퍼스를 필요로 하지 않는 CTC 알고리즘을 도입함과 동시에, RNN 에 비해 더 학습 속도가 빠르고 더 적은 데이터로도 학습이 가능한 합성곱 신경망(CNN)을 사용하여 딥 러닝 모델을 구축하여 한국어 음소 인식을 수행하여 보고자 하였다. 이 모델을 통해 본 연구에서는 한국어에 존재하는 49 가지의 음소를 추출하는 세 종류의 음소 인식기를 제작하였으며, 최종적으로 선정된 음소 인식 모델의 PER(phoneme Error Rate)은 9.44 로 나타났다. 선행 연구 사례와 간접적으로 비교하였을 때, 이 결과는 제안하는 모델이 기존 연구 사례와 대등하거나 조금 더 나은 성능을 보인다고 할 수 있다.
https://doi.org/10.3745/PKIPS.y2018m05a.398 인용 PDF

Hyperparameter experiments on end-to-end automatic speech recognition

Yang, Hyungwon;Nam, Hosung
- 말소리와 음성과학
- /
- 제13권1호
- /
- pp.45-51
- /
- 2021
End-to-end (E2E) automatic speech recognition (ASR) has achieved promising performance gains with the introduced self-attention network, Transformer. However, due to training time and the number of hyperparameters, finding the optimal hyperparameter set is computationally expensive. This paper investigates the impact of hyperparameters in the Transformer network to answer two questions: which hyperparameter plays a critical role in the task performance and training speed. The Transformer network for training has two encoder and decoder networks combined with Connectionist Temporal Classification (CTC). We have trained the model with Wall Street Journal (WSJ) SI-284 and tested on devl93 and eval92. Seventeen hyperparameters were selected from the ESPnet training configuration, and varying ranges of values were used for experiments. The result shows that "num blocks" and "linear units" hyperparameters in the encoder and decoder networks reduce Word Error Rate (WER) significantly. However, performance gain is more prominent when they are altered in the encoder network. Training duration also linearly increased as "num blocks" and "linear units" hyperparameters' values grow. Based on the experimental results, we collected the optimal values from each hyperparameter and reduced the WER up to 2.9/1.9 from dev93 and eval93 respectively.
https://doi.org/10.13064/KSSS.2021.13.1.045 인용 PDF KSCI

가상 데이터와 융합 분류기에 기반한 얼굴인식 (Face Recognition based on Hybrid Classifiers with Virtual Samples)

류연식;오세영
- 전자공학회논문지CI
- /
- 제40권1호
- /
- pp.19-29
- /
- 2003
본 논문은 인위적으로 생성된 가상 학습 데이터와 융합 분류기를 이용한 얼굴인식 알고리즘을 제안한다. 특징공간에서의 최근접 특징 선택 방법과 연결주의 모델에 기반한 서로 다른 형태의 분류기를 융합하여 통합효과를 얻도록 하였다. 두 분류기는 모두 학습 데이터의 공간적인 분포에 따라 생성된 가상 학습데이터를 이용하여 학습되고 이용된다. 첫째로, 특징 공간에서의 각 정보(Angular Infnrmation) 를 이용하는 최근접특징각(the Nearest Feature Angle : NFA)을 이용하여 저장된 학습데이터와 가장 근접한 것을 찾고, 둘째로, 질의(Query) 얼굴 특징 정보를 정면얼굴 영상의 특징정보로 투영하여 얻은 정보에 기반한 분류기의 결과를 이용한다. 정면영상 특징정보로의 투영은 다층 신경망을 이용하여 정면 회상망(Frontal Recall Network)을 구현하였고, 이것을 여러 개 묶어 앙상블 네트웍으로 구성한 Ensemble 회상망(Ensemble Recall Network)을 사용하여 일반화 성능을 향상시켰다. 끝으로, 각 분류기의 결과에 따라 융합 분류기가 최종 결과를 선택하도록 하였다. 제안된 알고리즘을 6 종류의 서고 다른 학습/시험데이터 군에 적용하여 평균 96.33%의 인식률을 얻었다. 이것은 특징라인에 기반한 방법(the Nearest Feature Line) 평균 에러율의 61.2% 이며, 단일 분류기를 사용한 경우 보다 안정된 견과를 얻고 있다.
PDF KSCI

검색결과 20건 처리시간 0.02초

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

자세히 찾기

이미지 검색 (β)