Search | Korea Science

An acoustic Doppler-based silent speech interface technology using generative adversarial networks (생성적 적대 신경망을 이용한 음향 도플러 기반 무 음성 대화기술)

Lee, Ki-Seung
- The Journal of the Acoustical Society of Korea
- /
- v.40 no.2
- /
- pp.161-168
- /
- 2021
In this paper, a Silent Speech Interface (SSI) technology was proposed in which Doppler frequency shifts of the reflected signal were used to synthesize the speech signals when 40kHz ultrasonic signal was incident to speaker's mouth region. In SSI, the mapping rules from the features derived from non-speech signals to those from audible speech signals was constructed, the speech signals are synthesized from non-speech signals using the constructed mapping rules. The mapping rules were built by minimizing the overall errors between the estimated and true speech parameters in the conventional SSI methods. In the present study, the mapping rules were constructed so that the distribution of the estimated parameters is similar to that of the true parameters by using Generative Adversarial Networks (GAN). The experimental result using 60 Korean words showed that, both objectively and subjectively, the performance of the proposed method was superior to that of the conventional neural networks-based methods.
https://doi.org/10.7776/ASK.2021.40.2.161 인용 PDF KSCI

Performance Comparision of Channel distortion Compensation Techniques in Keyword Spotting System over the Telephone Network (전화망을 통한 핵심어 검출 시스템에서의 채널왜곡 보상벙법의 성능비교)

이교혁
- Proceedings of the Acoustical Society of Korea Conference
- /
- 1996.10a
- /
- pp.56-60
- /
- 1996
본 논문에서 핵심어 검출(Keyword spotting ) 시스템에서의 채널 왜곡에 대한 보상방법등의 성능을 비교하였다. 훈련을 음성과 인식실험용 음성은 서로 다른 환경에서 수집되었으며, 특별히 인식실험용 음성으로는 전화망을 통한 음성 데이터를 이용하였다. 전화망을 통한 음성인식에서는 채널왜곡과 부가잡음에 의해서 음성신호에 왜곡이 생기므로 이들에 대한 적적한 보상이 필요하다. 본 논문에서는 채널 왜곡보상을 위한 처리방법으로 널리 사용되고 있는 global cepstral mean substraction (GCMS), local cepstral mean subtraction(LCMS) 그리고 RASTA processing을 적용하였다. 그리고 인식성능의 개선을 위해 이들 방법을 likelihood ration scorning 에 의한 후처리 과정을 적용하였다. 인식실험결과 이들 방법 모두 채널왜곡 보상을 하지 않았을 경우와 비교하여 더 좋은 인식성능을 얻을 수 있었으며, 그 중 후처리를 적용한 LCMS 방법이 가장 우수한 성능을 나타내었다.
PDF

Traffic Management of Integrated Services using ATM Networks (ATM 망을 이용한 통합서비스의 트래픽 관리)

Kim, Hoon;Park, Jong-Dae;Nam, Sang-Shic;Park, Kwang-Chae
- Proceedings of the Korea Information Processing Society Conference
- /
- 2001.10b
- /
- pp.1477-1480
- /
- 2001
기존 통신사업자가 급변하는 통신시장에 대응하기 위한 구체적 접근방법에 초점을 맞추어 통신기술의 변화와 이에 따른 기존망을 어떻게 개선하여야만 수익성에 차질을 빚지 않을 수 있느냐가 전재 조건이 된다. 먼저 통신기술의 변화에 따른 망의 진화방향을 음성의 패킷화 실현, 망 구조의 단순화 및 통합화를 통한 운용비용의 절감, 향후 신규서비스의 수용에 용이한 방향이 있어야 한다. 본 논문에서는 ATM을 중심으로 한 차세대 교환망에서 음성과 데이터가 동일 패킷망을 사용하므로서 망 대역폭을 효율적으로 활용하는 방법과 유효 대역 사용률을 향상하는 유연한 대역관리 방법에 대해 개괄적으로 논하였으며, 이를 바탕으로 대역폭 할당 프로토콜을 분석한 수 있는 모델을 제안하고, 주어진 음성 및 데이터 트래픽의 요구와 제약을 조건으로 시스템 파라미터를 최적화하기 위해 update interval 시간과 음성과 데이터 트래픽에 예약된 슬롯의 수를 사용하였다. 분석적인 모델은 성능에 관한 트래픽 유형들의 영향뿐만 아니라 혼합 트래픽 시스템의 동적 할당 방법과 대역관리 방법을 제공한다.
PDF

Parkinson's disease diagnosis using speech signal and deep residual gated recurrent neural network (음성 신호와 심층 잔류 순환 신경망을 이용한 파킨슨병 진단)

Shin, Seung-Su;Kim, Gee Yeun;Koo, Bon Mi;Kim, Hyoung-Gook
- The Journal of the Acoustical Society of Korea
- /
- v.38 no.3
- /
- pp.308-313
- /
- 2019
Parkinson's disease, one of the three major diseases in old age, has more than 70 % of patients with speech disorders, and recently, diagnostic methods of Parkinson's disease through speech signals have been devised. In this paper, we propose a method of diagnosis of Parkinson's disease based on deep residual gated recurrent neural network using speech features. In the proposed method, the speech features for diagnosing Parkinson's disease are selected and applied to the deep residual gated recurrent neural network to classify Parkinson's disease patients. The proposed deep residual gated recurrent neural network, an algorithm combining residual learning with deep gated recurrent neural network, has a higher recognition rate than the traditional method in Parkinson's disease diagnosis.
https://doi.org/10.7776/ASK.2019.38.3.308 인용 PDF KSCI HTML

Development of the Weather Forecasting Service System Based on AIN (지능망 기반 음성인식 일기예보 서비스 시스템 개발)

Park Sung-Joon;Kim Jae-In;Koo Myoung-Wan
- 한국정보통신설비학회:학술대회논문집
- /
- 2004.08a
- /
- pp.262-265
- /
- 2004
본 논문에서는 음성인식을 이용한 일기예보 서비스 시스템을 소개한다. 이 서비스는 사용자가 지역명을 말하면 음성인식을 통해 그 지역명을 인식하여 일기예보를 들려주며, 차세대 지능망(AIN: Advanced Intelligent Network)에 구현되었다. 음성인식은 IP(Intelligent Peripheral)에서 이루어지며. 음성인식 실험 결과, 실험실과 시스템 상에서 각각 95.04%와 93.81%의 인식율을 보여 주었다.
PDF

Implementation of VoIP Service in Hybrid Fiber Coaxial Network (Hybrid Fiber Coaxial망에서 VoIP 서비스 구현)

Ju, Jae-han
- Journal of Advanced Navigation Technology
- /
- v.21 no.1
- /
- pp.113-118
- /
- 2017
As interest in mobile devices and networks has increased recently, voice over internet protocol (VoIP) service, which is a technology for transmitting voice data using an existing internet protocol (IP) network, has rapidly spread, Cheap voice call service has become possible. As the digital broadcasting service becomes popular, hybrid fiber coaxial (HFC) network technology, which uses broadband cable network through fusion of broadcasting and communication, utilizes existing communication system and network equipment to provide various new services such as interactive broadcasting service. Therefore, if UGS-AD is applied to VoCM and RTPS is applied to MTA in order to guarantee the quality of voice data in actual HFC Internet service network, it is possible to smoothly perform voice data transmission in narrow upstream band which is a problem in actual commercial HFC network We also proposed a method to improve VoIP service by improving QoS of voice data in HFC Internet service network.
https://doi.org/10.12673/jant.2017.21.1.113 인용 PDF KSCI

Performance comparison of various deep neural network architectures using Merlin toolkit for a Korean TTS system (Merlin 툴킷을 이용한 한국어 TTS 시스템의 심층 신경망 구조 성능 비교)

Hong, Junyoung;Kwon, Chulhong
- Phonetics and Speech Sciences
- /
- v.11 no.2
- /
- pp.57-64
- /
- 2019
In this paper, we construct a Korean text-to-speech system using the Merlin toolkit which is an open source system for speech synthesis. In the text-to-speech system, the HMM-based statistical parametric speech synthesis method is widely used, but it is known that the quality of synthesized speech is degraded due to limitations of the acoustic modeling scheme that includes context factors. In this paper, we propose an acoustic modeling architecture that uses deep neural network technique, which shows excellent performance in various fields. Fully connected deep feedforward neural network (DNN), recurrent neural network (RNN), gated recurrent unit (GRU), long short-term memory (LSTM), bidirectional LSTM (BLSTM) are included in the architecture. Experimental results have shown that the performance is improved by including sequence modeling in the architecture, and the architecture with LSTM or BLSTM shows the best performance. It has been also found that inclusion of delta and delta-delta components in the acoustic feature parameters is advantageous for performance improvement.
https://doi.org/10.13064/KSSS.2019.11.2.057 인용 PDF KSCI

A Study of Speech Recognition Web Services Environment for Voice Browser (Voice Browser를 위한 음성 인식 웹서비스 환경에 관한 연구)

Hong, In-Suk;Kim, Yoon-Joong
- Proceedings of the Korea Information Processing Society Conference
- /
- 2009.04a
- /
- pp.142-145
- /
- 2009
음성인터페이스 관련 표준화는 음성 대화, 음성인식/합성, 전화망 등의 접속망을 상호 분리하여 음성정보시스템 구성요소들 각각의 상호 독립적인 개발을 보장해 주며, 각 요소의 이해가 없이도 음성정보시스템을 개발할 수 있도록 함으로써 음성정보기술의 보급 및 확산에 크게 기여하고 있다. 이에 W3C에서는 Voice Browser에 대한 표준화를 현재 진행 중에 있으며 Vocie Browser WG에서 Voice Browser를 위한 SIF(Speech Interface Framework)를 제안하였다. 제안된 SIF에서 Voice Browser가 음성인식을 실행하기 위해서는 많은 자원의 소요와 부하가 생길 수 있다. 이러한 문제점을 해결하기 위해 본 논문에서는 음성인식 웹 서비스를 기존의 SIF에 추가한 새로운 형태의 SIF를 제안하고자 한다. 음성인식은 원격 시스템에서 수행하고 그 결과를 Voice Browser가 사용할 수 있도록 음성인식 웹서비스 환경을 구축하였다. 그리고, XML-SRGS 포멧의 grammar를 음성인식기가 사용하는 EBNF 포멧의 grammar로 변환시키는 변환기를 구현하였다.
https://doi.org/10.3745/PKIPS.y2009m04a.142 인용 PDF

Intelligent Peripheral 의 기능과 Signaling

신석현;권은희;홍성주
- Information and Communications Magazine
- /
- v.11 no.3
- /
- pp.66-73
- /
- 1994
IP(intelligent peripheral)은 AIN(Advanced Intelligent Network)을 구성하는 망 요소들 중의 하나로 사용자와의 상호동작을 제공하여 지능망 서비스를 보다 유연하고 다양하게 하는 지능망 지원 시스팀인 동시에, 자체적으로 새로운 서비스를 제공할 수도 있는 시스팀이다. IP는 AIN 교환기에 ISDN PRI(primary rate access interface)를 통하여 연결되어지고, SCP(Service Control Point)/Adjunct의 통제하에 음성안내, 음성인식, 테스트/음성 변환 등의 기술을 이용하여 사용자로부터 정보를 수집하고, 저장된 정보를 전달하는 기능을 수행한다. 본 고에서는 IP 시스팀의 하드웨어/소프트웨어적인 구조, 기능, 사용자와의 IP를 연결시켜주는 신호방식, IP가 서비스 제공을 위해 사용자와 어떻게 연결되는지를 프로토콜 측면에서 예를 들어 설명하고, 마지막으로 이러한 IP 개발 관련 소요기술과 활용방안을 고찰한다.
PDF

A Study on Deep Neural Network based Speech Enhancement (심화 신경망 기반의 음성 향상 기법에 관한 연구)

Lee, Moa;Chang, Joon-Hyuk
- Proceedings of the Korea Information Processing Society Conference
- /
- 2018.05a
- /
- pp.342-343
- /
- 2018
본 논문에서는, 적층형 심화 신경망 회귀 모델을 도입하여 잡음이 포함된 입력 신호의 특징벡터로부터 깨끗한 입력 신호의 특징벡터를 추정함으로써 음성 향상 성능을 개선 시켰다. 제안된 방법은 기존의 단일 심화신경망 기법 보다 음성인식 성능 향상에 더욱 효과가 있었다.
https://doi.org/10.3745/PKIPS.y2018m05a.342 인용 PDF

Search Result 874, Processing Time 0.03 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)