통합 검색 | Korea Science

특징 맵 중요도 기반 어텐션을 적용한 복소 스펙트럼 기반 음성 향상에 관한 연구 (A study on speech enhancement using complex-valued spectrum employing Feature map Dependent attention gate)

정재희;김우일
- 한국음향학회지
- /
- 제42권6호
- /
- pp.544-551
- /
- 2023
잡음 음성의 지각적 품질과 명료도 향상을 위해 활용되는 음성 향상은 크기 스펙트럼을 이용한 방법에서 크기와 위상을 같이 향상시킬 수 있는 복소 스펙트럼을 이용한 방법으로 연구되어왔다. 본 논문에서는 잡음 음성의 명료도와 품질을 더욱 향상시키기 위해 복소 스펙트럼 기반 음성 향상 시스템에 어텐션 기법을 적용하는 방안에 관해 연구를 수행하였다. 어텐션 기법은 additive attention을 기반으로 수행하며 복소 스펙트럼의 특성을 고려하여 어텐션 가중치를 계산할 수 있도록 하였다. 또한 특징 맵의 중요도를 고려하기 위해 전역 평균 풀링 연산을 같이 사용하였다. 복소 스펙트럼 기반 음성 향상은 Deep Complex U-Net(DCUNET) 모델을 기반으로 수행하였으며, additive attention은 Attention U-Net 모델에서 제안된 방법을 기반으로 연구를 수행하였다. 거실 환경의 잡음 데이터에 대해 음성 향상을 수행한 결과, 제안한 방법이 Source to Distortion Ratio(SDR), Perceptual Evaluation of Speech Quality(PESQ), Short Time Objective Intelligibility(STOI) 평가 지표에서 기준 모델보다 개선된 성능을 보였으며, 낮은 Signal-to-Noise Ratio(SNR) 조건의 다양한 배경 잡음 환경에 대해서도 일관된 성능 향상을 보였다. 이를 통해 제안한 음성 향상 시스템이 효과적으로 잡음 음성의 명료도와 품질을 향상시킬 수 있음을 보여주었다.
https://doi.org/10.7776/ASK.2023.42.6.544 인용 PDF

Simultaneous Spectral Resolution and Sensitivity Enhancement in MR spectrum: Maximum Likelihood Deconvolution Reconstruction

Jeong, Gwang-Woo;Jeong, Jenny Eunice;Kang, Heoung-Keun
- 한국자기공명학회논문지
- /
- 제15권2호
- /
- pp.157-174
- /
- 2011
Although the use of apodization functions in connection with postprocessing of a 2D NMR spectrum proves improved spectral quality, there is usually a trade-off between resolution enhancement and noise suppression due to a classical "uncertainty principle." In this study, therefore, a mathematical deconvolution technique called "Maximum Likelihood Deconvolution (MLD)" was adopted to achieve the spectral resolution and sensitivity enhancement simultaneously. The MLD technique greatly facilitates visualization and restoration of the genuine spectral information from complex 2D NMR spectra that would be problematic with the conventional apodization/FT processing. In particular, application of the MLD to the 2D-NOE spectrum would be very useful to derive the important proton connectivities, which are essential to achieve elucidating the 3D molecular structure.
https://doi.org/10.6564/JKMRS.2011.15.2.157 인용 PDF KSCI

복소 스펙트럼 기반 음성 향상의 성능 향상을 위한 time-frequency self-attention 기반 skip-connection 기법 연구 (A study on skip-connection with time-frequency self-attention for improving speech enhancement based on complex-valued spectrum)

정재희;김우일
- 한국음향학회지
- /
- 제42권2호
- /
- pp.94-101
- /
- 2023
음성 향상에서 많이 사용되는 U-Net과 같이 인코더와 디코더로 구성된 심층 신경망 모델은 skip-connection을 통해 인코더의 특징을 디코더에 연결하는 구조로 구성되어 있다. Skip-connection은 디코더에서 향상된 스펙트럼을 재구성하는데 도움을 주며 인코더를 통해 손실된 정보를 보완해줄 수 있다. 이때 skip-connection을 통해 연결되는 인코더의 특징과 디코더의 특징의 의미는 서로 다르다. 본 논문에서는 복소 스펙트럼 기반 음성 향상의 성능 향상을 위해 디코더에 연결되는 인코더의 특징을 디코더 특징의 의미에 가깝게 변환해주도록 skip-connection에 Self-Attention(SA)을 적용하는 방안을 연구하였다. SA는 시퀀스-시퀀스 문제에서 출력 시퀀스를 생성할 때, 입력 시퀀스의 가중 산술 평균을 이용하여 결정적인 부분을 집중해서 볼 수 있도록 하는 기법으로, 음성 향상 분야에서도 이를 적용함으로써 성능 향상에 효과적임을 입증하는 연구가 진행되었다. SA를 skip-connection에 적용하기 위해 인코더 특징과 디코더 특징을 이용하는 총 3가지의 방법에 대해 연구하였다. TIMIT 데이터베이스를 이용한 음성 향상 실험 결과, 제안하는 방법이 기존 skip-connection으로만 연결된 Deep Complex U-Net(DCUNET)과 비교하여 모든 성능 평가 지표에서 향상된 결과를 보였다.
https://doi.org/10.7776/ASK.2023.42.2.094 인용 PDF

퓨리에 변환을 이용한 지문영상의 개선에 관한 연구 (A study on the fingerpring enhancement using the fourier transform)

곽윤식
- 한국통신학회논문지
- /
- 제21권8호
- /
- pp.1897-1904
- /
- 1996
This study intends to extract the efficient spectrum characteristics of the fingerpriint image in the fourier domain and to apply them for image enhancement. In order to effectively acquire the spectrum characteristics of the fingerprint in the fourier domain, I set up a 1*64 window as a processing unit and, combining various kinds of the record and overlap lengths, made the power spectrum density estimate for each of those combinations. each spectrum characeristic acquired was applied to a re-synthesis process of the fingerprint image, and, through comparisons and evaluations of the resultant images, an improved gray scale image could be obtained. The validity of this algorithm could be confirmed by the comparison and evaluation fo the binary images which were grained on the established method and the one I used in this experiment.
PDF

HeMOSU-2 관측 자료를 이용한 파랑 스펙트럼 매개변수 추정 및 분석 (Estimation and Analysis of Wave Spectrum Parameter using HeMOSU-2 Observation Data)

이욱재;고동휘;김지영;조홍연
- 한국해안·해양공학회논문집
- /
- 제33권6호
- /
- pp.217-225
- /
- 2021
본 연구에서는 국내 서해안에 설치된 HeMOSU-2 기상타워에서 5 Hz 간격으로 관측한 수면변동자료를 이용하여 파랑 스펙트럼 정보를 산정하였으며, 이를 통해 파랑 매개변수를 추정하였다. 모든 유의 파고 범위에 대하여 관측 스펙트럼을 기준으로 JONSWAP 스펙트럼의 첨두증대계수(γ_opt)와 수정 BM 스펙트럼의 척도계수(α) 및 형상계수(β)를 추정하였으며, 각각의 매개변수 분포를 확인하였다. 분석 결과, JONSWAP 스펙트럼의 첨두증대계수(γ_opt)는 기존에 제안되고 있는 3.3에 비해 매우 낮은 수준인 1.27로 산정됐으며, 전체 파고 범위에서 첨두증대계수(γ_opt)의 분포는 확률질량함수와 확률밀도함수의 결합형태로 나타났다. 또한, 수정 BM 스펙트럼의 척도계수(α) 및 형상계수(β)는 기존 [0.300, -1.098]에 비해 낮은 수준인 [0.253, -1.377]로 추정됐으며, 두 매개변수간 선형 상관관계 분석 결과 β = -3.86α로 나타났다.
https://doi.org/10.9765/KSCOE.2021.33.6.217 인용 PDF KSCI

효과적인 복소 스펙트럼 기반 음성 향상을 위한 시간과 주파수 영역 손실함수 조합에 관한 연구 (A study on loss combination in time and frequency for effective speech enhancement based on complex-valued spectrum)

정재희;김우일
- 한국음향학회지
- /
- 제41권1호
- /
- pp.38-44
- /
- 2022
잡음에 오염된 음성의 명료도와 음질을 향상시키고자 음성 향상을 수행한다. 본 연구에서는 복소값 스펙트럼을 이용한 마스크기반 음성 향상에서 시간 영역 손실함수와 주파수 영역 손실함수에 따른 학습 결과를 비교하였다. 시간 영역의 음성 파형과 주파수 영역의 스펙트럼의 세부정보를 고려해 두 영역의 장점을 활용할 수 있도록 손실함수 조합에 관해 연구를 진행하였다. 시간 영역 손실함수는 Scale Invariant-Source to Noise Ratio(SI-SNR)을 이용해 계산하고, 주파수 영역 손실함수는 복소값 스펙트럼과 크기 스펙트럼을 Mean Squared Error(MSE)로 계산하여 사용하였고, sin 함수를 이용해 위상에 대한 손실함수를 계산하였다. 손실함수 조합은 시간 영역 손실함수인 SI-SNR과 각 주파수 영역 손실함수를 조합하였다. 또한 크기 값과 위상 값을 모두 고려할 수 있도록 SI-SNR과 크기 스펙트럼, 위상에 관련된 손실함수들도 조합하여 실험을 진행하였다. 음성 향상 결과는 Source-to-Distortion Ratio(SDR), Perceptual Evaluation of Speech Quality(PESQ), Short-Time Objective Intelligibility(STOI)를이용해 성능 비교 평가를 진행하였다. 음성 향상 결과를 확인해보기 위해 스펙트럼 상에서 비교를 진행하였다. TIMIT 데이터베이스를 이용한 실험 결과, 시간 영역 또는 주파수 영역 손실함수보다 SI-SNR과 크기 스펙트럼을 조합한 손실함수를 사용하여 음성 향상을 학습했을 때 가장 높은 성능을 보였다.
https://doi.org/10.7776/ASK.2022.41.1.038 인용 PDF KSCI

Noise Suppression Using Normalized Time-Frequency Bin Average and Modified Gain Function for Speech Enhancement in Nonstationary Noisy Environments

Lee, Soo-Jeong;Kim, Soon-Hyob
- The Journal of the Acoustical Society of Korea
- /
- 제27권1E호
- /
- pp.1-10
- /
- 2008
A noise suppression algorithm is proposed for nonstationary noisy environments. The proposed algorithm is different from the conventional approaches such as the spectral subtraction algorithm and the minimum statistics noise estimation algorithm in that it classifies speech and noise signals in time-frequency bins. It calculates the ratio of the variance of the noisy power spectrum in time-frequency bins to its normalized time-frequency average. If the ratio is greater than an adaptive threshold, speech is considered to be present. Our adaptive algorithm tracks the threshold and controls the trade-off between residual noise and distortion. The estimated clean speech power spectrum is obtained by a modified gain function and the updated noisy power spectrum of the time-frequency bin. This new algorithm has the advantages of simplicity and light computational load for estimating the noise. This algorithm reduces the residual noise significantly, and is superior to the conventional methods.
PDF KSCI

확산망을 이용한 음성인식 (The Speech Recognition Using the Diffusion Network)

허만택
- 한국음향학회:학술대회논문집
- /
- 한국음향학회 1996년도 영남지부 학술발표회 논문집 Acoustic Society of Korean Youngnam Chapter Symposium Proceedings
- /
- pp.70-75
- /
- 1996
In this paper, the pre-precessing method for the recognition of single vowels by use of spectrum envelope is presented , we use new method of an extrating spectrum envelope using the diffusion filter bank. We reduced the total processing time, and got higher enhancement of discrimination . By getting 88.3% of average recognition rate for single vowels of real voice through computer simulation, we confirmed it to be useful for speech recongition which use spectrum analysis for voice signal to have many frequency components.
PDF

인간의 청각 메커니즘을 적용한 웨이블렛 분석을 통한 음성 향상에 대한 연구 (A study of speech. enhancement through wavelet analysis using auditory mechanism)

이준석;길세기;홍준표;홍승홍
- 대한전자공학회:학술대회논문집
- /
- 대한전자공학회 2002년도 하계종합학술대회 논문집(4)
- /
- pp.397-400
- /
- 2002
This paper has been studied speech enhancement method in noisy environment. By mean of that we prefer human auditory mechanism which is perfect system and applied wavelet transform. Multi-resolution of wavelet transform make possible multiband spectrum analysis like human ears. This method was verified very effective way in noisy speech enhancement.
PDF

잡음환경에서 음성인식 성능향상을 위한 바이너리 마스크를 이용한 스펙트럼 향상 방법 (Method for Spectral Enhancement by Binary Mask for Speech Recognition Enhancement Under Noise Environment)

최갑근;김순협
- 한국음향학회지
- /
- 제29권7호
- /
- pp.468-474
- /
- 2010
음성인식의 실용화에 가장 저해되는 요소는 배경잡음과 채널잡음에 의한 왜곡이다. 일반적으로 배경잡음은 음성인식 시스템의 성능을 저하시키고 이로 인해 사용 장소의 제약을 받게 한다. DSR (Distributed Speech Recognition) 기반의 음성인식 역시 이와 같은 문제로 성능 향상에 어려움을 겪고 있다. 이러한 문제를 해결하기 위해 다양한 잡음제거 알고리듬이 사용되고 있으나 낮은 SNR환경에서 부정확한 잡음추정으로 발생하는 스펙트럼 손상과 잔존 잡음은 음성인식기의 인식환경과 학습 환경의 불일치를 만들게 되어 인식률을 저하시키는 원인이 된다. 본 논문에서는 이와 같은 문제를 해결하기 위해 잡음제거 알고리듬으로 MMSE-STSA 방법을 사용하였고 손상된 스펙트럼을 보상하기 위해 Ideal Binary Mask를 이용하였다. 잡음환경 (SNR 15 ~ 0 dB)에 따른 실험결과 제안된 방법을 사용했을 때 향상된 스펙트럼을 얻을 수 있었고 향상된 인식성능을 확인했다.
https://doi.org/10.7776/ASK.2010.29.7.468 인용 PDF KSCI

검색결과 221건 처리시간 0.022초

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

자세히 찾기

이미지 검색 (β)