Search | Korea Science

Intelligibility Enhancement of Multimedia Contents Using Spectral Shaping (스펙트럼 성형기법을 이용한 멀티미디어 콘텐츠의 명료도 향상)

Ji, Youna;Park, Young-cheol;Hwang, Young-su
- Journal of the Institute of Electronics and Information Engineers
- /
- v.53 no.11
- /
- pp.82-88
- /
- 2016
In this paper, we propose an intelligibility enhancement algorithm for multimedia contents using spectral shaping. The dialogue signals is essential to understand the plot of audio-visual media contents such as movie and TV. However, the non-dialogue components as like sound effects and background music often degrade the dialogue clarity. To overcome this problem, this paper tries to improves the dialogue clarity of audio soundtracks which contain important cues for the visual scenes. In the proposed method, the dialogue components are first detected by soft masker based on speech presence probability (SPP) which is widely used in speech enhancement field. Then, extracted dialogue signals are applied to the spectral shaping method. It reallocate the spectral-temporal energy of speech to enhanced the intelligibility. The total energy is maintained as unchanged via a loudness normalization process to prevent saturation. The algorithm was evaluated using the modeled and real movie soundtracks and it was shown that the proposed algorithm enhances the dialogue clarity while preserving the total audio power.
https://doi.org/10.5573/ieie.2016.53.11.082 인용 PDF KSCI

Robust Speech Enhancement Based on Soft Decision Employing Spectral Deviation (스펙트럼 변이를 이용한 Soft Decision 기반의 음성향상 기법)

Choi, Jae-Hun;Chang, Joon-Hyuk;Kim, Nam-Soo
- Journal of the Institute of Electronics Engineers of Korea SP
- /
- v.47 no.5
- /
- pp.222-228
- /
- 2010
In this paper, we propose a new approach to noise estimation incorporating spectral deviation with soft decision scheme to enhance the intelligibility of the degraded speech signal in non-stationary noisy environments. Since the conventional noise estimation technique based on soft decision scheme estimates and updates the noise power spectrum using a fixed smoothing parameter which was assumed in stationary noisy environments, it is difficult to obtain the robust estimates of noise power spectrum in non-stationary noisy environments that spectral characteristics of noise signal such as restaurant constantly change. In this paper, once we first classify the stationary noise and non-stationary noise environments based on the analysis of spectral deviation of noise signal, we adaptively estimate and update the noise power spectrum according to the classified noise types. The performances of the proposed algorithm are evaluated by ITU-T P. 862 perceptual evaluation of speech quality (PESQ) under various ambient noise environments and show better performances compared with the conventional method.
PDF KSCI

Complex nested U-Net-based speech enhancement model using a dual-branch decoder (이중 분기 디코더를 사용하는 복소 중첩 U-Net 기반 음성 향상 모델)

Seorim Hwang;Sung Wook Park;Youngcheol Park
- The Journal of the Acoustical Society of Korea
- /
- v.43 no.2
- /
- pp.253-259
- /
- 2024
This paper proposes a new speech enhancement model based on a complex nested U-Net with a dual-branch decoder. The proposed model consists of a complex nested U-Net to simultaneously estimate the magnitude and phase components of the speech signal, and the decoder has a dual-branch decoder structure that performs spectral mapping and time-frequency masking in each branch. At this time, compared to the single-branch decoder structure, the dual-branch decoder structure allows noise to be effectively removed while minimizing the loss of speech information. The experiment was conducted on the VoiceBank + DEMAND database, commonly used for speech enhancement model training, and was evaluated through various objective evaluation metrics. As a result of the experiment, the complex nested U-Net-based speech enhancement model using a dual-branch decoder increased the Perceptual Evaluation of Speech Quality (PESQ) score by about 0.13 compared to the baseline, and showed a higher objective evaluation score than recently proposed speech enhancement models.
https://doi.org/10.7776/ASK.2024.43.2.253 인용 PDF

A Spectral Compensation Method for Noise Robust Speech Recognition (잡음에 강인한 음성인식을 위한 스펙트럼 보상 방법)

Cho, Jung-Ho
- 전자공학회논문지 IE
- /
- v.49 no.2
- /
- pp.9-17
- /
- 2012
One of the problems on the application of the speech recognition system in the real world is the degradation of the performance by acoustical distortions. The most important source of acoustical distortion is the additive noise. This paper describes a spectral compensation technique based on a spectral peak enhancement scheme followed by an efficient noise subtraction scheme for noise robust speech recognition. The proposed methods emphasize the formant structure and compensate the spectral tilt of the speech spectrum while maintaining broad-bandwidth spectral components. The recognition experiments was conducted using noisy speech corrupted by white Gaussian noise, car noise, babble noise or subway noise. The new technique reduced the average error rate slightly under high SNR(Signal to Noise Ratio) environment, and significantly reduced the average error rate by 1/2 under low SNR(10 dB) environment when compared with the case of without spectral compensations.
PDF KSCI

Tradeoff between Energy-Efficiency and Spectral-Efficiency by Cooperative Rate Splitting

Yang, Chungang;Yue, Jian;Sheng, Min;Li, Jiandong
- Journal of Communications and Networks
- /
- v.16 no.2
- /
- pp.121-129
- /
- 2014
The trend of an increasing demand for a high-quality user experience, coupled with a shortage of radio resources, has necessitated more advanced wireless techniques to cooperatively achieve the required quality-of-experience enhancement. In this study, we investigate the critical problem of rate splitting in heterogeneous cellular networks, where concurrent transmission, for instance, the coordinated multipoint transmission and reception of LTE-A systems, shows promise for improvement of network-wide capacity and the user experience. Unlike most current studies, which only deal with spectral efficiency enhancement, we implement an optimal rate splitting strategy to improve both spectral efficiency and energy efficiency by exploring and exploiting cooperation diversity. First, we introduce the motivation for our proposed algorithm, and then employ the typical cooperative bargaining game to formulate the problem. Next, we derive the best response function by analyzing the dual problem of the defined primal problem. The existence and uniqueness of the proposed cooperative bargaining equilibrium are proved, and more importantly, a distributed algorithm is designed to approach the optimal unique solution under mild conditions. Finally, numerical results show a performance improvement for our proposed distributed cooperative rate splitting algorithm.
https://doi.org/10.1109/JCN.2014.000022 인용 PDF KSCI

Speech Enhancement Using Multiresolutional Signal Analysis Methods (다해상도 신호해석 방법을 이용한 음성개선)

Seok, Jong-Won;Han, Mi-Kyung;Bae, Keun-Sung
- Journal of the Korean Institute of Telematics and Electronics S
- /
- v.36S no.7
- /
- pp.134-135
- /
- 1999
This paper presents a speech enhancement method with spectral subtraction using wavelet, wavelet packet and cosine packet transforms which are known as multiresolutional signal analysis method. The performance of each method is compared with the conventional spectral subtraction method. Performance assessments based on average SNR, cepstral distance and informal subjective listening test are carried out. Experimental result demonstrate that cosine packet shows the best result in objective performance measure as well as subjective shows less musical noise than the conventional spectral subtraction method after removing the noise components.
PDF

The Determination method of Available Bandwidth for Automation of the Split-Spectrum Processing (스플릿-스펙트럼 처리의 자동화를 위한 가용대역폭의 결정방법)

Ko, Dae-Sik
- The Journal of the Acoustical Society of Korea
- /
- v.14 no.6
- /
- pp.27-31
- /
- 1995
In this paper, the determination method of available bandwidth for automation of the split-spectrum processing(SSP) has been studied. The SSP is used for the visibility enhancement of the ultrasonic signal with grain noise. Even though the SSP has proved useful in signal-to-noise ratio enhancement, its application and automation have been limited due to ambiguity in the determination of available bandwidth. Until recently, it is the usual practice to optimize the available bandwidth by trial and error. The spectral histogram is the statistical distribution of the spectral windows that is selected by the minimization algorithm with the whole band of the spectrum of the received ultrasonic signal. Since the available bandwidth can be determined adaptively using spectral histogram, this method can be used for automation of the SSP. In order to evaluate the determination technique of the available bandwidth using spectral histogram, this method is applied to experimental ultrasonic data. The experimental results show that the spectral histogram is an efficient method for determination of the available bandwidth and automation of the SSP.
PDF

Performance Evaluation of Pansharpening Algorithms for WorldView-3 Satellite Imagery

Kim, Gu Hyeok;Park, Nyung Hee;Choi, Seok Keun;Choi, Jae Wan
- Journal of the Korean Society of Surveying, Geodesy, Photogrammetry and Cartography
- /
- v.34 no.4
- /
- pp.413-423
- /
- 2016
Worldview-3 satellite sensor provides panchromatic image with high-spatial resolution and 8-band multispectral images. Therefore, an image-sharpening technique, which sharpens the spatial resolution of multispectral images by using high-spatial resolution panchromatic images, is essential for various applications of Worldview-3 images based on image interpretation and processing. The existing pansharpening algorithms tend to tradeoff between spectral distortion and spatial enhancement. In this study, we applied six pansharpening algorithms to Worldview-3 satellite imagery and assessed the quality of pansharpened images qualitatively and quantitatively. We also analyzed the effects of time lag for each multispectral band during the pansharpening process. Quantitative assessment of pansharpened images was performed by comparing ERGAS (Erreur Relative Globale Adimensionnelle de Synthèse), SAM (Spectral Angle Mapper), Q-index and sCC (spatial Correlation Coefficient) based on real data set. In experiment, quantitative results obtained by MRA (Multi-Resolution Analysis)-based algorithm were better than those by the CS (Component Substitution)-based algorithm. Nevertheless, qualitative quality of spectral information was similar to each other. In addition, images obtained by the CS-based algorithm and by division of two multispectral sensors were shaper in terms of spatial quality than those obtained by the other pansharpening algorithm. Therefore, there is a need to determine a pansharpening method for Worldview-3 images for application to remote sensing data, such as spectral and spatial information-based applications.
https://doi.org/10.7848/ksgpc.2016.34.4.413 인용 PDF KSCI KPUBS HTML

Combining deep learning-based online beamforming with spectral subtraction for speech recognition in noisy environments (잡음 환경에서의 음성인식을 위한 온라인 빔포밍과 스펙트럼 감산의 결합)

Yoon, Sung-Wook;Kwon, Oh-Wook
- The Journal of the Acoustical Society of Korea
- /
- v.40 no.5
- /
- pp.439-451
- /
- 2021
We propose a deep learning-based beamformer combined with spectral subtraction for continuous speech recognition operating in noisy environments. Conventional beamforming systems were mostly evaluated by using pre-segmented audio signals which were typically generated by mixing speech and noise continuously on a computer. However, since speech utterances are sparsely uttered along the time axis in real environments, conventional beamforming systems degrade in case when noise-only signals without speech are input. To alleviate this drawback, we combine online beamforming algorithm and spectral subtraction. We construct a Continuous Speech Enhancement (CSE) evaluation set to evaluate the online beamforming algorithm in noisy environments. The evaluation set is built by mixing sparsely-occurring speech utterances of the CHiME3 evaluation set and continuously-played CHiME3 background noise and background music of MUSDB. Using a Kaldi-based toolkit and Google web speech recognizer as a speech recognition back-end, we confirm that the proposed online beamforming algorithm with spectral subtraction shows better performance than the baseline online algorithm.
https://doi.org/10.7776/ASK.2021.40.5.439 인용 PDF KSCI

Spectral Subtraction Usnig Whitening Filter for Reducing Residual Noise (잔류잡음 감소를 위한 백색화 스펙트럼 차감법)

오태호
- Proceedings of the Acoustical Society of Korea Conference
- /
- 1998.06e
- /
- pp.411-414
- /
- 1998
음성의 음질 향상(Speech Enhancement)을 위한 여러 가지 방법 중에서 주파수 차감법(Spectral Subtraction)은 계산량이 적기 때문에 현재 실시간으로 Speech Enhancement를 할 수 있는 가장 적절한 방법이다. 그러나, 이 방법은 원래의 입력음성에 없던 새로운 잡음을 만들어내는 큰 단점이 있는데, 이를 제거하기 위해 많은 연구가 되어오고 있다. 이러한 연구의 방향은 대부분 주변프레임 또는 주변의 주파수 성분과의 평균을 통해 피크값을 무디게 해 줌으로써 새로 생긴 튀는 잡음을 감소시키는 것이다. 이런 방법은 음성자체의 정보 또한 평균이 되어버리게 하는 새로운 단점을 낳는데, 이런 현상은 무성음구간에서 특히 심각해진다. 본 논문에서는 입력음성의 LPC 분석으로 백색필터(Whitening Filter)를 구성하여 이를 통과시킨 잔류신호(Residual)를 주파수 차감하여 얻은 새로운 잔류신호를 역 필터링하여(Synthesis Filter) 개선된 음성을 얻는 방법을 제안하였다. 제안된 알고리듬은, 주파수 차감시 포만트(Formant)의 정보가 더 유지 될 수 있기 때문에 잔류잡음을 줄일 수 있다. 청취 테스트 결과 제안한 방법이 기존의 방법보다 잔류잡음을 더 줄이는 사실을 확인할 수 있었다.
PDF

Search Result 208, Processing Time 0.023 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)