• Title/Summary/Keyword: Overlap and Add

Search Result 49, Processing Time 0.027 seconds

SWAPPING NATIVE AND NON-NATIVE SPEAKERS' PROSODY USING THE PSOLA ALGORITHM

  • Yoon Kyu-Chul
    • Proceedings of the KSPS conference
    • /
    • 2006.05a
    • /
    • pp.77-81
    • /
    • 2006
  • This paper presents a technique of imposing the prosodic features of a native speaker's utterance onto the same sentence uttered by a non-native speaker. Three acoustic aspects of the prosodic features were considered: the fundamental frequency (F0) contour, segmental durations, and the intensity contour. The fundamental frequency contour and the segmental durations of the native speaker's utterance were imposed on the non-native speaker's utterance by using the PSOLA (pitch-synchronous overlap and add) algorithm [1] implemented in Praat[2]. The intensity contour transfer was also done in Praat. The technique of transferring one or more of these prosodic features was elaborated and its implications in the area of language education were discussed.

  • PDF

A Wavelet-Domain IKONOS Satellite Image Fusion Algorithm Considering the Spectrum Range of Multispectral Images (다중분광 영상의 색상별 스펙트럼 영역을 고려한 웨이블릿 변역 IKONOS 위성영상 융합 알고리즘)

  • Lee, Young-Gun;Kuk, Jung-Gap;Cho, Nam-Ik
    • Journal of Broadcast Engineering
    • /
    • v.16 no.1
    • /
    • pp.14-22
    • /
    • 2011
  • The conventional satellite image fusion methods usually add the same amount of higher frequency components extracted from the panchromatic image to all the multispectral images. However, it is noted that each of multispectral images has different amount of overlap with the panchromatic image in terms of its spectrum, and also has different intensities. Thus giving the same amount of high frequency contents to all the spectral bands does not match with this observation, which causes color distortion in the fused image. In this paper, we propose a new wavelet-domain satellite image fusion algorithm that can compensate for these differences in intensity and spectrum overlap. For the compensation of intensity differences, we first estimate the high resolution multispectral images from P, considering the relative intensity ratios. For the compensation of the amount of spectral overlap, their wavelet coefficients are appended to the conventional wavelet-domain method where the coefficients for the addition is determined by the amount of spectrum overlap. Experiments are conducted for the IKONOS satellite images whose spectrums are well known, and the results show that the proposed algorithm gives higher PSNR and correlation coefficients compared to the conventional methods.

A Study on the Subband Acoustic Echo Canceller Using Weighted Overlap-Add SSB and QMF Filter Banks (중첩가산방식의 SSB 필터뱅크와 QMF 필터뱅크를 이용한 서브밴드 음향 반향 신호 제거기에 관한 연구)

  • 차경환;심동연;김천덕
    • Journal of the Korean Institute of Telematics and Electronics S
    • /
    • v.36S no.4
    • /
    • pp.93-100
    • /
    • 1999
  • 확성회의 시스템에서 응용되는 반향신호 제거기는 긴 잔향시간을 갖는 실내 공간의 환경변화에 따라 필터 계수의 갱신에 많은 시간이 요구되어 실시간 처리에 문제점으로 지적되고 있다. 본 논문에서는 연산량 저감을 통한 실시간 처리를 위하여 중첩가산방식의 SSB(Single Side Band) 필터뱅크를 사용한 서브밴드 적응 신호처리법을 제안한다. 이 방법은 입력과 출력의 스펙트럼을 몇 개의 주파수 밴드로 분할하여, 각 밴드를 ES-NLMS(Exponential Step-Normalized Least Mean Square) 알고리즘을 이용하여 적응 처리하는 것이다. 시뮬레이션 결과 중첩가산방식의 SSB 필터뱅크가 풀밴드 보다 ERLE(Echo Return Loss Enhancement)가 1∼2㏈ 정도 작을 때 연산량이 풀밴드 보다 약95%, QMF(Quadrature Mirror Filter)필터뱅크보다 약50% 정도 감소하여 우수한 것으로 나타났다.

  • PDF

Fast Time-Scale Modification of Speech Using Nonlinear Clipping Methods

  • Jung, Ho-Young;Kim, Hyung-Soon;Lee, Sung-Joo
    • MALSORI
    • /
    • no.59
    • /
    • pp.69-87
    • /
    • 2006
  • Among the conventional time-scale modification (TSM) methods, the synchronized overlap and add (SOLA) method is widely used due to its good performance relative to computational complexity But the SOLA method remains complex due to its synchronization procedure using the normalized cross-correlation function. In this paper, we introduce a computationally efficient SOLA method utilizing 3 level center clipping method, as well as zero-crossing and level-crossing information. The result of subjective preference test indicates that the proposed method can reduce the computational complexity by over 80% compared with the conventional SOLA method without serious degradation of synthesized speech quality.

  • PDF

Real-time Voice Change System using Pitch Change (피치 변환을 사용한 실시간 음성 변환 시스템)

  • 김원구
    • Proceedings of the Korean Institute of Intelligent Systems Conference
    • /
    • 2004.04a
    • /
    • pp.466-469
    • /
    • 2004
  • In this paper, real-time voice change method using pitch change technique is proposed to change one's voice to the other voice. For this purpose, sampling rate change method using DFT (Discrete Fourier Transform) method and time scale modification method using SOLA (Synchronized Overlap and Add) method is combined to change pitch. In order to evaluate the performance of the proposed method, voice transformation experiments were conducted. Experimental results showed that original speech signal is changed to the other speech signal in which original speaker's identity is difficult to find. The system is implemented using TI TMS320C6711DSK board to verify the system runs in real time.

  • PDF

Improved Time Domain Aliasing Cancellation Filter Bank Based on the Multiple Overlap-add Structure (다중 중첩-합 구조에 기반한 개선된 시간 영역 엘리어싱 제거 필터 뱅크)

  • 유철재;김형명
    • The Journal of Korean Institute of Communications and Information Sciences
    • /
    • v.26 no.8B
    • /
    • pp.1057-1069
    • /
    • 2001
  • 오디오 부호화 시스템에 널리 쓰이는 시간 영역 엘리어싱 제거(TDAC) 필터 뱅크의 성능 개선을 위하여, 다중 중첩-합 구조를 바탕으로 한 개선된 TDAC 필터 뱅크를 제안하였다. 제안된 구조는 필터 뱅크의 분해 부분과 합성 부분 사이에서 발생하는 양자화 잡음의 효과를 줄이도록 제안되었다. 모의 실험을 통해 같은 양자화 비트 수를 사용하는 경우에 제안한 시스템이 SNR 측면에서 보다 나은 성능을 나타냄을 보였으며, 망 전송 데이터 양을 같게 한 경우에도 제안한 시스템이 더 적은 데이터 양의 블록 단위를 가질 수 있으므로 데이터 망의 혼잡 제어에 있어 보다 유리할 수 있음을 보였다.

  • PDF

Implementation of the Variable Bit Rate Vocoder Using G.729 Vocoder (G.729 음성 보코더를 이용한 가변 전송율 보코더 구현)

  • Ham MyungKyu;Bae MyungJin
    • Proceedings of the Acoustical Society of Korea Conference
    • /
    • spring
    • /
    • pp.73-76
    • /
    • 2002
  • 본 논문에서는 8kbps의 전송율을 가진 ITU G.729 보코더와 PSOLA(Pitch Synchronized Overlap -Add) 알고리즘을 적용하여 전송율을 6kbps와 4kbp까지 낮출 수 있는 가변 전송율 보코더를 구현하였다. 제안한 방법은 4kbps일 경우에 G.729의 부호화전에 PSOLA를 적용하여 피치의 주기를 반으로 줄여 부호화한다. 이렇게 부호화된 데이터는 G.729의 복호화를 거치고 다시 PSOLA를 통해 음성의 피치 주기를 2배로 늘려주어 원음성을 합성하게된다. 기존의 Bkbp의 전송율을 갖는 G.729는 음성의 크기가 반으로 줄어 부호화되므로 전송율이 4kpb로 줄어들게 된다. 실험의 평가는 MOS 테스트를 통해 수행되었으며 4kbp에서 MOS값이 3.37정도로 측정되었다. 또한 처리해야할 음성의 길이가 줄어들게 되므로 계산시간도 줄어들게 된다.

  • PDF

Algorithm for Concatenating Multiple Phonemic Units for Small Size Korean TTS Using RE-PSOLA Method

  • Bak, Il-Suh;Jo, Cheol-Woo
    • Speech Sciences
    • /
    • v.10 no.1
    • /
    • pp.85-94
    • /
    • 2003
  • In this paper an algorithm to reduce the size of Text-to-Speech database is proposed. The algorithm is based on the characteristics of Korean phonemic units. From the initial database, a reduced phoneme unit set is induced by articulatory similarity of concatenating phonemes. Speech data is read by one female announcer for 1000 phonetically balanced sentences. All the recorded speech is then segmented by phoneticians. Total size of the original speech data is about 640 MB including laryngograph signal. To synthesize wave, RE-PSOLA (Residual-Excited Pitch Synchronous Overlap and Add Method) was used. The voice quality of synthesized speech was compared with original speech in terms of spectrographic informations and objective tests. The quality of the synthesized speech is not much degraded when the size of synthesis DB was reduced from 320 MB to 82 MB.

  • PDF

Wavelet-based Pitch Detector for 2.4 kbps Harmonic-CELP Coder (2.4 kbps 하모닉-CELP 코더를 위한 웨이블렛 피치 검출기)

  • 방상운;이인성;권오주
    • The Journal of the Acoustical Society of Korea
    • /
    • v.22 no.8
    • /
    • pp.717-726
    • /
    • 2003
  • This paper presents the methods that design the Wavelet-based pitch detector for 2,4 kbps Harmonic-CELP Coder, and that achieve the effective waveform interpolation by decision window shape of the transition region, Waveform interpolation coder operates by encoding one pitch-period-sized segment, a prototype segment, of speech for each frame, generate the smooth waveform interpolation between the prototype segments for voiced frame, But, harmonic synthesis of the prototype waveforms between previous frame and current frame occur not only waveform errors but also discontinuity at frame boundary on that case of pitch halving or doubling, In addtion, in transition region since waveform interpolation coder synthesizes the excitation waveform by using overlap-add with triangularity window, therefore, Harmonic-CELP fail to model the instantaneous increasing speech and synthesis waveform linearly increases, First of all, in order to detect the precise pitch period, we use the hybrid 1st pitch detector, and increse the precision by using 2nd ACF-pitch detector, Next, in order to modify excitation window, we detect the onset, offset of frame by GCI, As the result, pitch doubling is removed and pitch error rate is decreased 5.4% in comparison with ACF, and is decreased 2,66% in comparison with wavelet detector, MOS test improve 0.13 at transition region.

Reconstruction of Overlapping Character in Thai Printed Documents

  • Nucharee Pemchaiswa;Wichian Premchaiswadi;Voravit Premratanachai;Seinosuke Narita
    • Proceedings of the IEEK Conference
    • /
    • 2000.07a
    • /
    • pp.31-34
    • /
    • 2000
  • This paper proposes a reconstruction scheme for overlapping characters in Thai printed document. Overlapping characters are characters that overlap with surrounding characters. The problem of overlapping characters is still an unsolved problem In commercially available software of Thai character recognition systems. The algorithm of reconstruction scheme is based on structural analysis of overlapping Thai printed characters. It consists of 2 steps: overlapping point determination and reconstruction of segmented characters. The overlapping point is defined as the intersection point between characters and can be determined by using templates. Then, an overlapping character is separated into segments at the intersection point. The structure of each segment may be an incomplete character and is not identical to the original one. Therefore, the reconstruction process is employed to add the incomplete part of these segments. The proposed scheme has been implemented and tested with 70 patterns of conventionally found in overlapping printed Thai characters with different typefaces and type sizes. The experimental results show that the proposed scheme can segment and reconstruct overlapping characters correctly. The proposed scheme can improve the recognition rate of commercially available software, ThaiOCR1.5 and ArnThai1.0, more than 60 percents

  • PDF