통합 검색 | Korea Science

Line Spectral Frequency와 음성신호의 주파수 분포에 관한 연구 (A Study on the Relation Between the LSF's and Spectral Distribution of Speech Signals)

이동수;김영화
- 대한전자공학회논문지
- /
- 제25권4호
- /
- pp.430-436
- /
- 1988
LSF(Line Spectral Frequency) derived from LPC has known as a very useful transmission parameter of speech signals, for it has a good linear interpolation characteristics and a low spectrum distortion at low bit rates coding. This paper presents that it is possible to extract directly the formant frequencies of speech signals from LSF parameter without application of FFT algorithm by comparing the distribution of LSF parameter with the frequency distribution of analysis filter. This paper suggests the advanced algorithm that results in improving the speed of convergence at analytic solution method. Also, for the flexibility of parameters, the process that transforms from LSF to LPC is presented.
PDF

Multi-frame AR model을 이용한 LPC 계수 양자화 (Quantization of LPC Coefficients Using a Multi-frame AR-model)

정원진;김무영
- 한국음향학회지
- /
- 제31권2호
- /
- pp.93-99
- /
- 2012
음성코딩 시 성도는 Linear Predictive Coding (LPC) 계수를 이용해서 모델링 한다. 일반적으로 LPC 계수는 양자화와 선형보간 관점에서 유리한 Line Spectral Frequency (LSF) 파라미터로 변경하여 사용한다. 10차 이상의 다차원 LSF 데이터를 벡터 양자화를 이용하여 직접 코딩하게 되면 벡터 내 상관관계 (intra-frame correlation)를 모두 이용할 수 있으므로 rate-distortion 관점에서는 높은 효율을 기대할 수 있다. 하지만, 계산량과 메모리 요구량이 높아져서 실제 코딩 시스템에서는 사용할 수 없게 되므로, 차원을 나누어 압축하는 Split Vector Quantization (SVQ)이 이용된다. 또한, LSF 데이터는 과거 벡터와의 벡터 간 상관관계 (inter-frame correlation)가 높으므로, 이를 이용한 Predictive Split Vector Quantization (PSVQ)이 사용되고 있다. PSVQ는 SVQ 보다 높은 rate-distortion 성능을 보인다. 본 논문에서는 음성 저장 장치를 위한 최적의 PSVQ를 구현하기 위해서 다수의 과거 프레임 정보와의 벡터 간상관관계 (inter-frame correlation)를 고려한 Multi-Frame AR-model 기반 SVQ (MF-AR-SVQ)를 제안하였다. 기존 PSVQ와 비교해 보았을 때, MF-AR-SVQ는 계산량과 메모리 요구량의 큰 증가 없이, 평균 spectral distortion 관점에서 약 1비트의 성능 향상을 보였다.
https://doi.org/10.7776/ASK.2012.31.2.093 인용 PDF KSCI

광대역 음성 부호화기용 선 스펙트럼 주파수 계수 양자화기 설계 (Design of the LSF Parameter Quantizer for the Wideband Speech Codec)

지상현;강상원;윤병식
- 한국음향학회지
- /
- 제20권4호
- /
- pp.29-34
- /
- 2001
본 논문에서는 고품질 음성 서비스를 가능하게 하는 광대역 음성 부호화기의 선 스펙트럼 주파수 (line spectral frequency: ISF) 계수 양자화기를 설계하였다. 광대역 음성 부호화기를 위한 효율적인 LSF 계수 양자화기를 설계하기 위하여, 인접 프레임간의 상관도를 이용하였으며, 각 해당 프레임의 ISF 계수에 대한 양자화를 인접 프레임간 상관도가 높은 프레임과 상관도가 낮은 프레임으로 나누어 독립적으로 수행하였다. 인접 프레임간 상관도가 높은 프레임의 LSF계수 양자화를 위하여 예측 피라미드형 벡터 양자화기 (predictive pyramid vector quantizer: PPVQ)를 사용하여 양자화하였고, 상관도가 낮은 프레임의 LSF 계수는 피라미드형 벡터 양자화기 (PVQ)를 사용하여 양자화 하였다. PPVQ에서 예측기로 1차 AR 예측기를 사용하였다. 광대역 음성 부호화기를 위해 본 논문에서 설계된 UF 계수양자화기를 평균스펙트럼 왜곡(spectral distortion: SD) 성능 관점에서 실험한 결과, LSF계수 양자화에 할당된 비트가 프레임당 40비트일 때, 평균 SD값이 1 dB 내외이고, 2 dB 이상 및 4 dB 이상 outlier가 각각 3.87%및 0.01%인 transparent한 성능을 얻을 수 있었다.
PDF

포만트 공간에서의 주파수 변환을 이용한 이중 언어 음성 변환 연구 (Bilingual Voice Conversion Using Frequency Warping on Formant Space)

채의근;윤영선;정진만;은성배
- 말소리와 음성과학
- /
- 제6권4호
- /
- pp.133-139
- /
- 2014
This paper describes several approaches to transform a speaker's individuality to another's individuality using frequency warping between bilingual formant frequencies on different language environments. The proposed methods are simple and intuitive voice conversion algorithms that do not use training data between different languages. The approaches find the warping function from source speaker's frequency to target speaker's frequency on formant space. The formant space comprises four representative monophthongs for each language. The warping functions can be represented by piecewise linear equations, inverse matrix. The used features are pure frequency components including magnitudes, phases, and line spectral frequencies (LSF). The experiments show that the LSF-based voice conversion methods give better performance than other methods.
https://doi.org/10.13064/KSSS.2014.6.4.133 인용 PDF KSCI

선 스펙트럼 주파수의 청각 적응 부호화 (Perceptual and Adaptive Quantization of Line Spectral Frequency Parameters)

한우진;김은경;오영환
- 한국음향학회지
- /
- 제19권8호
- /
- pp.68-77
- /
- 2000
선 스펙트럼 주파수를 양자화하기 위한 대부분의 방법들이 가중 유클리드 거리에 기반하고 있는 반면, 본 논문에서는 청각 마스킹 효과에 기반한 에러 척도를 사용하여 선 스펙트럼 주파수를 효과적으로 양자화하는 방법을 제안하였다. 제안한 방법에서는 noise-to-mask ratio (NMR)를 선 스펙트럼 주파수의 양자화에 적합하도록 변형한 새로운 에러 척도를 유도하고, 이를 사용하여 선 스펙트럼 주파수를 양자화한다. 한편, 본 논문에서는 양자화하고자 하는 음성 프레임이 갖는 청각적인 특성을 고려하여 동적으로 비트를 할당하는 적응 양자화 알고리즘을 제안하였다. 성능 평가를 위해서 11948 프레임의 테스트 자료를 기존의 방법과 제안한 방법으로 각자 양자화하고 perceptually transparent frame의 비운 및 이때의 평균 비트율을 비교한 결과, 기존의 방법이 1800 bps의 비트율에서 89.9%의 perceptually transparent frame을 얻은 데 비해, 제안한 방법은 770 bps의 평균 비트율에서 95.5%의 perceptually transparent frame을 얻음으로써 제안한 방법이 효과적임을 보였다.
PDF

Block Constrained Trellis Coded Vector Quantization of LSF Parameters for Wideband Speech Codecs

Park, Jung-Eun;Kang, Sang-Won
- ETRI Journal
- /
- 제30권5호
- /
- pp.738-740
- /
- 2008
In this paper, block constrained trellis coded vector quantization (BC-TCVQ) is presented for quantizing the line spectrum frequency parameters of the wideband speech codec. Both a predictive structure and a safety-net concept are combined into BC-TCVQ to develop the predictive BC-TCVQ. The performance of this quantization is compared with that of the linear predictive coding vector quantizer used in the AMRWB codec, demonstrating reductions in spectral distortion.
PDF

G.723.1 음성 부호화기의 LSE 계수 양자화를 위한 고속화 알고리즘 연구 (A study on a fast algorithm for the LSP coefficient quantization of G. 723.1 speech codec)

송창용;성호상;강상원;성유나
- 한국음향학회:학술대회논문집
- /
- 한국음향학회 2000년도 하계학술발표대회 논문집 제19권 1호
- /
- pp.153-156
- /
- 2000
본 논문에서는 멀티미디어 서비스들 중에서 음성 또는 오디오 신호를 저속으로 압축할 때 사용되는 G.723.1 부호화기의 line spectral frequency(LSF) 계수 양자화 방식을 고속으로 처리하는 알고리즘을 제안하였다. 제안된 고속탐색 방법은 LSF 계수의 순서성질을 이용하여 코드북의 탐색 범위를 줄임으로써 계산량을 크게 감소시킨다. 제안된 고속탐색 방법을 predictive split VQ(PSVQ) 구조를 갖는 G.723.1 에 적용한 결과 spectral distortion(SD) 성능 감쇄 및 추가적인 메모리 증가 없이 최적 코드벡터를 찾기 위한 코드북 탐색 과정에서 코드북의 평균 탐색 범위가 $20.1\%$ 감소했으며, 이는 additions, subtractions, multiplies 및 comparisons 수가 각각 $19.1\%$, $20.1\%$, $19.4\%$ 및 $12.2\% 감소하는 결과를 얻었다.
PDF

VoIP 손실 환경에 강인한 저지연 LSF FEC 기법 (Low-Delay LSF FEC Technique Robust in Lossy VoIP Environment)

양해용;이경훈;황인호
- 대한전자공학회논문지SP
- /
- 제39권6호
- /
- pp.687-695
- /
- 2002
VoIP 음성 패킷 손실에 대한 대응 방안으로 제시되고 있는 매체 종속 FEC 기법은 통화 품질을 개선시키는 효과를 갖는데 반하여 한 프레임에 해당하는 추가지연이 발생하는 단점을 갖는다. 본 논문에서는 패킷 손실 복원에 사용되는 잉여 정보로 미래 프레임의 LSF 성분을 사용함으로써, 전송 지연을 줄이고 통화 품질을 개선할 수 있는 LSF FEC 기법을 제안하고 그 성능을 평가한다. 성능 평가를 위해서 VoIP에서 사용하는 ITU-T G.723.1, G.729 코덱을 Gilbert 손실 모델에 적용하고, PESQ 음질 측정 알고리즘을 사용하여 각 손실률 별로 MOS를 추정하는 방법을 사용한다. 본 논문에서 제안한 기법은 기존의 매체 종속 FEC 기법에 비해서 6.5ms∼27ms 이상의 지연 감소 효과를 가지고 있는 것으로 나타났으며, FEC를 적용하지 않은 경우와의 복원 음성 품질 비교 시뮬레이션 결과, 5% 정도의 현실적인 손실 환경에서 MOS 0.1 이상의 음질 개선 효과를 보였다.
PDF KSCI

A Line Spectrum Frequency Pairs Representation for Spectral Envelop Quantization

Park, Youngho;Lee, Won-Cheol;Bae, Myung-Jin
- 대한전자공학회:학술대회논문집
- /
- 대한전자공학회 2000년도 제13회 신호처리 합동 학술대회 논문집
- /
- pp.787-790
- /
- 2000
This paper introduces a new type of representation of the LSPs as a promising alternative used for transmitting the LPC parameters. Major contribution in this paper is that the vocal track information embedded on the spectral envelope can be represented in terms of the reduced number of LSF compared tn the conventional. Hence, it provides a possibility that LPC parameters could be quantized at a reduced bit rate without causing any major spectral distortion. The simulation result illustrates the capability of the proposed LSPs representation as an efficient quantization method via a proper rejection of the redundant pairs of pole and zero along the unit circle.
PDF

남녀 음성 변환 기술연구 (A Study On Male-To-Female Voice Conversion)

최정규;김재민;한민수
- 한국음향학회:학술대회논문집
- /
- 한국음향학회 2000년도 하계학술발표대회 논문집 제19권 1호
- /
- pp.115-118
- /
- 2000
Voice conversion technology is essential for TTS systems because the construction of speech database takes much effort. In this paper. male-to-female voice conversion technology in Korean LPC TTS system has been studied. In general. the parameters for voice color conversion are categorized into acoustic and prosodic parameters. This paper adopts LSF(Line Spectral Frequency) for acoustic parameter, pitch period and duration for prosodic parameters. In this paper. Pitch period is shortened by the half, duration is shortened by $25\%, and LSFs are shifted linearly for the voice conversion. And the synthesized speech is post-filtered by a bandpass filter. The proposed algorithm is simpler than other algorithms. for example, VQ and Neural Net based methods. And we don't even need to estimate formant information. The MOS(Mean Opinion Socre) test for naturalness shows 2.25 and for female closeness, 3.2. In conclusion, by using the proposed algorithm. male-to-female voice conversion system can be simply implemented with relatively successful results.
PDF

검색결과 10건 처리시간 0.021초

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

자세히 찾기

이미지 검색 (β)