Search | Korea Science

Real-time Implementation of the AMR Speech Coder Using $OakDSPCore^{\circledR}$ ($OakDSPCore^{\circledR}$를 이용한 적응형 다중 비트 (AMR) 음성 부호화기의 실시간 구현)

이남일;손창용;이동원;강상원
- The Journal of the Acoustical Society of Korea
- /
- v.20 no.6
- /
- pp.34-39
- /
- 2001
An adaptive multi-rate (AMR) speech coder was adopted as a standard of W-CDMA by 3GPP and ETSI. The AMR coder is based on the CELP algorithm operating at rates ranging from 12.2 kbps down to 4.75 kbps, and it is a source controlled codec according to the channel error conditions and the traffic loading. In this paper, we implement the DSP S/W of the AMR coder using OakDSPCore. The implementation is based on the CSD17C00A chip developed by C&S Technology, and it is tested using test vectors, for the AMR speech codec, provided by ETSI for the bit exact implementation. The DSP B/W requires 20.6 MIPS for the encoder and 2.7 MIPS for the decoder. Memories required by the Am coder were 21.97 kwords, 6.64 kwords and 15.1 kwords for code, data sections and data ROM, respectively. Also, actual sound input/output test using microphone and speaker demonstrates its proper real-time operation without distortions or delays.
PDF

Acoustic Characteristics of Korean Spoken by the Women Immigrants from Japan and Philippine (여성 결혼이민자들의 한국어 조음에 나타나는 음향음성학 특성 연구 - 일본과 필리핀 출신 여성 결혼이민자들을 대상으로)

Jo, Seon-Hui;Kim, Hyun-Gi;Kim, Sun-Jun
- Speech Sciences
- /
- v.15 no.3
- /
- pp.203-217
- /
- 2008
The number of Asian women immigrants in Korea is getting bigger and it's important to note that their communication problem in Korean causes not only the difficulty of adapting to Korean society but their children's speech-language disorder. To date there is little research on their acoustics characters and articulatory errors. Therefore, this study focuses on acoustic characters and articulatory error patterns of the women immigrants from Japan and Philippine based on the theory of "contrastive analysis". The subjects were 16 Japanese women immigrants(age: 42.5$\pm$4.4) and 14 Philippine women immigrants(age: 31.64$\pm$6.7) and control group consisted of 10 Korean women(age: 28.3$\pm$1.2). Speech and hearing of all subjects and control group were within normal limits. Speech samples were analyzed in a computer using CSL and data analysis was done on FFT widow for F1, F2, F3 of vowels and on wideband spectrogram for VOT of plosives and africatives. The results of this study were like this; For Japanese women immigrants, they had different articulatory patterns of /e/, /a/, /u/, /o/, /$\varepsilon$/, /m/ from those of Koreans and showed articulatory errors on the fortis and aspirated sounds. The reason is Japanese has only two distinctive characters for plosives and affricates; voicing and voiceless. The Philippine women immigrants also showed the same error patterns as the Japanese women immigrants. Especially the errors on aspirated sounds were prominent because their mother tongue has no distinctive characters about aspirated sounds. For vowels, they showed errors of /a/, /o/, /c/.
PDF

Audio /Speech Codec Using Variable Delay MDCT/IMDCT (가변 지연 MDCT/IMDCT를 이용한 오디오/음성 코덱)

Sangkil Lee;In-Sung Lee
- The Journal of Korea Institute of Information, Electronics, and Communication Technology
- /
- v.16 no.2
- /
- pp.69-76
- /
- 2023
A high-quality audio/voice codec using the MDCT/IMDCT process can perfectly restore the current frame through an overlap-add process with the previous frame. In the overlap-add process, an algorithm delay equal to the frame length occurs. In this paper, we propose a MDCT/IMDCT process that reduces algorithm delay by using a variable phase shift in MDCT/IMDCT process. In this paper, a low-delay audio/speech codec was proposed by applying the low delay MDCT/IMDCT algorithm to the ITU-T standard codec G.729.1 codec. The algorithm delay in the MDCT/IMDCT process can be reduced from 20 ms to 1.25 ms. The performance of the decoded output signal of the audio/speech codec to which low-delay MDCT/IMDCT is applied is evaluated through the PESQ test, which is an objective quality test method. Despite of the reduction in transmission delay, it was confirmed that there is no difference in sound quality from the conventional method.
https://doi.org/10.17661/jkiiect.2023.16.2.69 인용 PDF HTML

Narrowband to Wideband Conversion of Speech using Modularized Neural Network (모듈화 된 신경 회로망을 이용한 음성의 Narrowband에서 Wideband로의 변환)

Woo Dong Hun;Ko Charm Han;Kang Hyun Min;Kim Yoo Shin;Kim Hyung Soon
- Proceedings of the Acoustical Society of Korea Conference
- /
- autumn
- /
- pp.21-24
- /
- 2001
본 논문은 신경 회로망을 이용하여, 전화망 대역의 음성, 즉, narrowband 음성에서 wideband 음성을 복원하고자 했다. BP 알고리즘을 사용하는 기존의 신경 회로망의 경우에는 음성과 같이 복잡하고 크기가 큰 훈련데이터에 대해서는 훈련이 제대로 되지 않는 단점이 있다. 그러므로 븐 논문에서는 이를 해결하기 위해 입력으로 들어온 LPC 켑스트럼 벡터를 k-means 알고리즘을 이용하여 미리 정한 개수의 cluster로 나눈 다음, 각각의 cluster에 대해 독립적인 신경 회로망을 적용했다 이로 인해 각각의 신경 회로망은 제한되고 서로 상관관계가 많은 음성들만 훈련하면 되므로, 기존의 신경 회로망에서 생기는 훈련의 정체를 개선할 수 있었다. 또 clustering 과정에서 생기는 오류를 보완하기 위해 후보신경 로망들의 출력에 fuzzy 개념을 적용해서 최종 출력을 내도록 했다 실험 결과에서, 제안한 알고리즘은 기존의 codebook mapping 알고리즘보다 스펙트럼 거리척도에 의한 비교 및 주관적인 음질 평가 양쪽에서 개선된 성능을 보였다.
PDF

A Real-time Implementation of G.729.1 Codec on an ARM Processor for the Improvement of VoWiFi Voice Quality (VoWiFi 음질 향상을 위한 G.729.1 광대역 코덱의 ARM 프로세서에의 실시간 구현)

Park, Nam-In;Kang, Jin-Ah;Kim, Hong-Kook
- 한국HCI학회:학술대회논문집
- /
- 2008.02a
- /
- pp.230-235
- /
- 2008
This paper addresses issues associated with the real-time implementation of a wideband speech codec such as ITU-T G. 729. 1 on an ARM processor in order to provide an improved voice quality of a VoWiFi service. The real-time implementation features in optimizing the C-source code of G.729. 1 and replacing several parts of the codec algorithm with faster ones. The performance of the implementation is measured by the CPU time spent for G.729.1 on the ARM926EJ processor that is used for a VoWiFi phone. It is shown from the experiments that the G.729.1 codec works in real-time with better voice quality than G 729 codec that is conventionally used for VoIP or VoWiFi phones.
PDF

Frequency Band Selection Exited Linear Prediction Wideband Speech/Audio Coding Using SBR (SBR을 이용한 주파수 밴드선택 여기 선형예측 광대역 음성/오디오 부호화)

Jang, Sunghoon;Lee, Insung
- The Journal of the Acoustical Society of Korea
- /
- v.32 no.6
- /
- pp.556-562
- /
- 2013
This paper is aimed to improve performance of Band-Selection speech/audio Coder reconstucted band spectrum that is not sent by the comfort noise. To improve the performance, we use the Spectral Band Replication(SBR) technique instead of substitution of Comfort noise. To synthesize SBR signal, the SBR algorithm is referenced in selected signals and the spectrum synthesized by SBR is injected to non-selected band. Each sub-band spectrum has been energy-weighted by real audio signal. We propose the enhanced the Band-Selection Coder that utilizes synthesized SBR signal from selected signal instead of comfort noise.
https://doi.org/10.7776/ASK.2013.32.6.556 인용 PDF KSCI

Design of Wideband Speech Coder Using the G.723-1,G.729 Combined with MLT (G.723.1,G.729 부호화기와 MLT 방법을 이용한 광대역 음성 부호화기 설계)

김정중;김종학;이인성
- Proceedings of the IEEK Conference
- /
- 2001.09a
- /
- pp.939-942
- /
- 2001
본 논문에서는 ITU-T G.723.1, G.729 부호화기와 MLT(Modulated Lapped Transform) 방법을 이용한 광대역 음성 부호화방법을 제안한다. 제안된 광대역 음성부호화 방법은 16 kHz로 샘플링된 입력신호를 QMF(Quadrature Mirror Filter)사용하여 저대역과 고대역으로 나누며, 각 대역은 8 kHz의 샘플링을 갖는 협대역 음성 신호로 변환된다. 고대역은 MLT변환 후 벡터 양자화하며 또한 MLT를 사용한 ATC(Adaptive Transform Coding)방법을 적용하여 표현하며 저대역은 G.723.1과 G.729 부호화기를 사용한다. 설계된 광대역 음성부호화기의 성능을 평가하기 위하여 MOS (Mean Opinion score)실험을 수행하였다. MOS 실험을 통해 16 kbps G.729-MLT VQ방식이 G.722 56kbps 와 비슷한 음질을 나타내었다.
PDF

Development of Wideband GSM-EFR Speech Coding Algorithm with Application of Wavelet Transform to High-Band Signal (High-Band 신호에 웨이브렛 변환을 적용한 광대역 GSM-EFR 음성부호화 알고리즘 개발)

이승원;배건성
- Proceedings of the IEEK Conference
- /
- 2000.09a
- /
- pp.783-786
- /
- 2000
본 논문에서는 웨이브렛 변환을 적용한 광대역 음성부호화 알고리즘을 제안하였다. 제안한 음성부호화 알고리즘은 split-band 구조를 가지며, 16 kHz로 sampling된 입력신호를 QMF를 이용해서 동일한 대역폭을 갖는 두 개의 subband 신호로 나누고 이를 8kHz의 sampling율을 갖도록 downsampling 한다. 그리고 저대역 신호는 GSM-EFR 음성부호화 알고리즘을 이용하여 부호화하고, 고대역 신호는 DWT(Discrete Wavelet Transform)을 적용하여 subband로 나누어 부호화하였다. 각 subband에서 양자화 된 파라미터는 IDWT(Inverse DWT)과정을 거쳐서 upsampling되고 합성 QMF를 통과시켜 최종 합성음을 구하였다. 제안한 음성부호화기는 저대역 신호의 GSM-EFR 부호화에 12.2 kbps, 웨이브렛 변환을 이용한 고대역 신호의 부호화에 7.8 kbps로 전체 20 kbps의 전송율을 가지면서 G.722 표준안의 56 kbps에서의 합성음과 비슷한 음질을 나타내었다.
PDF

Design of Multi Rate Wideband Speech Coder Using the AMR(Adaptive Multi-Rate) Coder (AMR 부호화기와 결합된 다전송률 광대역 음성부호화기 설계)

김은주;이호창;이인성
- Proceedings of the IEEK Conference
- /
- 2000.09a
- /
- pp.755-758
- /
- 2000
본 논문에서는 AMR(Adaptive Multi-Rate)를 이용하여 광대역 음성부호화기를 설계하였다. 16kHz로 샘플링 된 입력 신호를 QMF 필터에 의해 두 개의 대역으로 나누어, 각각 decimation하여 두 개의 8kHz 샘플링 신호로 변환시킨 후 저대역(0Hz-3400Hz)의 신호와 고대역(3400Hz -7000Hz)의 신호로 나누어 각각 부호화한다. 나누어진 두 개의 협대역 음성신호는 AMR(Adaptive Multi-Rate)과 ATC(Adaptive Transform Coding)을 사용하여 각각 부호화되어 전송된다. 두 대역으로부터 부호화된 정보는 20.2kbps에서 12.75kbps까지의 전송률을 갖고, 수신단에서는 각 대역을 AMR과ATC방법으로 역부호화하여 음성신호를 합성한다. 설계된 광대역 음성부호화기의 성능을 평가하기 위해 ITU-T의 표준안인 G.722를 포함하여 MOS 시험을 하였다.
PDF

Quantization on Wideband Speech Codec for Next Generation Packet Phone (차세대 패킷 전화용 광대역 음성 부호화기의 양자화에 대한 연구)

Kim Youngvo;Jeong Byounghak;Park Hochong
- Proceedings of the Acoustical Society of Korea Conference
- /
- autumn
- /
- pp.81-84
- /
- 2004
패킷망을 통한 음성 통신이 발달됨에 따라 패킷 스위칭 채널 환경에서 계층적 구조를 가지는 광대역 음성 부호화기의 개발에 대한 요구가 늘어나고 있다. 본 논문에서는 이러한 차세대 패킷 전화용 광대역 음성 부호화기의 상위 대역에 대해서 효율적인 양자화 방법을 제안한다. 먼저 전체 프레임을 다수의 짧은 부프레임으로 구분하고, 각각의 부프레임에 MLT(Modulated Lapped Transform)변환을 적용하여 주파수 영역으로 변환하여 2차원 구조의 데이터 행렬을 생성한다. 이러한 2차원 구조의 데이터를 크기와 부호로 분리하고, 크기는 2차원 DCT를 사용하여 시간과 주파수 영역에서의 신호 압축을 동시에 얻을 수 있게 하였다. 이와 같은 새로운 구조를 활용하여 기존의 방법보다 Energy Compaction 효과를 높이고 양자화 성능을 향상시킬 수 있었다. 또한 Core Layer의 부호화된 파라미터를 상위 대역의 양자화에 이용함으로써 그 성능을 향상시킬 수 있는 방법을 제안한다.
PDF

Search Result 57, Processing Time 0.024 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)