Search | Korea Science

CNN based Complex Spectrogram Enhancement in Multi-Rotor UAV Environments (멀티로터 UAV 환경에서의 CNN 기반 복소 스펙트로그램 향상 기법)

Kim, Young-Jin;Kim, Eun-Gyung
- Journal of the Korea Institute of Information and Communication Engineering
- /
- v.24 no.4
- /
- pp.459-466
- /
- 2020
The sound collected through the multi-rotor unmanned aerial vehicle (UAV) includes the ego noise generated by the motor or propeller, or the wind noise generated during the flight, and thus the quality is greatly impaired. In a multi-rotor UAV environment, both the magnitude and phase of the target sound are greatly corrupted, so it is necessary to enhance the sound in consideration of both the magnitude and phase. However, it is difficult to improve the phase because it does not show the structural characteristics. in this study, we propose a CNN-based complex spectrogram enhancement method that removes noise based on complex spectrogram that can represent both magnitude and phase. Experimental results reveal that the proposed method improves enhancement performance by considering both the magnitude and phase of the complex spectrogram.
https://doi.org/10.6109/jkiice.2020.24.4.459 인용 PDF KSCI

Influence of the Shear Property of Seabed Appearing in the Striation Pattern of the Spectrogram of Ship-radiated Noise Measured in a Shallow Sea (천해에서 측정한 선박 방사소음 스펙트로그램의 줄무늬 패턴에 나타나는 해저면 전단성 영향)

Lee, Seong-Wook;Hahn, Joo-Young;Baek, Woon;Na, Jung-Yul
- The Journal of the Acoustical Society of Korea
- /
- v.23 no.3
- /
- pp.197-205
- /
- 2004
This paper represents the results of interpretation on the cause of sign changing of the striation slopes appearing in the range-frequency domain spectrogram of ship-radiated noise measured in a shallow sea. Striation patterns and dispersion characteristics simulated from a numerical model based on mode theory at various seabed conditions show that the sign changing of the striation slopes appearing in measured signal is caused by the shear property of seabed. more specifically by the shear property of the basement lying below the sediment which is estimated about 3±1m thick.
PDF KSCI

A Visual Study of the Phonemic Awareness (음소인지에 관한 시각적 연구)

Park, Heesuk
- Journal of Digital Contents Society
- /
- v.16 no.2
- /
- pp.219-225
- /
- 2015
This experimental study aims at understanding the Korean subjects' phonemic awareness in the English minimal pairs. For the purpose of the experiment, English listening comprehension tests were designed using minimal pairs and conducted among subjects, and the results of the tests were analyzed with the help of spectrogram. From the results of this study, I could find out three important things: First, subjects have difficulty in understanding and distinguishing English vowel minimal pairs. Second, among the English vowel minimal pairs, they had much difficulty in distinguishing between /ə:/ and /ɔ:/. Third, subjects could recognize the semivowel /w/ in words without any difficulty. In addition to this, I tried to analyze the results using the spectrogram, which helps to educate students effectively.
https://doi.org/10.9728/dcs.2015.16.2.219 인용 PDF KSCI

A Method of Sound Segmentation in Time-Frequency Domain Using Peaks and Valleys in Spectrogram for Speech Separation (음성 분리를 위한 스펙트로그램의 마루와 골을 이용한 시간-주파수 공간에서 소리 분할 기법)

Lim, Sung-Kil;Lee, Hyon-Soo
- The Journal of the Acoustical Society of Korea
- /
- v.27 no.8
- /
- pp.418-426
- /
- 2008
In this paper, we propose an algorithm for the frequency channel segmentation using peaks and valleys in spectrogram. The frequency channel segments means that local groups of channels in frequency domain that could be arisen from the same sound source. The proposed algorithm is based on the smoothed spectrum of the input sound. Peaks and valleys in the smoothed spectrum are used to determine centers and boundaries of segments, respectively. To evaluate a suitableness of the proposed segmentation algorithm before that the grouping stage is applied, we compare the synthesized results using ideal mask with that of proposed algorithm. Simulations are performed with mixed speech signals with narrow band noises, wide band noises and other speech signals.
https://doi.org/10.7776/ASK.2008.27.8.418 인용 PDF KSCI

Differentiation of Adductor-Type Spasmodic Dysphonia from Muscle Tension Dysphonia Using Spectrogram (스펙트로그램을 이용한 내전형 연축성 발성 장애와 근긴장성 발성 장애의 감별)

Noh, Seung Ho;Kim, So Yean;Cho, Jae Kyung;Lee, Sang Hyuk;Jin, Sung Min
- Journal of the Korean Society of Laryngology, Phoniatrics and Logopedics
- /
- v.28 no.2
- /
- pp.100-105
- /
- 2017
Background and Objectives : Adductor type spasmodic dysphonia (ADSD) is neurogenic disorder and focal laryngeal dystonia, while muscle tension dysphonia (MTD) is caused by functional voice disorder. Both ADSD and MTD may be associated with excessive supraglottic contraction and compensation, resulting in a strained voice quality with spastic voice breaks. The aim of this study was to determine the utility of spectrogram analysis in the differentiation of ADSD from MTD. Materials and Methods : From 2015 through 2017, 17 patients of ADSD and 20 of MTD, underwent acoustic recording and phonatory function studies, were enrolled. Jitter (frequency perturbation), Shimmer (amplitude perturbation) were obtained using MDVP (Multi-dimensional Voice Program) and GRBAS scale was used for perceptual evaluation. The two speech therapist evaluated a wide band (11,250 Hz) spectrogram by blind test using 4 scales (0-3 point) for four spectral findings, abrupt voice breaks, irregular wide spaced vertical striations, well defined formants and high frequency spectral noise. Results : Jitter, Shimmer and GRBAS were not found different between two groups with no significant correlation (p>0.05). Abrupt voice breaks and irregular wide spaced vertical striations of ADSD were significantly higher than those of MTD with strong correlation (p<0.01). High frequency spectral noise of MTD were higher than those of ADSD with strong correlation (p<0.01). Well defined formants were not found different between two groups. Conclusion : The wide band spectrograms provided visual perceptual information can differentiate ADSD from MTD. Spectrogram analysis is a useful diagnostic tool for differentiating ADSD from MTD where perceptual analysis and clinical evaluation alone are insufficient.
PDF

A Comparison Study on the Speech Signal Parameters for Chinese Leaners' Korean Pronunciation Errors - Focused on Korean /ㄹ/ Sound (중국인 학습자의 한국어 발음 오류에 대한 음성 신호 파라미터들의 비교 연구 - 한국어의 /ㄹ/ 발음을 중심으로)

Lee, Kang-Hee;You, Kwang-Bock;Lim, Ha-Young
- Asia-pacific Journal of Multimedia Services Convergent with Art, Humanities, and Sociology
- /
- v.7 no.6
- /
- pp.239-246
- /
- 2017
This paper compares the speech signal parameters between Korean and Chinese for Korean pronunciation /ㄹ/, which is caused many errors by Chinese leaners. Allophones of /ㄹ/ in Korean is divided into lateral group and tap group. It has been investigated the reasons for these errors by studying the similarity and the differences between Korean /ㄹ/ pronunciation and its corresponding Chinese pronunciation. In this paper, for the purpose of comparison the speech signal parameters such as energy, waveform in time domain, spectrogram in frequency domain, pitch based on ACF, Formant frequencies are used. From the phonological perspective the speech signal parameters such as signal energy, a waveform in the time domain, a spectrogram in the frequency domain, the pitch (F0) based on autocorrelation function (ACF), Formant frequencies (f1, f2, f3, and f4) are measured and compared. The data, which are composed of the group of Korean words by through a philological investigation, are used and simulated in this paper. According to the simulation results of the energy and spectrogram, there are meaningful differences between Korean native speakers and Chinese leaners for Korean /ㄹ/ pronunciation. The simulation results also show some differences even other parameters. It could be expected that Chinese learners are able to reduce the errors considerably by exploiting the parameters used in this paper.
https://doi.org/10.14257/ajmahs.2017.06.56 인용

Spontaneous Speech Emotion Recognition Based On Spectrogram With Convolutional Neural Network (CNN 기반 스펙트로그램을 이용한 자유발화 음성감정인식)

Guiyoung Son;Soonil Kwon
- The Transactions of the Korea Information Processing Society
- /
- v.13 no.6
- /
- pp.284-290
- /
- 2024
Speech emotion recognition (SER) is a technique that is used to analyze the speaker's voice patterns, including vibration, intensity, and tone, to determine their emotional state. There has been an increase in interest in artificial intelligence (AI) techniques, which are now widely used in medicine, education, industry, and the military. Nevertheless, existing researchers have attained impressive results by utilizing acted-out speech from skilled actors in a controlled environment for various scenarios. In particular, there is a mismatch between acted and spontaneous speech since acted speech includes more explicit emotional expressions than spontaneous speech. For this reason, spontaneous speech-emotion recognition remains a challenging task. This paper aims to conduct emotion recognition and improve performance using spontaneous speech data. To this end, we implement deep learning-based speech emotion recognition using the VGG (Visual Geometry Group) after converting 1-dimensional audio signals into a 2-dimensional spectrogram image. The experimental evaluations are performed on the Korean spontaneous emotional speech database from AI-Hub, consisting of 7 emotions, i.e., joy, love, anger, fear, sadness, surprise, and neutral. As a result, we achieved an average accuracy of 83.5% and 73.0% for adults and young people using a time-frequency 2-dimension spectrogram, respectively. In conclusion, our findings demonstrated that the suggested framework outperformed current state-of-the-art techniques for spontaneous speech and showed a promising performance despite the difficulty in quantifying spontaneous speech emotional expression.
https://doi.org/10.3745/TKIPS.2024.13.6.284 인용 PDF

Audio Genre Classification based on Deep Learning using Spectrogram (스펙트로그램을 이용한 딥 러닝 기반의 오디오 장르 분류 기술)

Jang, Woo-Jin;Yun, Ho-Won;Shin, Seong-Hyeon;Park, Ho-chong
- Proceedings of the Korean Society of Broadcast Engineers Conference
- /
- 2016.06a
- /
- pp.90-91
- /
- 2016
본 논문에서는 스펙트로그램을 이용한 딥 러닝 기반의 오디오 장르 분류 기술을 제안한다. 기존의 오디오 장르 분류는 대부분 GMM 알고리즘을 이용하고, GMM의 특성에 따라 입력 성분들이 서로 직교한 성질을 갖는 MFCC를 오디오의 특성으로 사용한다. 그러나 딥 러닝을 입력의 성질에 제한이 없으므로 MFCC보다 가공되지 않은 특성을 사용할 수 있고, 이는 오디오의 특성을 더 명확히 표현하기 때문에 효과적인 학습을 할 수 있다. 본 논문에서는 딥 러닝에 효과적인 특성을 구하기 위하여 스펙트로그램(spectrogram)을 사용하여 오디오 특성을 추출하는 방법을 제안한다. 제안한 방법을 사용한면 MFCC를 특성으로 하는 딥 러닝보다 더 높은 인식률을 얻을 수 있다.
PDF

Fluidic velocity sensing with a speaker based optical doppler tomography (유속 센싱을 위한 스피커형 광학적 유체 단층촬영 기술)

Lee, Chang-Ho;Kim, Jee-Hyun
- Journal of Sensor Science and Technology
- /
- v.17 no.4
- /
- pp.317-324
- /
- 2008
This paper presents an optical doppler tomography(ODT) system using a speaker as a method to achieve depth measurement in a flowing sample. The use of the speaker provides easy implementation with a low cost. The nonlinear characteristics of the speaker has hindered its adaptation because it produces inconsistent fringe frequencies at different depths. This paper reports an adaptive algorithm to compensate the nonlinear characteristics, and could, resultantly, acquire the Doppler frequency shift caused by the sample. The experiment utilizes a flowing scattering particle solution in a capillary tube at a certain flow rate. The Doppler frequency profile over the lumen was calculated by using spectrogram method. and we obtained the velocity image of the sample.
https://doi.org/10.5369/JSST.2008.17.4.317 인용 PDF KSCI

A Study on the Pulse Doppler System with M-mode Image and Spectrum Analyzer (주파수 해석기와 M-mode 영상을 갖는 펄스 도플러 장치의 개발에 관한 연구)

Jeong, Taek-Seob;Park, Sei-Hyun;Kim, Young-Kil
- Proceedings of the KIEE Conference
- /
- 1987.07b
- /
- pp.1217-1220
- /
- 1987
We have developed a Ultra Sound Pulsed Doppler System with two-dimensional M-mode image and Spectrum analyzer. The image of the M-mode is composed of time and depth axes. The Spectrum analyzer shows the spectrum of Doppler signal which represents the velocity component of time dependent blood-flow behavior. The spectrogram using Spectrum analyzer is composed of frequency and amplitude axes. The outputs of the system are audio signals, velocity curves, velocity profiles, M-mode images and spectrogram.
PDF

Search Result 236, Processing Time 0.032 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)