• Title/Summary/Keyword: Speech detection

Search Result 471, Processing Time 0.03 seconds

Pitch Detection by Synchronizing the Phase of Noise-Corrupted Speech Signals (위상 동기화에 의한 잡음 음성의 피치 검출)

  • 이병국;배명진;안수길
    • The Journal of the Acoustical Society of Korea
    • /
    • v.11 no.1E
    • /
    • pp.42-49
    • /
    • 1992
  • 시간 영역에서 음성의 피치 정보를 추출하는 새로운 알고리즘을 제안한다. 이 알고리즘은, 위상 이 일치하는 고조파 성분의 합은 위상이 일치하지 않는 고조파 성분의 합의 경우보다 주기 정보를 분명 히 나타낸다는 사실을 이용한 것이다. 즉, 음성 신호의 위상 성분을 0으로 되도록 하여 실질적으로 기본 파와 모든 고조파 성분의 위상을 일치시킨다. 이 알고리즘은 잡음이 없는 음성의 경우 0.18%의 조오류 를 보이며, 0dB 눅의 경우에도 3.63%의 조오류를 보임으로써 잡음에 강건한 성질이 있음을 알 수 있다. 또한 시간 영역에서의 결정 논리를 사용하므로 피치 해상도가 우수하다. 전반적인 실험결과는 제안된 알고리즘이 피치 검출에 상당히 효율적임을 나타낸다.

  • PDF

Impostor Detection in Speaker Recognition Using Confusion-Based Confidence Measures

  • Kim, Kyu-Hong;Kim, Hoi-Rin;Hahn, Min-Soo
    • ETRI Journal
    • /
    • v.28 no.6
    • /
    • pp.811-814
    • /
    • 2006
  • In this letter, we introduce confusion-based confidence measures for detecting an impostor in speaker recognition, which does not require an alternative hypothesis. Most traditional speaker verification methods are based on a hypothesis test, and their performance depends on the robustness of an alternative hypothesis. Compared with the conventional Gaussian mixture model-universal background model (GMM-UBM) scheme, our confusion-based measures show better performance in noise-corrupted speech. The additional computational requirements for our methods are negligible when used to detect or reject impostors.

  • PDF

A Comparison and Analysis of Deep Learning Framework (딥 러닝 프레임워크의 비교 및 분석)

  • Lee, Yo-Seob;Moon, Phil-Joo
    • The Journal of the Korea institute of electronic communication sciences
    • /
    • v.12 no.1
    • /
    • pp.115-122
    • /
    • 2017
  • Deep learning is artificial intelligence technology that can teach people like themselves who need machine learning. Deep learning has become of the most promising in the development of artificial intelligence to understand the world and detection technology, and Google, Baidu and Facebook is the most developed in advance. In this paper, we discuss the kind of deep learning frameworks, compare and analyze the efficiency of the image and speech recognition field of it.

The role of prosodic phrasing in Korean word segmentation (음운 구조가 한국어 단어 분절에 미치는 영향)

  • Kim, Sa-Hyang
    • Proceedings of the KSPS conference
    • /
    • 2007.05a
    • /
    • pp.114-118
    • /
    • 2007
  • The current study investigates the degree to which various prosodic cues at the boundaries of a prosodic phrase in Korean (Accentual Phrase) contributed to word segmentation. Since most phonological words in Korean are produced as one AP, it was hypothesized that the detection of acoustic cues at AP boundaries would facilitate word segmentation. The prosodic characteristics of Korean APs include initial strengthening at the beginning of the phrase and pitch rise and final lengthening at the end. A perception experiment revealed that the cues that conform to the above-mentioned prosodic characteristics of Korean facilitated listeners' word segmentation. Results also showed that duration and amplitude cues were more helpful in segmentation than pitch. Further, the results showed that a pitch cue that did not conform to the Korean AP interfered with segmentation.

  • PDF

Effects of Multi-modality Cues on Personal Navigation in Wearable Computing (웨어러블 컴퓨터 환경의 개인 네비게이션 수행에 다중양식 단서가 미치는 영향)

  • Jeon, Ha-Young;Chae, Haeng-Suk;Hong, Ji-Young;Han, Kwang-Hee
    • Journal of the Ergonomics Society of Korea
    • /
    • v.26 no.4
    • /
    • pp.1-7
    • /
    • 2007
  • Navigation system or way finding in Wearable computer help disabled and impaired persons and it is impossible to be safe and efficient for drivers as well as pedestrian. Wearable computing situation must be multi-tasking simultaneously and users need minimal attention. In this paper, we used virtual environment as real way-finding similarly. The direction cues of navigation system are investigated as visual only, visual & auditory, and visual & speech. In the paper, the trial demonstrates the difference of performance in detection of directing and performance of motor and subjective satisfaction of user.

The Perceptual effect of 'Prosodic vs. Semantic' Focus Representation in Phoneme Detecting (음소 지각에 대한 초점의 운율적 실현과 의미적 실현의 효과(I))

  • Kim Hee-Sung;Jo Min-Ha;Kim Kee-Ho
    • Proceedings of the KSPS conference
    • /
    • 2006.05a
    • /
    • pp.71-74
    • /
    • 2006
  • The purpose of this study is to observe how Korean listeners detect a target phoneme with 'Focus' represented by prosodic prominence and question-induced semantic emphasis. According to the automated phoneme detection task using E-Prime, Korean listeners detected phoneme targets more rapidly when the target-bearing words were in prominence position and in question-induced position. However, when phoneme targets were in prominence position, response time was much faster than in question-induced position. The results suggest that the prosodic prominence which is explicit method of focus representation be more effective than question-inducing, implicit method of it, in phoneme detecting.

  • PDF

Weighted QPSK/PCM Speech Signal Detection with the Erasure Zone (가중치를 부여한 QPSK/PCM 음성신호의 소거대역 설정에 의한 신호수신)

  • Ahn, Seung-Choon;Lee, Moon-Ho
    • Proceedings of the KIEE Conference
    • /
    • 1988.07a
    • /
    • pp.179-182
    • /
    • 1988
  • Since the bits in any encoded PCM word are of different importance to the bit positions, in order to improve the signal to noise ratio the technique that the encoded signal bits are weighted for the QPSK transmission system, is presented. Also the erasure zone is established at the detector, such that if the output falls into the erasure zone, the regenerated sample is replaced by interpolation. Two weighting methods are shown here. One is the method that the same weighting profile is used to Q and I dimension in QPSK signal constellations. The other is diferent weighting to Q and I dimension. The gains of this new technique in overall signal s/n compared to conventional QPSK transmission system were 5 db and 2db, respectively.

  • PDF

An Information Transmission for Intelligent Train Operation (인텔리전트 열차운전을 위한 정보 전송)

  • Ahn, Sang-Kwon;Choi, Gui-Man;Kim, Yang-Mo
    • Proceedings of the KIEE Conference
    • /
    • 1997.07a
    • /
    • pp.339-341
    • /
    • 1997
  • This study is presenting the method for an effective data transmission in MAGLEV which is now tested and intends to provide for an intelligent operation of signal system in future. To exchange a lot of information, it is ideal to adopt a digital system and a micro-based system is essential for these purposes. FSK modulation and HDLC protocol are adopted on this study and information line assembly which is used as the information exchange, as the speech communication, and as the detection of speed and position is constructed in one unit. Actually this study is produced academic achievements of the data transmission system of MAGLEV train and an advanced method of intelligent operation in future railway system.

  • PDF

Design of Voice Activity Detection Algorithm for Variable Rate Speech Coders (가변전송률 음성부호화기 적용을 위한 음성활성도 측정 알고리즘 설계)

  • 김재원
    • The Journal of Korean Institute of Communications and Information Sciences
    • /
    • v.26 no.9A
    • /
    • pp.1451-1458
    • /
    • 2001
  • 디지털 이동통신 시스템에서 가장 빈번하게 발생하는 음성 서비스의 궁극적인 목표는 양호한 음성 품질과 높은 주파수 효율의 제공에 있다. 음성은 묵음 구간에 의하여 구분되어진 짧고 간헐적인 음성 에너지의 반복으로 표현 가능하며 실제 음성 통화중 활성 음성이 존재하는 구간은 약 40%, 나머지 60% 구간은 묵음 또는 상대방의 음성을 듣는 구간이다. 이 묵음 구간을 효율적으로 활용함에 의해 시스템의 스펙트럼 이득을 얻을 수 있다. 본 논문에서는 디지털 이동통신 시스템과 같이 다양하게 변화하는 주변 잡음 환경에서도 강건하게 동작 가능하여 10msec 프레임 크기를 갖는 음성부호화기에 적용 가능한 음성 활성도 측정 방안을 설계하였다. 설계된 알고리즘은 음성에너지, 스펙트럼 분포, 영교차율, 그리고 LPC 잔여신호의 Peakiness 측정값을 이용하였다.

  • PDF

On a Duration Control Method of Speech Waveform by an Automatic Pitch Point Detection (자동 피치시점 검출에 의한 음성신호의 지속시간 조절 법에 관한 연구)

  • Park Won;Park HyungBin;Bae MyungJin
    • Proceedings of the Acoustical Society of Korea Conference
    • /
    • autumn
    • /
    • pp.217-220
    • /
    • 2000
  • 일반적으로 고음질 음성합성을 하기 위해서는 합성음의 지속 시간을 변경하여 줌으로써 운율을 조절하는 기법이 필요하다 이에 먼저 고음질용 음성부호화법을 선정하여야 하고 정확한 피치와 피치시점검출을 통해서 음원분류가 되어야한다. 본 논문에서는 제안한 자동 피치시점 검출을 적용해서 운율조절에 필요한 지속시간 조절 법을 제안하고자 한다. 제안한 방법은 시간영역에서 직접 처리하기 때문에 피치동기분석이 용이하고 다른 영역으로의 변환과정이 불필요하다. 결과적으로 파형부호화법을 적용하고 제안한 자동 피치서점 검출에 의한 지속시간 조절법을 적용하였을 때 비교적 우수한 결과를 얻을 수 있었다.

  • PDF