Search | Korea Science

Bird sounds classification by combining PNCC and robust Mel-log filter bank features (PNCC와 robust Mel-log filter bank 특징을 결합한 조류 울음소리 분류)

Badi, Alzahra;Ko, Kyungdeuk;Ko, Hanseok
- The Journal of the Acoustical Society of Korea
- /
- v.38 no.1
- /
- pp.39-46
- /
- 2019
In this paper, combining features is proposed as a way to enhance the classification accuracy of sounds under noisy environments using the CNN (Convolutional Neural Network) structure. A robust log Mel-filter bank using Wiener filter and PNCCs (Power Normalized Cepstral Coefficients) are extracted to form a 2-dimensional feature that is used as input to the CNN structure. An ebird database is used to classify 43 types of bird species in their natural environment. To evaluate the performance of the combined features under noisy environments, the database is augmented with 3 types of noise under 4 different SNRs (Signal to Noise Ratios) (20 dB, 10 dB, 5 dB, 0 dB). The combined feature is compared to the log Mel-filter bank with and without incorporating the Wiener filter and the PNCCs. The combined feature is shown to outperform the other mentioned features under clean environments with a 1.34 % increase in overall average accuracy. Additionally, the accuracy under noisy environments at the 4 SNR levels is increased by 1.06 % and 0.65 % for shop and schoolyard noise backgrounds, respectively.
https://doi.org/10.7776/ASK.2019.38.1.039 인용 PDF KSCI HTML

OnExpo HOT&COOL / COOL COMPANY 어뮤즈텍

O, Suk-Hyeon
- Digital Contents
- /
- no.12 s.127
- /
- pp.78-79
- /
- 2003
음악을 더 재미있게! 더 편리하게! 종이악보에도 디지털 바람이 불었다. 그리고 그 바람의 선두에는 뮤즈북 스코어(www.musebook.co.kr)가 있다. 뮤즈북 스코어는 최근 개발 된 태블릿 PC를 이용해 수만 장의 악보를 저장해 연주할 수 있는 프로그램으로 음악인식기술을 이용한‘음악인식 전자악보’이다. 피아노 악보대에 종이악보 대신 태블릿 PC를 올려놓고 하드디스크에 저장된 MusicXLM 전자악보들을 마음대로 불러서 사용하는 뮤즈북 스코어는 피아노를 연주 하면 태블릿 PC의 마이크로 소리를 듣고 분석해 자동으로 페이지를 넘겨준다. 사용자의 연주속도 가 빨라지거나 느려지더라도 현재 연주위치를 계 속 추적하기 때문에 사용자는 연주에만 집중할 수 있다.
PDF

Knowledge Representation Method for Dynamic Gesture Recognition (동적 제스쳐 인식을 위한 지식 표현 기법)

고일주;최형일
- Proceedings of the Korean Institute of Intelligent Systems Conference
- /
- 1995.10b
- /
- pp.293-299
- /
- 1995
본 논문은 컴퓨터 시각을 이용하여 동적 제스쳐를 인식하기 위한 효율적인 지식 표 현 기법의 개발을 목표로 한다. 제스쳐란 시각적인 언어로서 소리를 대신하여 몸짓이나 손 짓을 통하여 자신의 생각이나 의도를 전달하는 보조적인 의사 전달 수단이다. 제안된 기법 은 여러 다양한 지식을 통합하여 총체적으로 표현하기에 적합한 프레임 구조를 기반으로 한 다. 프레임 지식을 물체의 특성을 표현하는 객체 지식, 물체의 움직임을 표현하는 행동 지 식, 그리고 객체 지식과 행동 지식의 순서화 된 집함으로써 동적인 제스쳐를 표현하는 스키 마로 분류한다.
PDF

QRAS-based Algorithm for Omnidirectional Sound Source Determination Without Blind Spots (사각영역이 없는 전방향 음원인식을 위한 QRAS 기반의 알고리즘)

Kim, Youngeon;Park, Gooman
- Journal of Broadcast Engineering
- /
- v.27 no.1
- /
- pp.91-103
- /
- 2022
Determination of sound source characteristics such as: sound volume, direction and distance to the source is one of the important techniques for unmanned systems like autonomous vehicles, robot systems and AI speakers. There are multiple methods of determining the direction and distance to the sound source, e.g., using a radar, a rider, an ultrasonic wave and a RF signal with a sound. These methods require the transmission of signals and cannot accurately identify sound sources generated in the obstructed region due to obstacles. In this paper, we have implemented and evaluated a method of detecting and identifying the sound in the audible frequency band by a method of recognizing the volume, direction, and distance to the sound source that is generated in the periphery including the invisible region. A cross-shaped based sound source recognition algorithm, which is mainly used for identifying a sound source, can measure the volume and locate the direction of the sound source, but the method has a problem with "blind spots". In addition, a serious limitation for this type of algorithm is lack of capability to determine the distance to the sound source. In order to overcome the limitations of this existing method, we propose a QRAS-based algorithm that uses rectangular-shaped technology. This method can determine the volume, direction, and distance to the sound source, which is an improvement over the cross-shaped based algorithm. The QRAS-based algorithm for the OSSD uses 6 AITDs derived from four microphones which are deployed in a rectangular-shaped configuration. The QRAS-based algorithm can solve existing problems of the cross-shaped based algorithms like blind spots, and it can determine the distance to the sound source. Experiments have demonstrated that the proposed QRAS-based algorithm for OSSD can reliably determine sound volume along with direction and distance to the sound source, which avoiding blind spots.
https://doi.org/10.5909/JBE.2022.27.1.91 인용 PDF KSCI KPUBS

An Arrangement Method of Voice and Sound Feedback According to the Operation : For Interaction of Domestic Appliance (조작 방식에 따른 음성과 소리 피드백의 할당 방법 가전제품과의 상호작용을 중심으로)

Hong, Eun-ji;Hwang, Hae-jeong;Kang, Youn-ah
- Journal of the HCI Society of Korea
- /
- v.11 no.2
- /
- pp.15-22
- /
- 2016
The ways to interact with digital appliances are becoming more diverse. Users can control appliances using a remote control and a touch-screen, and appliances can send users feedback through various ways such as sound, voice, and visual signals. However, there is little research on how to define which output method to use for providing feedback according to the user' input method. In this study, we designed an experimental study that seeks to identify how to appropriately match the output method - voice and sound - based on the user input - voice and button. We made four types of interaction with two kinds input methods and two kinds of output methods. For the four interaction types, we compared the usability, perceived satisfaction, preference and suitability. Results reveals that the output method affects the ease of use and perceived satisfaction of the input method. The voice input method with sound feedback was evaluated more satisfying than with the voice feedback. However, the keying input method with voice feedback was evaluated more satisfying than with sound feedback. The keying input method was more dependent on the output method than the voice input method. We also found that the feedback method of appliances determines the perceived appropriateness of the interaction.
PDF KSCI

음성인식기술의 오늘과 내일

Hyeon, Won-Bok
- The Science & Technology
- /
- v.31 no.4 s.347
- /
- pp.75-80
- /
- 1998
기계가 말하고 알아듣는 시대가 빠른 걸음으로 다가오고 있다. 21세기 초에는 외국어를 모르는 사람들도 외국인과 대화할 때 거추장스럽게 통역관을 내세우지 않아도 된다. 통역용 소프트웨어가 내장된 전화기가 등장하는가 하면 포켓용 통역장치가 그때그때 대화내용을 우리말과 외국어로 옮겨 합성소리로 알려준다.
PDF

뉴밀레니엄의 비전(2) - 21세기의 사무실

Korean Federation of Science and Technology Societies
- The Science & Technology
- /
- v.33 no.9 s.376
- /
- pp.32-33
- /
- 2000
21세기의 화이트칼라는 아침에 출근하면 보안이 잘 된 지능형 문으로 걸어 들어와서 데스크탑의 가상조수가 그날의 스케줄을 큰 소리로 읽는 것을 듣느다. 그리고 지능형 의자에 앉아서 그날의 할 일을 음성인식장치 PC로 챙긴다. 평판스크린으로 된 벽의 영상과 데이터를 보면서... 멀리 있는 동료들과 얘기하려면 실물 그대로 입체 비디오 회의시스템에 불러낸다.
PDF

디지털 경제를 주도할 디지털 컨텐츠 산업의 육성방향

박영일
- Proceedings of the Korea Database Society Conference
- /
- 1999.10a
- /
- pp.1-11
- /
- 1999
o 디지털컨텐츠(멀티미디어컨텐츠)란 무엇인가\ulcorner 멀티미디어 : 기존 아날로그 기술에서 개별적으로 성장했던 문자, 음성, 사진, 비디오, 애니메이션의 미디어 영역들이 디지털 기술이 발달하면서 통합된 미디어를 말함. 디지털화는 글, 소리, 그림, 영상, 숫자 등의 온갖 정보들을 컴퓨터가 인식할 수 있는 신호(2진수 코드)로 바꾸는 것임. (중략)
PDF

A Development of Infant Education Content for Animal Study (동물모형 학습을 위한 유아교육 콘텐츠 개발)

Lee, Kwang-Hyoung;Kim, Jung-Jae
- Journal of the Korea Academia-Industrial cooperation Society
- /
- v.11 no.9
- /
- pp.3510-3516
- /
- 2010
In this paper to make young children to learn habits of the animals, crying, features, and English and Korean language, The system was developed to target the zoo various animals exist. If young child places a doll on the front of interesting animal, then young child can learn to look through the display connected to the model. The zoo is reducing the current appearance of the zoo, sensors that can recognize animals are attached to each cage. Attached to each sensor has a unique ID, If this approach recognizes a doll baby and will transmit a unique ID to the handler. Transmitted ID search the matched value sent from the database to retrieve the content and then the content is to be output through the output device. Also if the doll near the animal's room, young children find out animal sound and basic learning by multimedia effects. At the same time Korean, English, Mathematics are learned.
https://doi.org/10.5762/KAIS.2010.11.9.3510 인용 PDF KSCI

Development of Sound-sensible Security Camera based on Raspberry Pi (라즈베리파이 기반 소리인식 보안카메라 개발)

Park, Dae-Bok;Kim, Sun-Hyuk;Kim, Ju-Young;Rho, Young J.
- Proceedings of the Korea Information Processing Society Conference
- /
- 2015.10a
- /
- pp.1563-1566
- /
- 2015
보안과 관련된 기술이 발전하여 대규모의 장소에 적합한 보안시스템들이 많이 개발되었다. 특히 CCTV를 이용한 감시카메라의 형태도 다양화되었다. 스마트폰의 어플리케이션이나 웹을 통해서 어디서든 감시할 수도 있어, 이를 통해 보안사고 시에 빠른 대처가 가능하다. 하지만 대규모 시스템이 아닌 경우에는 침입자 발견이 늦고, 뒤늦은 대처로 인해 큰 피해가 발생할 수 있다. 라즈베리파이, 실드 보드 등 기타 하드웨어들을 통하여 침입자를 스스로 감지하여 사용자에게 즉시 알림을 전송함으로써 보안사고에 대한 대처를 빠르고 효율적으로 할 수 있는 보안카메라를 구현하였다. 본 보안 시스템은 소리의 방향을 계산하고 정확한 방향으로의 보정을 통하여 최초 침입자를 인식한다. 이후 이미지트래킹을 통하여 침입자를 추적한다. 무선 네트워크를 이용하기 때문에 네트워크가 지원되는 어느 장소에서든지 사용이 가능하다. 대규모 보안시스템을 설치할 여건이 되기 어려운 작은 공장, 상가, 사무실 등에서 보안시스템으로 사용되면 유용할 것이다. 자세한 개발 내용은 본문에 기술한다.
https://doi.org/10.3745/PKIPS.y2015m10a.1563 인용 PDF

Search Result 212, Processing Time 0.045 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)