Search | Korea Science

One-shot multi-speaker text-to-speech using RawNet3 speaker representation (RawNet3를 통해 추출한 화자 특성 기반 원샷 다화자 음성합성 시스템)

Sohee Han;Jisub Um;Hoirin Kim
- Phonetics and Speech Sciences
- /
- v.16 no.1
- /
- pp.67-76
- /
- 2024
Recent advances in text-to-speech (TTS) technology have significantly improved the quality of synthesized speech, reaching a level where it can closely imitate natural human speech. Especially, TTS models offering various voice characteristics and personalized speech, are widely utilized in fields such as artificial intelligence (AI) tutors, advertising, and video dubbing. Accordingly, in this paper, we propose a one-shot multi-speaker TTS system that can ensure acoustic diversity and synthesize personalized voice by generating speech using unseen target speakers' utterances. The proposed model integrates a speaker encoder into a TTS model consisting of the FastSpeech2 acoustic model and the HiFi-GAN vocoder. The speaker encoder, based on the pre-trained RawNet3, extracts speaker-specific voice features. Furthermore, the proposed approach not only includes an English one-shot multi-speaker TTS but also introduces a Korean one-shot multi-speaker TTS. We evaluate naturalness and speaker similarity of the generated speech using objective and subjective metrics. In the subjective evaluation, the proposed Korean one-shot multi-speaker TTS obtained naturalness mean opinion score (NMOS) of 3.36 and similarity MOS (SMOS) of 3.16. The objective evaluation of the proposed English and Korean one-shot multi-speaker TTS showed a prediction MOS (P-MOS) of 2.54 and 3.74, respectively. These results indicate that the performance of our proposed model is improved over the baseline models in terms of both naturalness and speaker similarity.
https://doi.org/10.13064/KSSS.2024.16.1.067 인용 PDF

Speech Intelligibility Analysis on the Laser Detected Sound of the Glass Windows (유리창의 레이저 탐지음에 대한 음성명료도 분석)

Kim, Seock-Hyun;Lee, Hyun-Woo;Kim, Hee-Dong
- The Journal of the Acoustical Society of Korea
- /
- v.28 no.2
- /
- pp.127-134
- /
- 2009
In this study, possibility of the laser eavesdropping is investigated on the window glasses with various thicknesses, Glass windows are excited by maximum length sequency (MLS) signal and the vibration sound is detected by a laser doppler vibrometer. From the detected sound, speech intelligibility is objectively estimated. Speech transmission index (STI), which is based on the modulation transfer function (MTF). is calculated for the estimation. Finally, disturbing wave effect on the speech intelligibility is analysed by using an outside speaker and a window shaker attached on the glass window. The purpose of the study is to estimate the possibility of remote eavesdropping by the laser sensor and to evaluate the performance of the homemade window shaker to protect from the remote eavesdropping.
https://doi.org/10.7776/ASK.2009.28.2.127 인용 PDF KSCI

A Study on the Comparison of Digital Speech Coding Performance (디지털 음성방식의 성능 비교에 대한 연구)

배철수
- The Journal of Korean Institute of Communications and Information Sciences
- /
- v.17 no.8
- /
- pp.881-890
- /
- 1992
Resonable speech quality assessment methodologies are required for speech quality assessment model which is used at speech system and communication network. There are objectlve measuies and subjective measures and subjective measure has the variousproblems in speech quality assessment methodologies. The objective of this study is to compare objective measures with subjective measure and obtain the objective measure as close as possible to subjective measure.
PDF

Assessment on the Speech Quality for Quantization Distortion (양자화 왜곡에 대한 음성품질 평가)

Kim, Jeong-Hwan
- Electronics and Telecommunications Trends
- /
- v.10 no.4 s.38
- /
- pp.129-142
- /
- 1995
본 고에서는, 음성을 디지털로 부호화하여 전송함으로써 발생되는 신호 대 양자화왜곡 비(Q)의 개념 및 CODEC과의 관계를 분석하고, MNRU를 디지털 회로로 구현하는데 필요한 입력음성 신호레벨, 잡음의 통계적 성질 및 진폭제한이 음성품질에 미치는 영향을 살펴보았다. 또한, 본 연구에서 구현한 MNRU의 성능에 대해 주관평가 실험을 실시하여, 다른 나라의 주관평가 결과와 비교/분석하였다.
https://doi.org/10.22648/ETRI.1995.J.100410 인용 PDF

음성총괄평가

정옥란
- Proceedings of the KSLP Conference
- /
- 1994.06a
- /
- pp.101-109
- /
- 1994
정상음성이란 개인의 음성 매개변수(vocal parameter), 즉 음도(pitch), 강도(loudness), 음질(quality), 유동성(flexibility) 등이 그 사람의 성, 연령, 환경, 체구 등에 적합한 음성을 말한다. 비정상적인 음성을 가진 음성자애 환자의 의뢰는 이비인후과 전문의에 의해 이루어지는 경우가 많고, 이 외에도 가족, 주변인, 환자의 교사 등에 의해 그리고 때때로 자가의뢰를 해오는 환자도 있다. (중략)
PDF

갑상선 수술 후 주관적 음성평가를 위한 설문지 유형 비교

Yun, Yeong-Seon;Son, Yeong-Ik
- Proceedings of the KSLP Conference
- /
- 2011.03a
- /
- pp.22-22
- /
- 2011
PDF

운율 분석용 DB 작성을 위한 자동 레이블러(Automatic labeler)의 성능 평가 및 유용성

강상훈;이항섭;김회린
- Proceedings of the KSPS conference
- /
- 1996.10a
- /
- pp.468-471
- /
- 1996
이 논문에서는 대량의 음성합성용 운율 DB를 용이하게 구축하기 위해 음성번역시스템을 이용한 자동 레이블러의 성능을 다양한 음성데이타를 대상으로 평가하였다. 실험 결과 FM radio news문장, 대화체 문장 및 낭독체 문장 등에는 레이블링 대상 음소의 약 80% 이상이 오류가 30msec 이내인 범위로 레이블링 되며, 고립단어에 대해서는 약 60%의 성능을 보여주고 있다. 현재 당 연구실에서는 자동 레이블러를 이용하여 합성용 운율 DB 및 합성단위를 작성하고 있으며. 자동 레이블러를 이용함으로서 일관성 있는 레이블링 결과를 얻을 수 있을 환 아니라 작성하는데 소요되는 시간도 줄일 수 있었다
PDF

Activities of Speech DB construction out of Countries (해외 음성 DB 구축 동향)

이용주
- Proceedings of the Acoustical Society of Korea Conference
- /
- 1995.06a
- /
- pp.253-260
- /
- 1995
음성정보처리 연구에 공통으로 이용 가능한 대량의 각종 음성 데이터를 수집, 편집, 배포하는 dfl은 연구 개발자의 입장에서는 분석, 합성, 인식등의 알고리즘 개발 평가에 이용 가능하며, 음성인식, 합성 시스템의 사용자 입장에서는 각종 시스템의 성능을 객관적으로 평가할 수 있다는 면에서 매우 중요하다. 본 논문에서는 국내 음성 DB 의 효율적인 구축을 위한 방안 도출에 참고하기 위하여 해외 각국의 구축 동향을 기관별, 형태별, 분야별로 구체적으로 정리하여 소개한다.
PDF

Spectrum Filter Algorithm based on Acoustic Model (음향학적 모델에 의한 스펙트럼 필터 알고리즘)

Choi, Jae-seung
- Proceedings of the Korean Institute of Information and Commucation Sciences Conference
- /
- 2016.10a
- /
- pp.770-772
- /
- 2016
본 논문에서는 음성신호처리 시스템에 유용하게 사용되는 음성신호의 특징 파라미터를 출력하는 스펙트럼 필터모델을 사용하여, 배경잡음 환경 하에서 음성신호 중의 잡음을 제거하는 알고리즘을 제안한다. 따라서 본 논문에서는 배경잡음을 제거할 때 고려해야 할 인간의 청각특성이 포함된 음성의 진폭 스펙트럼에 의한 청각필터의 특성을 도입한다. 본 논문의 실험에서 사용한 성능평가의 방법으로는 음절 명료도의 테스트에 적합한 주관적인 평가인 주파수 영역에서의 스펙트럼 왜곡률(Spectral Distortion, SD)을 사용하여 실험결과를 비교하고 고찰한다.
PDF

Review of Standard Sound Quality Assessment Methods for the Transmitted and Processed Sounds (음질 평가법의 표준과 연구 동향 - 전송 처리음 분야)

Oh, Wongeun
- The Journal of the Acoustical Society of Korea
- /
- v.32 no.3
- /
- pp.214-226
- /
- 2013
Assessing the quality of audio signals is an important consideration in making high quality sounds and various methods have been developed. This paper provides a general framework of sound quality and a technical overview of the international standard methods which are described in ITU-T, ITU-R, IEC and ANSI Recommendations in the speech intelligibility, speech quality, and audio quality areas. In addition, some recent findings and future works are included.
https://doi.org/10.7776/ASK.2013.32.3.214 인용 PDF KSCI

Search Result 1,638, Processing Time 0.035 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)