• Title/Summary/Keyword: word length

Search Result 229, Processing Time 0.023 seconds

Improvements on Phrase Breaks Prediction Using CRF (Conditional Random Fields) (CRF를 이용한 운율경계추성 성능개선)

  • Kim Seung-Won;Lee Geun-Bae;Kim Byeong-Chang
    • MALSORI
    • /
    • no.57
    • /
    • pp.139-152
    • /
    • 2006
  • In this paper, we present a phrase break prediction method using CRF(Conditional Random Fields), which has good performance at classification problems. The phrase break prediction problem was mapped into a classification problem in our research. We trained the CRF using the various linguistic features which was extracted from POS(Part Of Speech) tag, lexicon, length of word, and location of word in the sentences. Combined linguistic features were used in the experiments, and we could collect some linguistic features which generate good performance in the phrase break prediction. From the results of experiments, we can see that the proposed method shows improved performance on previous methods. Additionally, because the linguistic features are independent of each other in our research, the proposed method has higher flexibility than other methods.

  • PDF

Performance of Pseudomorpheme-Based Speech Recognition Units Obtained by Unsupervised Segmentation and Merging (비교사 분할 및 병합으로 구한 의사형태소 음성인식 단위의 성능)

  • Bang, Jeong-Uk;Kwon, Oh-Wook
    • Phonetics and Speech Sciences
    • /
    • v.6 no.3
    • /
    • pp.155-164
    • /
    • 2014
  • This paper proposes a new method to determine the recognition units for large vocabulary continuous speech recognition (LVCSR) in Korean by applying unsupervised segmentation and merging. In the proposed method, a text sentence is segmented into morphemes and position information is added to morphemes. Then submorpheme units are obtained by splitting the morpheme units through the maximization of posterior probability terms. The posterior probability terms are computed from the morpheme frequency distribution, the morpheme length distribution, and the morpheme frequency-of-frequency distribution. Finally, the recognition units are obtained by sequentially merging the submorpheme pair with the highest frequency. Computer experiments are conducted using a Korean LVCSR with a 100k word vocabulary and a trigram language model obtained by a 300 million eojeol (word phrase) corpus. The proposed method is shown to reduce the out-of-vocabulary rate to 1.8% and reduce the syllable error rate relatively by 14.0%.

The Acoustic Realization of Phrasal Verb vs. Verb-preposition (구절 동사와 전치사 수반동사의 의미에 따른 음성적 실현)

  • Kim, Hee-Sung;Song, Ji-Yeon;Kim, Kee-Ho
    • MALSORI
    • /
    • no.63
    • /
    • pp.67-84
    • /
    • 2007
  • Verb phrase could have two different meanings according to which is followed after verb; adverb or preposition. The meaning of 'verb+adverb' is deduced from a figurative meaning which is idiomatic expression, and 'verb+preposition' is interpreted as the literal meaning. The purpose of this study is to observe how English native speakers and Korean leaners of English distinguish two sentences of the same word strings with acoustic cues like pause and duration. According to the result, as pause was used for meaning distinction, it was likely that the pause length preceding prepositions was longer than that of following adverbs. To distinguish two sentences of the same word strings, all participants seemed to use pause, verb lengthening and adverb/preposition lengthening. Among them, there is a hierarchical significance; in sequence, pause, verb lengthening, adverb/preposition lengthening.

  • PDF

Word Similarity Calculation by Using the Edit Distance Metrics with Consonant Normalization

  • Kang, Seung-Shik
    • Journal of Information Processing Systems
    • /
    • v.11 no.4
    • /
    • pp.573-582
    • /
    • 2015
  • Edit distance metrics are widely used for many applications such as string comparison and spelling error corrections. Hamming distance is a metric for two equal length strings and Damerau-Levenshtein distance is a well-known metrics for making spelling corrections through string-to-string comparison. Previous distance metrics seems to be appropriate for alphabetic languages like English and European languages. However, the conventional edit distance criterion is not the best method for agglutinative languages like Korean. The reason is that two or more letter units make a Korean character, which is called as a syllable. This mechanism of syllable-based word construction in the Korean language causes an edit distance calculation to be inefficient. As such, we have explored a new edit distance method by using consonant normalization and the normalization factor.

A Study on vowel length of Korean monophthong (한국어의 세대별 음향 연구 -단순모음을 중심으로-)

  • Lee JaeKang
    • Proceedings of the Acoustical Society of Korea Conference
    • /
    • spring
    • /
    • pp.325-328
    • /
    • 2000
  • According to H.B.Lee(1993), standard Korean vowel qualities are as follows: in /i/, /e/, $/\epsilon/$, /a/, /o/, /w/, they have 4 qualities each other and in /er/ there are 3 qualities. The environments of 4 qualities are iong and stressed vowel in word initial, short and stressed vowel in word initial, unstressed vowel in word initial, unstressed vowel in word finial. The aim of this study is to seek and compare with H.B.Lee(1993). Conclusively I could not find on the whole any pattern of the same types of H.B.Lee(1993) in this study And especially in Fl vowel formant values of /er/and /w/, I never found any pattern of the same types of H.B.Lee(1993). Also F2 vowel formant values of $/\varepsilon/$ and /w/ do not have any kind of pattern of the same types of H.B.Lee(1993), between them, the patternize of F2 vowel formant values in /w / is especially difficult. It is the same story of Jaekang Lee(1998). But in some case, the patternize could be done. among the whole vowels, analysis environment b has the wide width on the change of the formant value. As the another result of the analysis It is to possible to make the pattern of the old male group. The old male group on the whole is analyzed to have the most low formant values and the old women group is analyzed to have the most high formants values, but in the most high formant valus there are young women group. And the formant values's rising in 2 cases of the formant value of /er/ is analyzed to have the same pattern of H.B.Lee(1993).

  • PDF

Comparison of Time and Frequency Resources of DFT-s-OFDM Systems Using the Zero-Tail and Unique Word (Zero Tail과 Unique Word를 사용하는 DFT-s-OFDM 시스템들의 시간과 주파수 자원 비교)

  • Kim, Byeongjae;Ryu, Heung-Gyoon
    • The Journal of Korean Institute of Communications and Information Sciences
    • /
    • v.41 no.12
    • /
    • pp.1715-1720
    • /
    • 2016
  • In the upcoming 5-generation mobile communication system, various techniques for improving the power efficiency and spectral efficiency have been proposed. 5G mobile communication system also have been studied a lot of multi-carrier-based modulation techniques like the 4G mobile communication system. In this paper, we analyzed the conventional system structure of the Zero-tail DFT-s-OFDM and UW (Unique Word) -DFT-s-OFDM system based on DFT-s-OFDM system in these techniques. UW and zero are added and used each system, and CP is removed. the result of quality of systems for simulation, OOB(Out of Band) power of Zero-tail DFT-s-OFDM and UW-DFT-s-OFDM use the less time resource as long as CP length, also both systems are reduced about 11dB than DFT-s-OFDM system. In these result, Zero-tail DFT-s-OFDM and UW-DFT-s-OFDM system are more effective than DFT-s-OFDM system.

Korean Head-Tail Tokenization and Part-of-Speech Tagging by using Deep Learning (딥러닝을 이용한 한국어 Head-Tail 토큰화 기법과 품사 태깅)

  • Kim, Jungmin;Kang, Seungshik;Kim, Hyeokman
    • IEMEK Journal of Embedded Systems and Applications
    • /
    • v.17 no.4
    • /
    • pp.199-208
    • /
    • 2022
  • Korean is an agglutinative language, and one or more morphemes are combined to form a single word. Part-of-speech tagging method separates each morpheme from a word and attaches a part-of-speech tag. In this study, we propose a new Korean part-of-speech tagging method based on the Head-Tail tokenization technique that divides a word into a lexical morpheme part and a grammatical morpheme part without decomposing compound words. In this method, the Head-Tail is divided by the syllable boundary without restoring irregular deformation or abbreviated syllables. Korean part-of-speech tagger was implemented using the Head-Tail tokenization and deep learning technique. In order to solve the problem that a large number of complex tags are generated due to the segmented tags and the tagging accuracy is low, we reduced the number of tags to a complex tag composed of large classification tags, and as a result, we improved the tagging accuracy. The performance of the Head-Tail part-of-speech tagger was experimented by using BERT, syllable bigram, and subword bigram embedding, and both syllable bigram and subword bigram embedding showed improvement in performance compared to general BERT. Part-of-speech tagging was performed by integrating the Head-Tail tokenization model and the simplified part-of-speech tagging model, achieving 98.99% word unit accuracy and 99.08% token unit accuracy. As a result of the experiment, it was found that the performance of part-of-speech tagging improved when the maximum token length was limited to twice the number of words.

Verification of the Usefulness of the Mock TOEIC Test using Corpus Indices : Focusing on the Analysis of Difficulty and Discrimination (코퍼스 지표를 활용한 모의 토익시험의 유용성 검증 : 난이도와 변별도 분석을 중심으로)

  • Lee, Yena
    • The Journal of the Korea Contents Association
    • /
    • v.21 no.10
    • /
    • pp.576-593
    • /
    • 2021
  • In this study, in order to investigate the factors that affect the percentage of correct answers and the degree of discrimination of the TOEIC test, a regression analysis was performed using corpus indicators that influence correct answer rate and the degree of discrimination for each part derived from the item analysis. The basic calculation word_length, consistency index LSA_overlap_adjacent_sentences, lexical diversity MTLD_VOCD, conjunction All_logical_causal_connectives_incidence, situational model casual_particles_causal_verbs_Ratio, syntactic complexity Left_embeddedness, and syntactic pattern density Infinitive_density were found to have negative effects. These factors that lower the correct answer rate can be utilized when setting learning goals. Vocabulary diversity index MTLD_VOCD, conjunction Additive_connectives_incidence, syntactic pattern density Infinitive_density, and lexical information person1_2_pronoun_incidence were found to have a positive effect. Factors influencing the increase in discrimination may provide important information for developing a learning program.

An Approach to Segmentation of Address Strings of unconstrained handwritten Hangul using Run-Length Code (Rum-Length code를 이용한 제약없이 쓰여진 한글 필기체 주소열 분할)

  • Kim, Gyeonghwan;Yoon, Jason-J
    • Journal of KIISE:Software and Applications
    • /
    • v.28 no.11
    • /
    • pp.813-821
    • /
    • 2001
  • While recognition of isolated units of writing, such as a character or a word, has been extensively studied, emphasis on the segmentation itself has been lacking. In this paper we propose an active segmentation method for handwritten Hangul address strings based on the Run-length code. A slant correction algorithm, which is considered as an important preprocessing step for the segmentation, is presented. Three fundamental candidate estimation functions are introduced to detect the clues on touching points, and the classification of touching types is attempted depending on the structural peculiarity of Hangul. Our experiments show segmentation performance of 88.2% on touching characters with minimal over-segmentation.

  • PDF

Implementation of Connected-Digit Recognition System Using Tree Structured Lexicon Model (트리 구조 어휘 사전을 이용한 연결 숫자음 인식 시스템의 구현)

  • Yun Young-Sun;Chae Yi-Geun
    • MALSORI
    • /
    • no.50
    • /
    • pp.123-137
    • /
    • 2004
  • In this paper, we consider the implementation of connected digit recognition system using tree structured lexicon model. To implement efficiently the fixed or variable length digit recognition system, finite state network (FSN) is required. We merge the word network algorithm that implements the FSN with lexical tree search algorithm that is used for general speech recognition system for fast search and large vocabulary systems. To find the efficient modeling of digit recognition system, we investigate some performance changes when the lexical tree search is applied.

  • PDF