• Title/Summary/Keyword: trigram

Search Result 39, Processing Time 0.023 seconds

Performance of speech recognition unit considering morphological pronunciation variation (형태소 발음변이를 고려한 음성인식 단위의 성능)

  • Bang, Jeong-Uk;Kim, Sang-Hun;Kwon, Oh-Wook
    • Phonetics and Speech Sciences
    • /
    • v.10 no.4
    • /
    • pp.111-119
    • /
    • 2018
  • This paper proposes a method to improve speech recognition performance by extracting various pronunciations of the pseudo-morpheme unit from an eojeol unit corpus and generating a new recognition unit considering pronunciation variations. In the proposed method, we first align the pronunciation of the eojeol units and the pseudo-morpheme units, and then expand the pronunciation dictionary by extracting the new pronunciations of the pseudo-morpheme units at the pronunciation of the eojeol units. Then, we propose a new recognition unit that relies on pronunciation by tagging the obtained phoneme symbols according to the pseudo-morpheme units. The proposed units and their extended pronunciations are incorporated into the lexicon and language model of the speech recognizer. Experiments for performance evaluation are performed using the Korean speech recognizer with a trigram language model obtained by a 100 million pseudo-morpheme corpus and an acoustic model trained by a multi-genre broadcast speech data of 445 hours. The proposed method is shown to reduce the word error rate relatively by 13.8% in the news-genre evaluation data and by 4.5% in the total evaluation data.

Computational Analysis of Neighboring Genes on Arabidopsis thaliana Chromosomes 4 and 5: Their Genomic Association as Functional Subunits

  • Goh, Sung-Ho;Kim, Tae-Hyung;Kim, Jee-Hyub;Nam, DouGu;Choi, Doil;Hur, Cheol-Goo
    • Genomics & Informatics
    • /
    • v.1 no.1
    • /
    • pp.40-49
    • /
    • 2003
  • The genes related to specific events or pathways in bacteria are frequently localized proximate to the genome of their neighbors, as with the structures known as operon, but eukaryotic genes seem to be independent of their neighbors, and are dispersed randomly throughout genomes. Although cases are rare, the findings from structures similar to prokaryotic operons in the nematode genome, and the clustering of housekeeping genes on human genome, lead us to assess the genomic association of genes as functional subunits. We evaluated the genomic association of neighboring genes on chromosomes 4 and 5 of Arabidopsis thaliana with and without respectively consideration of the scaffold/matrix­attached regions (S/MAR) loci. The observed number of functionally identical bigrams and trig rams were significantly higher than expected, and these results were verified statistically by calculating p-values for weighted random distributions. The observed frequency of functionally identical big rams and trig rams were much higher in chromosome 4 than in chromosome 5, but the frequencies with, and without, consideration of the S/MAR in each chromosome were similar. In this study, a genomic association among functionally related neighboring genes in Arabidopsis thaliana was suggested.

Feature Weighting for Opinion Classification of Comments on News Articles (뉴스 댓글의 감정 분류를 위한 자질 가중치 설정)

  • Lee, Kong-Joo;Kim, Jae-Hoon;Seo, Hyung-Won;Rhyu, Keel-Soo
    • Journal of Advanced Marine Engineering and Technology
    • /
    • v.34 no.6
    • /
    • pp.871-879
    • /
    • 2010
  • In this paper, we present a system that classifies comments on a news article into a user opinion called a polarity (positive or negative). The system is a kind of document classification system for comments and is based on machine learning techniques like support vector machine. Unlike normal documents, comments have their body that can influence classifying their opinions as polarities. In this paper, we propose a feature weighting scheme using such characteristics of comments and several resources for opinion classification. Through our experiments, the weighting scheme have turned out to be useful for opinion classification in comments on Korean news articles. Also Korean character n-grams (bigram or trigram) have been revealed to be helpful for opinion classification in comments including lots of Internet words or typos. In the future, we will apply this scheme to opinion analysis of comments of product reviews as well as news articles.

The Selection of House Site and Its Architectural Expression in the Chosun Dynasty : A Case Study of Confucianist Lee-sik's Taegpoongdang in Yangpyung, Kyungki-do (조선 중기 유가(儒家)의 세계관이 반영된 집터 선정과 건축적 표현 -양평군 소재 택당 이식의 택풍당을 중심으로-)

  • Sung, Dong-Hwan;Cho, In-Chul
    • Journal of the Korean association of regional geographers
    • /
    • v.11 no.3
    • /
    • pp.367-380
    • /
    • 2005
  • This paper aims to investigate the characteristics of house site selection and its expression of building through manuscript of Taegdanggip which was authored by Lee-sik in the middle of Chosun dynasty. Its results are summarized in the following. Firstly, as a Confucianist, Lee-sik selected his ancestor's grave site as well as his house site by means of divination sign. And then he interpreted the characteristics of the location from feng-shui perspective. Secondly, he built Taepoongdang(literally 'pond and wind house') as his house for retirement based on a trigram from the Book of Changes. He reflected the divination sign in consturcting his house Taekpoongdang. Finally, the location of Taekpoongdang and Baekagog village was well suitable to feng-shui theory.

  • PDF

Part-Of-Speech Tagging using multiple sources of statistical data (이종의 통계정보를 이용한 품사 부착 기법)

  • Cho, Seh-Yeong
    • Journal of the Korean Institute of Intelligent Systems
    • /
    • v.18 no.4
    • /
    • pp.501-506
    • /
    • 2008
  • Statistical POS tagging is prone to error, because of the inherent limitations of statistical data, especially single source of data. Therefore it is widely agreed that the possibility of further enhancement lies in exploiting various knowledge sources. However these data sources are bound to be inconsistent to each other. This paper shows the possibility of using maximum entropy model to Korean language POS tagging. We use as the knowledge sources n-gram data and trigger pair data. We show how perplexity measure varies when two knowledge sources are combined using maximum entropy method. The experiment used a trigram model which produced 94.9% accuracy using Hidden Markov Model, and showed increase to 95.6% when combined with trigger pair data using Maximum Entropy method. This clearly shows possibility of further enhancement when various knowledge sources are developed and combined using ME method.

A Study on Performance Evaluation of Hidden Markov Network Speech Recognition System (Hidden Markov Network 음성인식 시스템의 성능평가에 관한 연구)

  • 오세진;김광동;노덕규;위석오;송민규;정현열
    • Journal of the Institute of Convergence Signal Processing
    • /
    • v.4 no.4
    • /
    • pp.30-39
    • /
    • 2003
  • In this paper, we carried out the performance evaluation of HM-Net(Hidden Markov Network) speech recognition system for Korean speech databases. We adopted to construct acoustic models using the HM-Nets modified by HMMs(Hidden Markov Models), which are widely used as the statistical modeling methods. HM-Nets are carried out the state splitting for contextual and temporal domain by PDT-SSS(Phonetic Decision Tree-based Successive State Splitting) algorithm, which is modified the original SSS algorithm. Especially it adopted the phonetic decision tree to effectively express the context information not appear in training speech data on contextual domain state splitting. In case of temporal domain state splitting, to effectively represent information of each phoneme maintenance in the state splitting is carried out, and then the optimal model network of triphone types are constructed by in the parameter. Speech recognition was performed using the one-pass Viterbi beam search algorithm with phone-pair/word-pair grammar for phoneme/word recognition, respectively and using the multi-pass search algorithm with n-gram language models for sentence recognition. The tree-structured lexicon was used in order to decrease the number of nodes by sharing the same prefixes among words. In this paper, the performance evaluation of HM-Net speech recognition system is carried out for various recognition conditions. Through the experiments, we verified that it has very superior recognition performance compared with the previous introduced recognition system.

  • PDF

Effect of syllable complexity on the visual span of Korean Hangul reading and its relation to reading abilities (한글 글자 유형이 시각 폭과 읽기 능력에 미치는 영향)

  • Choi, Youngon;Kim, Tae Hoon
    • Korean Journal of Cognitive Science
    • /
    • v.27 no.2
    • /
    • pp.325-353
    • /
    • 2016
  • The visual span refers to the number of letters that can be accurately recognized without moving one's eyes. The size of the visual span is affected by sensory factors such as perimetric complexity, crowding, and mislocation of letters. Korean Hangul utilizes rather unique alphabetic-syllabary writing system, quite different from English and Chinese writing systems. Due to this combinatorial nature of the script, the visual span for Hangul characters can also be affected by the letter type (e.g., CV vs CVCC). The present study examined the effect of syllable complexity on the visual span for Hangul by comparing letter recognition accuracy across four letter type conditions (C only, CV, CVC, and CVCC). We also aimed to determine the meaningful letter type(s) that is associated with differences in reading abilities in Korean. Using a trigram presentation method, we found that overall recognition accuracy declined as syllable complexity increased. However, the visual span for CVC type was greater than that for CV type, suggesting that the effect is not necessarily linear, and that there might be other factors affecting the visual span for these types of letters. C and CV type showed fairly strong positive correlations with reading comprehension, suggesting that these might be the meaningful units for measuring visual span in relating to reading abilities.

Noju Oh Hui-sang's ConfucianismDoctrine and its Characteristics (노주(老洲) 오희상(吳熙常)의 경설(經說)과 그 특징(特徵))

  • Kim, Young-ho
    • The Journal of Korean Philosophical History
    • /
    • no.38
    • /
    • pp.129-162
    • /
    • 2013
  • Noju Oh Hui-sang was a Confucian who was active during the reign of King Sunjo in late Joseon Dynasty and he also was a master of the Sallim faction. Though he is known as an eclectic Neo-Confucian, he had profound knowledge in the study of Confucian classics as well through succeeding the family study handed down by his father Oh Jae-sun and his oldest brother Oh Yun-sang. This thesis hereby examines Noju's Confucianism doctrine and its characteristics. Noju's Confucianism doctrine is characterized significantly with the following aspects. First, its analyses are detailed overall and it annotates chapters and verses mostly related to Neo-Confucian theories on interpretation of the Confucian classics. Second, it conducts in-depth study not only on Chu Hsi's annotation but also on the small commentaries (小注) in Compendium of the Commentaries on Four Chinese Classics (四書集註大全). In terms of Chu Hsi's theory, however, Noju interprets Confucian classics while supplementing shortcomings on Chu Hsi's theory rather than opposing it. For opinions of all philosophers and scholars on small commentaries, it expresses rather critical theories than supporting ones. Third, it quotes many theories not only of Chinese Confucians but also of Korean ones. It mainly introduces theories of Namdang Han Won-jin, including those of Yi Yulgok. Among them, it particularly has frequent quotations from Han Won-jin's Kyoungyigimunrok (經義記聞錄). Fourth, Noju actively acknowledges senior Confucians' theories many times in quoting them but he also daringly points out their errors when a theory is thought not to be appropriate. He indicates errors one by one in theories not only of Uam and Yulgok but even of Mencius. Fifth, it especially discusses Book of Changes (周易) in depth. It tends to criticize Chengzi's I-Chuan (易傳) but accept Chu Hsi's Benyi (本義). It roughly explains Book of Changes in general but seldom directly accounts for trigrams of it other than Qian trigram and it has detailed explanation especially on Xicizhuan (繫辭傳).

The Concept of 'the Former World and the Later World' in Daesoon Thought as Introduced via the Diagrams of The Comprehensive Mirror of Taegeukdo (『태극도통감』의 도상을 통해 본 대순사상의 '선·후천' 개념)

  • Lee Bong-ho
    • Journal of the Daesoon Academy of Sciences
    • /
    • v.47
    • /
    • pp.65-103
    • /
    • 2023
  • In The Canonical Scripture (典經), the core scripture of Daesoon Thought, the Former World and the Later World are divided into the Era of Mutual Contention and the Era of Mutual Beneficence. This concept of the Former World and the Later World appears in diagrams on I-Ching Studies (易學) in the text titled, The Comprehensive Mirror of Taegeukdo (太極道通鑑). In I-Ching Studies, Anterior Heaven (先天) and Posterior Heaven (後天) are the main concepts in Song Dynasty diagram books on I-Ching Studies. Among the diagrams of I-Ching Studies, Fuxi's Diagram of the Sequence of the Eight Trigrams, Fuxi's Diagram of the Positions of the Eight Trigrams, Fuxi's Diagram of the Sequence of the 64 Hexagrams, and Fuxi's Diagram of the Positions of the 64 Hexagrams correspond to the Anterior Heaven, and King Wen's Diagram of the Sequence of the Eight Trigrams and King Wen's Diagram of the Positions of the Eight Trigrams correspond to Posterior Heaven. In The Comprehensive Mirror of Taegeukdo, the diagrams of I-Ching Studies are reinterpreted according to Daesoon Thought. The Diagram of the Eight Trigrams of King Wen's Era corresponds to King Wen's Diagram of the Eight Trigrams in I-Ching Studies. This diagram was drawn according to the text in Chapter Five of the Treatise of Remarks on the Trigrams. This diagram corresponds to "the Era of the Nobility of Earth (地尊時代)" centered on the trigram kun (坤 / ☷ ground). Fuxi's Diagram of the Positions of the Eight Trigrams in I-Ching Studies corresponds to The Diagram of the Positions of the Eight Trigrams of Fuxi's Era in Daesoon Thought. The most significant feature of this diagram is that the trigrams assigned to the directions of north and south match the hexagram indicating the obstruction of Heaven and Earth. This is hexagram 12 (否), meaning "obstruction," and it depicts no exchange or communication between Yin and Yang. Naturally, this symbolizes mutual destruction overtaking Yin and Yang. Daesoon Thought expresses this as "the Era of the Nobility of Heaven (天尊時代)." The most significant feature of The Diagram of the Eight Trigrams of the Corrected Book of Changes in The Comprehensive Mirror of Taegeukdo is that the trigrams assigned to the directions of south and north are indicative of hexagram 11, Peace on Earth and in Heaven (泰). This is a diagram in which mutual destruction is resolved through the Five Phases because the trigrams for water (坎 / ☵) and fire (離 / ☲) are in a corrected orientation. Therefore, this diagram symbolizes a world "free from Mutual Contention" and "the Era of Human Nobility (人尊時代)." According to the contents of The Canonical Scripture, the Supreme God performed the Reordering Works of the Three Realms to correct the Mutual Contention of the Former World, and as a result, the Mutual Contention of the Former World will give way to the implementation of the Dao of Mutual Beneficence. The Supreme God's Reordering Works of the Three Realms have been completed in the realm of divine beings, but in the Later World, they appear as an Earthly Paradise where the Dao of Mutual Beneficence is realized. The diagram depicting the Later World is The Diagram of the Eight Trigrams of the Era of the Corrected Book of Changes in The Comprehensive Mirror of Taegeukdo.