• Title/Summary/Keyword: Similar Words

Search Result 612, Processing Time 0.022 seconds

Realtime Word Filtering System against Variations of Censored Words in Korean (변형된 한글 금칙어에 대한 실시간 필터링 시스템)

  • Kim, ChanWoo;Sung, Mee Young
    • Journal of Korea Multimedia Society
    • /
    • v.22 no.6
    • /
    • pp.695-705
    • /
    • 2019
  • The level of psychological damage caused by verbal abuse among cyberbully victims is very serious. It is going to introduce a system that determines the level of sanctions against chatting in real time using the automatic prohibited words filtering based on artificial neural network. In this paper, we propose a keyword filtering method that detects the modified prohibited words and determines whether the corresponding chat should be sanctioned in real time, and a real-time chatting screening system using it. The accuracy of filtering through machine learning was improved by processing data in advance through coding techniques that express consonants and vowels of similar pronunciation at close distances. After comparing and analyzing Mahalanobis-based clustering algorithms and artificial neural network-based algorithms, algorithms that utilize artificial neural networks showed high performance. If it is applied to Internet chatting, comments or online games, it is expected that it will be able to filter more effectively than the existing filtering method and that this will ease communication inconvenience due to existing indiscriminate filtering methods.

The Effect of Acoustic Correlates of Domain-initial Strengthening in Lexical Segmentation of English by Native Korean Listeners

  • Kim, Sa-Hyang;Cho, Tae-Hong
    • Phonetics and Speech Sciences
    • /
    • v.2 no.3
    • /
    • pp.115-124
    • /
    • 2010
  • The current study investigated the role of acoustic correlates of domain-initial strengthening in lexical segmentation of a non-native language. In a series of cross-modal identity-priming experiments, native Korean listeners heard English auditory stimuli and made lexical decision to visual targets (i.e., written words). The auditory stimuli contained critical two word sequences which created temporal lexical ambiguity (e.g., 'mill#company', with the competitor 'milk'). There was either an IP boundary or a word boundary between the two words in the critical sequences. The initial CV of the second word (e.g., [$k_{\Lambda}$] in 'company') was spliced from another token of the sequence in IP- or Wd-initial positions. The prime words were postboundary words (e.g., company) in Experiment 1, and preboundary words (e.g., mill) in Experiment 2. In both experiments, Korean listeners showed priming effects only in IP contexts, indicating that they can make use of IP boundary cues of English in lexical segmentation of English. The acoustic correlates of domain-initial strengthening were also exploited by Korean listeners, but significant effects were found only for the segmentation of postboundary words. The results therefore indicate that L2 listeners can make use of prosodically driven phonetic detail in lexical segmentation of L2, as long as the direction of those cues are similar in their L1 and L2. The exact use of the cues by Korean listeners was, however, different from that found with native English listeners in Cho, McQueen, and Cox (2007). The differential use of the prosodically driven phonetic cues by the native and non-native listeners are thus discussed.

  • PDF

A Study on the Readability of Elementary School Science Textbooks (초등학교 과학 교과서의 이독성 연구)

  • Koh, Han-Joong;Song, Jeong-Mee;Kang, Suk-Jin
    • Journal of Korean Elementary Science Education
    • /
    • v.29 no.2
    • /
    • pp.134-143
    • /
    • 2010
  • The purpose of this study is to devise a new method for examining the readabilities of textbooks and to compare the readabilities of elementary school science textbooks. Third and sixth grade science textbooks were compared in terms of word, sentence, and paragraph in this study. In the word analyses, criterion suggested by Kim (2003) who classified about 238,000 words into seven categories according to their educational importances was adopted. In this study, the words from 3rd and 6th grade science textbooks were classified into four categories, and then the kinds and frequencies of words in each category were investigated. In the sentence analyses, sentences were classified either a simple sentence or a compound/complex sentence, and the ratios of each type were calculated. The average number of words in a sentence was also calculated in the sentence analyses. The ratios of conjunctions and demonstratives were examined in the paragraph analyses. The results indicated that both the kinds and frequencies of words in 3rd grade science textbook were smaller than those of 6th grade one. However, both science textbooks were similar in the distributions of words across the four categories. The ratio of simple sentences in 3rd grade science textbook was higher than that of 6th grade one, and the length of a sentence in 3rd grade science textbook was also shorter than that of 6th grade one. Both the ratios of conjunctions and demonstratives in 3rd grade science textbook were lower than those of 6th grade one.

  • PDF

Development of Similar Bibliographic Retrieval System based on Neighboring Words and Keyword Topic Information (인접한 단어와 키워드 주제어 정보에 기반한 유사 문헌 검색 시스템 개발)

  • Kim, Kwang-Young;Kwak, Seung-Jin
    • Journal of Korean Library and Information Science Society
    • /
    • v.40 no.3
    • /
    • pp.367-387
    • /
    • 2009
  • The similar bibliographic retrieval system follows whether it selects a thing of the extracted index term and or not the difference in which the similar document retrieval system There be many in the search result is generated. In this research, the method minimally making the error of the selection of the extracted candidate index term is provided In this research, the word information in which it is adjacent by using candidate index terms extracted from the similar literature and the keyword topic information were used. And by using the related author information and the reranking method of the search result, the similar bibliographic system in which an accuracy is high was developed. In this paper, we conducted experiments for similar bibliographic retrieval system on a collection of Korean journal articles of science and technology arena. The performance of similar bibliographic retrieval system was proved through an experiment and user evaluation.

  • PDF

A Study on the Textile Terminologies of the Chosun Period (朝鮮時代 服飾用語 硏究II-織物關聯用語를 中心으로-)

  • 김진구
    • The Research Journal of the Costume Culture
    • /
    • v.9 no.3
    • /
    • pp.532-536
    • /
    • 2001
  • This study is concerned with the textile related terminologies of the Chosun period. The purpose of this study was to trace and to examine some textile related terms such as goro, mooruwi, modan, shiok, jal, gaam, and chien. These words were examined and analyzed in terms of the origins, meanings, and neighbouring languages. The results of this research can be summarized as follows: The results of this study revealed that the word goro of the Chosun period was derived from the Chinese ku lo 羅 or (Equations. See Full-text). Korean goro or goroi is a transliteration of the Chinese moolo 霧羅. The word modan 帽緞 was a kind of rich silk fabric. Manchurian kamku 帽緞 was derived from Arabic word kamkha. The word shiok, shiok, shiuk, shiurk, or shiu 시으 means felt in Korean. Similar words to Korean shiok was found in Afro-Asiatic family such as Egyptian, Hebrew, and Assyrians. Egyptian shiu means a seep or a goat. The word jal meaning black sable was found was originated in the Chinese tzuerl 子兒皮, black sable. The word Korean gaam 가암, 가음, was similar to Mongorian k∂m meaning a material. Also Iraq-Arabian xaam meaning raw, unworked, unprocessed, had the same meaning as the Korean gaam. Xaam and gaam have almost the same phonetical sounds. The Korean gaam was derived from the xaam of Iraq-Arabian. Korean chien meaning cloth was derived from the Chinese chyan or chien (Equations. See Full-text).

  • PDF

The final stop consonant perception in typically developing children aged 4 to 6 years and adults (4-6세 정상발달아동 및 성인의 종성파열음 지각력 비교)

  • Byeon, Kyeongeun;Ha, Seunghee
    • Phonetics and Speech Sciences
    • /
    • v.7 no.1
    • /
    • pp.57-65
    • /
    • 2015
  • This study aimed to identify the development pattern of final stop consonant perception using the gating task. Sixty-four subjects participated in the study: 16 children aged 4 years, 16 children aged 5 years, 17 children aged 6 years, and 15 adults. One-syllable words with consonant-vowel-consonant(CVC) structure, mokㄱ-motㄱ and papㄱ-patㄱ were used as stimuli in order to remove the redundancy of acoustic cues in stimulus words, 40ms-length (-40ms) and 60ms-length (-60ms) from the entire duration of the final consonant were deleted. Three conditions (the whole word segment, -40ms, -60ms) were used for this speech perception experiment. 48 tokens (4 stimuli ${\times}3$ conditions ${\times}4$ trials) in total were provided for participants. The results indicated that 5 and 6 year olds showed final consonant perception similar to adults in stimuli, papㄱ-patㄱ and only the 6-year-old children showed perception similar to adults in stimuli, 'mokㄱ-motㄱ. The results suggested that younger typically developing children require more acoustic information to accurately perceive final consonants than older children and adults. Final consonant perception ability may become adult-like around 6 years old. The study provides fundamental data on the development pattern of speech perception in normal developing children, which can be used to compare to those of children with communication disorders.

Automatic extraction of similar poetry for study of literary texts: An experiment on Hindi poetry

  • Prakash, Amit;Singh, Niraj Kumar;Saha, Sujan Kumar
    • ETRI Journal
    • /
    • v.44 no.3
    • /
    • pp.413-425
    • /
    • 2022
  • The study of literary texts is one of the earliest disciplines practiced around the globe. Poetry is artistic writing in which words are carefully chosen and arranged for their meaning, sound, and rhythm. Poetry usually has a broad and profound sense that makes it difficult to be interpreted even by humans. The essence of poetry is Rasa, which signifies mood or emotion. In this paper, we propose a poetry classification-based approach to automatically extract similar poems from a repository. Specifically, we perform a novel Rasa-based classification of Hindi poetry. For the task, we primarily used lexical features in a bag-of-words model trained using the support vector machine classifier. In the model, we employed Hindi WordNet, Latent Semantic Indexing, and Word2Vec-based neural word embedding. To extract the rich feature vectors, we prepared a repository containing 37 717 poems collected from various sources. We evaluated the performance of the system on a manually constructed dataset containing 945 Hindi poems. Experimental results demonstrated that the proposed model attained satisfactory performance.

Researcher and Research Area Recommendation System for Promoting Convergence Research Using Text Mining and Messenger UI (텍스트 마이닝 방법론과 메신저UI를 활용한 융합연구 촉진을 위한 연구자 및 연구 분야 추천 시스템의 제안)

  • Yang, Nak-Yeong;Kim, Sung-Geun;Kang, Ju-Young
    • The Journal of Information Systems
    • /
    • v.27 no.4
    • /
    • pp.71-96
    • /
    • 2018
  • Purpose Recently, social interest in the convergence research is at its peak. However, contrary to the keen interest in convergence research, an infrastructure that makes it easier to recruit researchers from other fields is not yet well established, which is why researchers are having considerable difficulty in carrying out real convergence research. In this study, we implemented a researcher recommendation system that helps researchers who want to collaborate easily recruit researchers from other fields, and we expect it to serve as a springboard for growth in the convergence research field. Design/methodology/approach In this study, we implemented a system that recommends proper researchers when users enter keyword in the field of research that they want to collaborate using word embedding techniques, word2vec. In addition, we also implemented function of keyword suggestions by using keywords drawn from LDA Topicmodeling Algorithm. Finally, the UI of the researcher recommendation system was completed by utilizing the collaborative messenger Slack to facilitate immediate exchange of information with the recommended researchers and to accommodate various applications for collaboration. Findings In this study, we validated the completed researcher recommendation system by ensuring that the list of researchers recommended by entering a specific keyword is accurate and that words learned as a similar word with a particular researcher match the researcher's field of research. The results showed 85.89% accuracy in the former, and in the latter case, mostly, the words drawn as similar words were found to match the researcher's field of research, leading to excellent performance of the researcher recommendation system.

Implementation of A Plagiarism Detecting System with Sentence and Syntactic Word Similarities (문장 및 어절 유사도를 이용한 표절 탐지 시스템 구현)

  • Maeng, Joosoo;Park, Ji Su;Shon, Jin Gon
    • KIPS Transactions on Software and Data Engineering
    • /
    • v.8 no.3
    • /
    • pp.109-114
    • /
    • 2019
  • The similarity detecting method that is basically used in most plagiarism detecting systems is to use the frequency of shared words based on morphological analysis. However, this method has limitations on detecting accurate degree of similarity, especially when similar words concerning the same topics are used, sentences are partially separately excerpted, or postpositions and endings of words are similar. In order to overcome this problem, we have designed and implemented a plagiarism detecting system that provides more reliable similarity information by measuring sentence similarity and syntactic word similarity in addition to the conventional word similarity. We have carried out a comparison of on our system with a conventional system using only word similarity. The comparative experiment has shown that our system can detect plagiarized document that the conventional system can detect or cannot.

Research on Comparing System with Syntactic-Semantic Tree in Subjective-type Grading (주관식 문제 채점에서의 구문의미트리 비교 시스템에 대한 연구)

  • Kang, WonSeog
    • The Journal of Korean Association of Computer Education
    • /
    • v.20 no.5
    • /
    • pp.79-88
    • /
    • 2017
  • To upgrade the subjective question grading, we need the syntactic-semantic analysis to analyze syntatic-semantic relation between words in answering. However, since the syntactic-semantic tree has structural and semantic relation between words, we can not apply the method calculating the similarity between vectors. This paper suggests the comparing system with syntactic-semantic tree which has structural and semantic relation between words. In this thesis, we suggest similarity calculation principles for comparing the trees and verify the principles through experiments. This system will help the subjective question grading by comparing the trees and be utilized in distinguishing similar documents.