• Title/Summary/Keyword: 어휘정보

Search Result 1,062, Processing Time 0.031 seconds

A Method of Identifying Ownership of Personal Information exposed in Social Network Service (소셜 네트워크 서비스에 노출된 개인정보의 소유자 식별 방법)

  • Kim, Seok-Hyun;Cho, Jin-Man;Jin, Seung-Hun;Choi, Dae-Seon
    • Journal of the Korea Institute of Information Security & Cryptology
    • /
    • v.23 no.6
    • /
    • pp.1103-1110
    • /
    • 2013
  • This paper proposes a method of identifying ownership of personal information in Social Network Service. In detail, the proposed method automatically decides whether any location information mentioned in twitter indicates the publisher's residence area. Identifying ownership of personal information is necessary part of evaluating risk of opened personal information online. The proposed method uses a set of decision rules that considers 13 features that are lexicographic and syntactic characteristics of the tweet sentences. In an experiment using real twitter data, the proposed method shows better performance (f1-score: 0.876) than the conventional document classification models such as naive bayesian that uses n-gram as a feature set.

Semantic Information Retrieval Based on User-Word Intelligent Network (U-WIN 기반의 의미적 정보검색 기술)

  • Im, Ji-Hui;Choi, Ho-Seop;Ock, Cheol-Young
    • Proceedings of the Korea Contents Association Conference
    • /
    • 2006.11a
    • /
    • pp.547-550
    • /
    • 2006
  • The criterion which judges an information retrieval system performance is to how many accurately retrieve an information that the user wants. The search result which uses only homograph has been appears the various documents that relates to each meaning of the word or intensively appears the documents that relates to specific meaning of it. So in this paper, we suggest semantic information retrieval technique using relation within User-Word Intelligent Network(U-WIN) to solve a disambiguation of query In our experiment, queries divide into two classes, the homograph used in terminology and the general homograph, and it sets the expansion query forms at "query + hypemym". Thus we found that only web document search's precision is average 73.5% and integrated search's precision is average 70% in two portal site. It means that U-WIN-Based semantic information retrieval technique can be used efficiently for a IR system.

  • PDF

An Investigation of Information Usefulness of Google Scholar in Comparison with Web of Science (Google Scholar의 학술정보 검색을 위한 정보 유용성 비교연구)

  • Kim, Hyunjung
    • Journal of the Korean BIBLIA Society for library and Information Science
    • /
    • v.25 no.3
    • /
    • pp.215-234
    • /
    • 2014
  • The purpose of this study is to investigate whether Google Scholar (GS) can substitute Web of Science (WoS) for those who don't have access to the subscription-based indexing service and if users feel GS is useful for scholarly information. To achieve the research purpose, the study evaluates both quantitative and qualitative aspects of the two databases. The major results through statistical analysis show that GS indexes much more records and citations for LIS journals than WoS(p < .01), but users' feedback about GS is not better than those about WoS.

Improving Performance of Search Engine Using Category based Evaluation (범주 기반 평가를 이용한 검색시스템의 성능 향상)

  • Kim, Hyung-Il;Yoon, Hyun-Nim
    • The Journal of the Korea Contents Association
    • /
    • v.13 no.1
    • /
    • pp.19-29
    • /
    • 2013
  • In the current Internet environment where there is high space complexity of information, search engines aim to provide accurate information that users want. But content-based method adopted by most of search engines cannot be used as an effective tool in the current Internet environment. As content-based method gives different weights to each web page using morphological characteristics of vocabulary, the method has its drawbacks of not being effective in distinguishing each web page. To resolve this problem and provide useful information to the users, this paper proposes an evaluation method based on categories. Category-based evaluation method is to extend query to semantic relations and measure the similarity to web pages. In applying weighting to web pages, category-based evaluation method utilizes user response to web page retrieval and categories of query and thus better distinguish web pages. The method proposed in this paper has the advantage of being able to effectively provide the information users want through search engines and the utility of category-based evaluation technique has been confirmed through various experiments.

Judgment about the Usefulness of Automatically Extracted Temporal Information from News Articles for Event Detection and Tracking (사건 탐지 및 추적을 위해 신문기사에서 자동 추출된 시간정보의 유용성 판단)

  • Kim Pyung;Myaeng Sung-Hyon
    • Journal of KIISE:Software and Applications
    • /
    • v.33 no.6
    • /
    • pp.564-573
    • /
    • 2006
  • Temporal information plays an important role in natural language processing (NLP) applications such as information extraction, discourse analysis, automatic summarization, and question-answering. In the topic detection and tracking (TDT) area, the temporal information often used is the publication date of a message, which is readily available but limited in its usefulness. We developed a relatively simple NLP method of extracting temporal information from Korean news articles, with the goal of improving performance of TDT tasks. To extract temporal information, we make use of finite state automata and a lexicon containing time-revealing vocabulary. Extracted information is converted into a canonicalized representation of a time point or a time duration. We first evaluated the extraction and canonicalization methods for their accuracy and investigated on the extent to which temporal information extracted as such can help TDT tasks. The experimental results show that time information extracted from text indeed helps improve both precision and recall significantly.

Development of Evaluation Tool for Educational Applications (교육용 앱 평가도구 개발 연구)

  • Lee, Jeong-Sook;Kim, Sung-Wan
    • Proceedings of the Korean Society of Computer Information Conference
    • /
    • 2013.01a
    • /
    • pp.149-152
    • /
    • 2013
  • 이 연구는 스마트교육환경에서의 교육용 앱을 평가하기 위한 신뢰롭고 타당한 도구를 개발하는데 있다. 기존 선행연구에 기초해서 교육용 앱의 평가를 위한 평가모형을 도출했으며, 이 모형은 4개의 평가영역(교수 학습측면, 화면디자인측면, 기술측면, 경제 윤리측면)과 13개 평가요소들로 구성되었다. 이 잠재모형의 통계적 타당성 검증을 위한 자료수집을 하고자, 경기도 소재 중학교 1곳과 고등학교 2곳의 학생 156명을 대상으로 교육용 앱을 평가하는데 있어서 각 평가문항이 갖는 중요도를 평가하였다. 수집된 자료를 탐색적 요인분석한 결과, 교육용 앱을 평가하는 영역으로 교수 학습(흥미성 자기주도성 실용성, 인지발달성), 화면디자인(디자인의 적합성, 어휘의 정확성), 기술(호환성, 안정성), 경제 윤리(경제성, 윤리성) 등 4개 영역이 제안되었다. 또한 문항내적일관성을 확인하고자 신뢰도 분석한 결과, 각 평가영역 별 Cronbach ${\alpha}$는 .88, .85, .82, .80으로 모두 적합한 수준을 보였다. 따라서, 이 연구를 통해 도출된 교육용 앱 평가도구는 통계적으로나 타당성과 신뢰성 측면에서 의미 있는 것으로 판단할 수 있다.

  • PDF

Analyzing and Extracting Relations between Topic Keywords Based on Word Formation (조어 중심적 주제어간 관계 추출 및 분석)

  • Jung, Han-Min;Lee, Mi-Kyoung;Sung, Won-Kyung
    • Proceedings of the Korean Society for Language and Information Conference
    • /
    • 2008.06a
    • /
    • pp.166-171
    • /
    • 2008
  • 본 연구는 기존에 잘 알려지고 널리 사용되고 있는 어휘 의미망이나 시소러스를 활용하기 어려운 과학 기술 분야, 특히 IT 분야에서 대용량 용어간 관계를 빠른 시간 내에 구축하여 검색 브라우징, 내비게이션 용도로 활용하는 것을 목표로 한다. 시소러스 구축 절차를 따르는 경우에 분야 전문가에 의한 정교한 작업과 고비용을 필요로 하여 충분한 구축 크기를 확보하는 것에 현실적인 어려움이 있다. 시소러스 자동 구축 방법론을 사용하는 경우에도 해당 용어들이 출현하는 방대한 말뭉치를 확보해야 하며 관계 구축 결과에 대한 직관적 이해가 쉽지 않다는 단점이 있다. 본 연구는 해외 학술 논문 말뭉치와 메타데이터에서 획득한 37만 여 주제어들을 이용하여 상 하위 관계, 관련어, 형제 관계를 추출하기 위해 조어적 기준에 근거한 규칙들을 이용한다. 이들 규칙을 이용하여 추출한 관계 수는 상 하위 관계 60여 만 개, 관련어 640여 만 개, 형제 관계 2,000여 만 개 등이다. 또한, 추출 결과 중 일부를 수작업으로 분석하여 단순한 추출 규칙에서 발생하는 오류 유형을 찾아내고 향후 과제에서 해결할 수 있는 방안에 대해 논하자고 한다.

  • PDF

A Study On Generation and Reduction of the Notation Candidate for the Notation Restoration of Korean Phonetic Value (한국어 음가의 표기 복원을 위한 표기 후보 생성 및 감소에 관한 연구)

  • Rhee, Sang-Burm;Park, Sung-Hyun
    • The KIPS Transactions:PartB
    • /
    • v.11B no.1
    • /
    • pp.99-106
    • /
    • 2004
  • The syllable restoration is a process restoring a phonetic value recognized in a speech recognition device with the notation form that a vocalization is former. In this paper a syllable restoration rule was composed of a based on standard pronunciation for a syllable restoration process. A syllable restoring regulation was used, and a generation method of a notation candidate set was researched. Also, A study is held to reduce the number of created notation candidate. Three phases of reduction processes were suggested. Reduction of a notation candidate has the non-notation syllable, non-vocabulary syllable and non-stem syllable. As a result of experiment, an average of 74% notation candidate decrease rates were shown.

Bibliography and the Cenventional Chinese Catalogue - Emphasis on the period prior to the Opium War- ('Bibliography'의 어휘와 '중국재래의 목록학' -특히 아편전쟁이전을 중심으로-)

  • Shim Woo-choon
    • Journal of the Korean Society for Library and Information Science
    • /
    • v.4
    • /
    • pp.27-42
    • /
    • 1975
  • Usage and scope of the word Bibliography in comparison with in conventional Chinese Catalogue (中國 在來 目錄學) (1) Usage of the word in connection with the study of books in the West has been changed from 'writing of books' (17th century) to the meaning of 'study of a book as an object'(l8th century), and this meaning of the 18th century has been transmitted up to the present. (2) In its scope, 14 branches(eight in physical aspect, six in content of books) were set up independently for the study of a book as an object. On the other hand, the term Textual Bibliography(校수學) was in use in China before the Opium War, however the word Catalogue (目錄學) has been a current word for the subject study as in the case of Bibliography in the West. And in the scope of study of a book as an object, although some of its aspect is somewhat similar to the Occidental Bibliorgraphy, the main stream of learning is regregarded as the root and the physical aspects as branches and lea leaves, thus the latter has been treated with less importance.

  • PDF

Textbook vocabulary analysis for Korean phonics program of 1st and 2nd graders (한글 파닉스 교육을 위한 초등 1-2학년 교과서 어휘 자소분석)

  • Lee, Daeun;Kim, Hyeji;Shin, Gayoung;Seol, Ahyoung;Pae, Soyeong;Kim, Mibae
    • 한국어정보학회:학술대회논문집
    • /
    • 2016.10a
    • /
    • pp.226-230
    • /
    • 2016
  • 본 연구는 초등 저학년 읽기부진아동을 위한 한글 파닉스 교육의 기반을 확립하고자 1-2학년 교과서 고빈도 어절 531개를 기반으로 자소 및 음운규칙을 분석하였다. 연구결과, 자소-음소 일치 어절을 기반으로 하였을 때 초성에서 50번 이상 나타난 자소는 /ㄱ/, /ㄹ/, /ㄴ/, /ㅅ/, /ㅎ/, /ㅈ/이다. 중성에서 50번 이상 나타난 자소는 /ㅏ/, /ㅣ/, /ㅗ/, /ㅡ/, /ㅜ/이다. 종성에서 50번 이상 나타난 자소는 /ㄹ/, /ㄴ/, /ㅇ/이다. 자소와 음소가 불일치 된 어절을 기반으로 하였을 때 가장 많이 출현하는 음운규칙은 연음화 규칙이었다. 본 연구결과를 바탕으로 교과서를 기반으로 한 한글 파닉스 교육에 유용하게 사용될 수 있을 것이다.

  • PDF