• Title/Summary/Keyword: 텍스트 연구

Search Result 3,492, Processing Time 0.034 seconds

A probabilistic information retrieval model by document ranking using term dependencies (용어간 종속성을 이용한 문서 순위 매기기에 의한 확률적 정보 검색)

  • You, Hyun-Jo;Lee, Jung-Jin
    • The Korean Journal of Applied Statistics
    • /
    • v.32 no.5
    • /
    • pp.763-782
    • /
    • 2019
  • This paper proposes a probabilistic document ranking model incorporating term dependencies. Document ranking is a fundamental information retrieval task. The task is to sort documents in a collection according to the relevance to the user query (Qin et al., Information Retrieval Journal, 13, 346-374, 2010). A probabilistic model is a model for computing the conditional probability of the relevance of each document given query. Most of the widely used models assume the term independence because it is challenging to compute the joint probabilities of multiple terms. Words in natural language texts are obviously highly correlated. In this paper, we assume a multinomial distribution model to calculate the relevance probability of a document by considering the dependency structure of words, and propose an information retrieval model to rank a document by estimating the probability with the maximum entropy method. The results of the ranking simulation experiment in various multinomial situations show better retrieval results than a model that assumes the independence of words. The results of document ranking experiments using real-world datasets LETOR OHSUMED also show better retrieval results.

A study of Big-data analysis for relationship between students (학생들의 관계성 파악을 위한 빅-데이터 분석에 관한 연구)

  • Hwang, Deuk-Young;Kim, Jin-Mook
    • Convergence Security Journal
    • /
    • v.15 no.4
    • /
    • pp.113-119
    • /
    • 2015
  • Recent, cyber violence is increasing in a school and the severity of the problems encountered day by day. In particular, the severity of the cyber force using the smart phone is recognized as a very high and great problems socially. Cyberbullying have long damage degree and a wide range time duration against of existed physical cyber violence. Then student's affects is very seriously. Therefore, we analyzes the relationship and languages in the classroom for students to use to identify signs of cyber violence that may occur between friends in the class. And we support this information to identified parent, classroom teachers and school sheriff for prevent cyberbullying accidents in the school. For this research, we will design and implement a messenger in the cyber classroom. It have many components that are Big-data vocabulary, analyzer, and communication interface. Our proposed messenger can analyze lingual sign and friendship between students using Big-data analysis method such as text mining. It can analysis relationship by per-student, per-classroom.

A Study on Equation Recognition Using Tree Structure (트리 구조를 이용한 수식 인식 연구)

  • Park, Byung-Joon;Kim, Hyun-Sik;Kim, Wan-Tae
    • The Journal of Korea Institute of Information, Electronics, and Communication Technology
    • /
    • v.11 no.4
    • /
    • pp.340-345
    • /
    • 2018
  • The Compared to general sentences, the Equation uses a complex structure and various characters and symbols, so that it is not possible to input all the character sets by simply inputting a keyboard. Therefore, the editor is implemented in a text editor such as Hangul or Word. In order to express the Equation properly, it is necessary to have the learner information which can be meaningful to interpret the syntax. Even if a character is input, it can be represented by another expression depending on the relationship between the size and the position. In other words, the form of the expression is expressed as a tree model considering the relationship between characters and symbols such as the position and size to be expressed. As a field of character recognition application, a technique of recognizing characters or symbols(code) has been widely known, but a method of inputting and interpreting a Equation requires a more complicated analysis process than a general text. In this paper, we have implemented a Equation recognizer that recognizes characters in expressions and quickly analyzes the position and size of expressions.

Costume Analysis through the Text and Characters of 'Nezimaki Tori Chronicle' ([태엽 감는 새 연대기]를 텍스트로한 캐릭터와 의상 관계 분석)

  • Lim, Chan;Yu, Dahye;Peak, Lora
    • The Journal of the Korea Contents Association
    • /
    • v.13 no.11
    • /
    • pp.617-625
    • /
    • 2013
  • Along with and <1Q84> Murakami Haruki sees through human nature that fades away with time in his latest work with 'color and memory'. This is the type of character or story to read with the deployment of instrument. Costume of characters conceived in the expansion of the body, the inner expression of personal notes. To others what you want to look, how they formed on the body, with a range of social, if you are claiming a statement. To form the body by forming self-practice is that there becomes. This paper the inside of the characters and the apparent strong interest in the set, costume characters, especially characters were recognized as representing the device of novel. Work directly in the results, or is described as a symbolic system of the body and clothing that embodies the meaning of the symbol you want to the structure of the exchange.

Design and Implementation of XML Web Agent for Data Exchange and Replication between Heterogeneous DBMSs (이기종 DBMS간 데이터 교환과 복제를 위한 XML 웹 에이전트 설계 및 구현)

  • Yu, Sun-Young;Lee, Chun-Keun;Yim, Jae-Hong
    • Journal of Korea Multimedia Society
    • /
    • v.7 no.7
    • /
    • pp.967-975
    • /
    • 2004
  • HTML is unstructured document because of using restricted tag. HTML is difficult to extract data from HTML document. But XML is able to use user definition tag, that is easy to store information. Also XML is easy to extract data from XML document. This is the reason why XML is a standard for data exchange format on the Internet, so XML is fitted to exchange data between heterogeneous DBMSs(DataBase Management System). In this paper, we designed and implemented of XML web agent for data replication between heterogeneous DBMSs. A XML web agent system controls data of DBMS, and generates a XML document from data of DBMS. Also XML web agent is data exchange or replication between heterogeneous DBMS by the medium of XML.

  • PDF

Investigation of the Possibility of Research on Medical Classics Applying Text Mining - Focusing on the Huangdi's Internal Classic - (텍스트마이닝(Text mining)을 활용한 한의학 원전 연구의 가능성 모색 -『황제내경(黃帝內經)』에 대한 적용례를 중심으로 -)

  • Bae, Hyo-jin;Kim, Chang-eop;Lee, Choong-yeol;Shin, Sang-won;Kim, Jong-hyun
    • Journal of Korean Medical classics
    • /
    • v.31 no.4
    • /
    • pp.27-46
    • /
    • 2018
  • Objectives : In this paper, we investigated the applicability of text mining to Korean Medical Classics and suggest that researchers of Medical Classics utilize this methodology. Methods : We applied text mining to the Huangdi's internal classic, a seminal text of Korean Medicine, and visualized networks which represent connectivity of terms and documents based on vector similarity. Then we compared this outcome to the prior knowledge generated through conventional qualitative analysis and examined whether our methodology could accurately reflect the keyword of documents, clusters of terms, and relationships between documents. Results : In the term network, we confirmed that Qi played a key role in the term network and that the theory development based on relativity between Yin and Yang was reflected. In the document network, Suwen and Lingshu are quite distinct from each other due to their differences in description form and topic. Also, Suwen showed high similarity between adjacent chapters. Conclusions : This study revealed that text mining method could yield a significant discovery which corresponds to prior knowledge about Huangdi's internal classic. Text mining can be used in a variety of research fields covering medical classics, literatures, and medical records. In addition, visualization tools can also be utilized for educational purposes.

Emotion-based Gesture Stylization For Animated SMS (모바일 SMS용 캐릭터 애니메이션을 위한 감정 기반 제스처 스타일화)

  • Byun, Hae-Won;Lee, Jung-Suk
    • Journal of Korea Multimedia Society
    • /
    • v.13 no.5
    • /
    • pp.802-816
    • /
    • 2010
  • To create gesture from a new text input is an important problem in computer games and virtual reality. Recently, there is increasing interest in gesture stylization to imitate the gestures of celebrities, such as announcer. However, no attempt has been made so far to stylize a gestures using emotion such as happiness and sadness. Previous researches have not focused on real-time algorithm. In this paper, we present a system to automatically make gesture animation from SMS text and stylize the gesture from emotion. A key feature of this system is a real-time algorithm to combine gestures with emotion. Because the system's platform is a mobile phone, we distribute much works on the server and client. Therefore, the system guarantees real-time performance of 15 or more frames per second. At first, we extract words to express feelings and its corresponding gesture from Disney video and model the gesture statistically. And then, we introduce the theory of Laban Movement Analysis to combine gesture and emotion. In order to evaluate our system, we analyze user survey responses.

Web Accessibility Evaluation of Professional Sports Clubs in Korea (프로스포츠 웹 사이트의 접근성 평가)

  • Choi, Kyoung-Ho;You, Kang-Soo
    • Journal of the Korea Institute of Information and Communication Engineering
    • /
    • v.16 no.3
    • /
    • pp.399-406
    • /
    • 2012
  • The government is supporting the law not to be uncomfortable in all departments of some sports and cultural activities for the handicapped, making the Welfare Law for People with Disability(Article 25) in Korea. Moreover web sites which are places of business more than 300 employees including other public organizations are making it mandatory to observe web accessibility for the handicapped. This study analyzed in statistical aspects to investigate systematically how professional sports clubs observe the accessibility of web site to some degree. As a result, it turned out that the compliance record on the items of the providing of text alternatives(44.92%) for non-text content and the keyboard accessible(46.79%) was low. However, by and large we are able to recognize that the compliance record of the web site is on an increasing trend with the course of time.

A design of Customized Community Service System based on user-behavior analysis on social network (소셜 네트워크 사용자 행위의 속성 분석을 통한 맞춤형 커뮤니티 서비스 시스템 설계)

  • Shin, Eun-se;Kim, Myung-june;Han, So-ra;Oh, Eun-ji;Lee, Kang-whan
    • Proceedings of the Korean Institute of Information and Commucation Sciences Conference
    • /
    • 2012.10a
    • /
    • pp.190-192
    • /
    • 2012
  • 최근 소셜 네트워크 서비스는 언제 어디서나 정보를 누구라도 손쉽게 전달하고 볼 수 있는 수단으로 각광받고 있다. 소셜 네트워크 서비스의 주요한 특징은 사람과 사람, 사람과 정보, 정보와 정보 간의 관계 네트워크로서, 사용자가 능동적으로 참여한다는 것이다. 하지만 범람하는 수많은 정보들 속에서 사용자가 직접 정보를 검색 및 분류해야 하는 과정은 사람과 정보간의 관계 네트워크 측면에서 소셜의 의미를 충족하지 못한다. 이러한 기존의 정보 활용법은 사용자의 선호도에 따른 맞춤형 정보의 수용과 공유를 제시하지 못하고 있다. 본 연구에서 설계된 사용자 맞춤형 서비스 시스템은 사용자의 상황인식 속성정보와 이에 따른 선호도를 평가하는 알고리즘을 기반으로 하여 보다 효율적인 커뮤니티 공간이 제공될 수 있는 맞춤형 커뮤니티 서비스 시스템을 설계 제안한다. 제안된 시스템에서는 소셜 네트워크 서비스에서 사용자가 텍스트를 읽거나 작성하는 행위를 바탕으로 사용자의 관심사를 제공된 알고리즘으로 분석하여 사용자의 선호도에 따른 정보를 분류하고, 사용자의 인적정보로부터 선별한 유사 사용자들을 통해 신뢰성이 높은 정보를 우선적으로 선출한다. 따라서 사용자의 속성과 선호도를 고려한 상황인식 정보를 제공함으로써 사용자가 직접 정보를 검색 및 분류하는 과정을 단축하고 정보의 신뢰성을 향상할 수 있는 방법을 제시한다. 이러한 상황인식 기반의 맞춤형 커뮤니티 서비스 시스템은 실시간으로 많은 정보가 공유되는 서비스에서 다양하게 적용되어 인터넷 신문, 타겟 마케팅 광고 등의 응용분야에서 다양한 정보제공 서비스 시스템으로 적용될 수 있을 것으로 본다.

  • PDF

E-Learning Content Search Support System Design for Self-Directed Learning (자기주도학습을 위한 이러닝 콘텐츠 검색 지원 시스템 설계)

  • Yong, Sung-Jung;Kim, Yu-Doo;Moon, Il-Young
    • Journal of Practical Engineering Education
    • /
    • v.12 no.1
    • /
    • pp.73-83
    • /
    • 2020
  • Recently, the importance of self-directed learning has emerged in the fields of public education, private education, lifelong education, and vocational training education, in which learners can actively cope with knowledge in an infusion-oriented way. However, there are various theoretical knowledge such as concepts and strategies for self-directed learning, but the situation is insufficient for a system where learners can easily receive content in the academic field they want, depending on the actual self-directed learning operation plan or learning area. Therefore, since it is important to provide various learning content in this paper, we utilize text mining techniques to obtain appropriate information and refine and categorize the meaning. On-line, they want to study a system that provides a variety of content in the academic field that learners are trying to acquire.