• Title/Summary/Keyword: Length of Document

Search Result 77, Processing Time 0.021 seconds

A Study on Text Summarize Automation Using Document Length Normalization (문서 길이 정규화를 이용한 문서 요약 자동화에 관한 연구)

  • 이재훈;김영천;이성주
    • Proceedings of the Korean Institute of Intelligent Systems Conference
    • /
    • 2001.05a
    • /
    • pp.228-230
    • /
    • 2001
  • WWW(World Wide Web)와 온라인 정보 서비스의 급속한 성장으로 인해, 보다 많은 정보가 온라인으로 이용 혹은 접근 가능해 졌다. 이런 정보홍수로 접근 가능한 정보들이 과잉되는 문제가 발생했다. 이러한 과잉 정보 현상으로 인하여 시간적 제약이 뒤따르며 이용 가능한 모든 정보를 근거로 중요한 의사 결정을 내려야 한다. 문서 요약 자동화(Text Summarize Automation)는 이 문제를 처리하는데 필수적이다. 본 논문에서는 정보 검색을 통해 획득한 문서들을 일차적으로 문서 길이 정규화를 이용하여 질의에 적합하고 신뢰도가 더욱 높은 문서 정보를 얻을 수 있음을 보인다.

  • PDF

An Enhanced Feature Selection Method Based on the Impurity of Words Considering Unbalanced Distribution of Documents (문서의 불균등 분포를 고려한 단어 불순도 기반 특징 선택 방법)

  • Kang, Jin-Beom;Yang, Jae-Young;Choi, Joong-Min
    • Journal of KIISE:Software and Applications
    • /
    • v.34 no.9
    • /
    • pp.804-816
    • /
    • 2007
  • Sample training data for machine learning often contain irrelevant information or redundant concept. It is also the case that the original data may include noise. If the information collected for constructing learning model is not reliable, it is difficult to obtain accurate information. So the system attempts to find relations or regulations between features and categories in the teaming phase. The feature selection is to remove irrelevant or redundant information before constructing teaming model. for improving its performance. Existing feature selection methods assume that the distribution of documents is balanced in terms of the number of documents for each class and the length of each document. In practice, however, it is difficult not only to prepare a set of documents with almost equal length, but also to define a number of classes with fixed number of document elements. In this paper, we propose a new feature selection method that considers the impurities among the words and unbalanced distribution of documents in categories. We could obtain feature candidates using the word impurity and eventually select the features through unbalanced distribution of documents. We demonstrate that our method performs better than other existing methods via some experiments.

The Kunjung-mun Sangryangmun of Kyunbok-koong Palace (경복궁(景福宮) 근정문(勤政門) 상량문(上樑門))

  • Seo, byung-pae
    • Korean Journal of Heritage: History & Science
    • /
    • v.34
    • /
    • pp.196-209
    • /
    • 2001
  • Kunjung-mun Gate the only existing multi-level palace gate from the Chosun Dynasty, is the main fate of Kungjung-jun, the central building of Kyungbok-koong Palace. Sangryangmun(a written record of the construction of the ridge beam) of Kunjung-mun Gate was discovered in a hole under its main beam during the renovation project on September 19th, 2000. At the time of discovery, Sangryangmun was found in its original state as a rolled up scroll. On a clould-patterned, red silk cloth, 78 cm in width and 1200 cm in length, each of all 92 lines of the Kunjung-mun Sangryangmun is comprised of either 7 or 11 brush-written, ornamental "Seal" characters. With an exception of its discoloration, the material is considered well preserved. After its discovery, the National Institute of Cultural Properties stored the document in an airtight container for a permanent preservation. In accordance to the Royal Command at the time, the Kunjung-mun Sangryangmun was composed by Kim Byungi, then written by Lee Donsang on January 19th 1867. This document records the meaning and the process of the repair effort of Kunjung-mun Gate includes the wish for peace and longevity of the Chosun Kingdom and its people.

Automatic In-Text Keyword Tagging based on Information Retrieval

  • Kim, Jin-Suk;Jin, Du-Seok;Kim, Kwang-Young;Choe, Ho-Seop
    • Journal of Information Processing Systems
    • /
    • v.5 no.3
    • /
    • pp.159-166
    • /
    • 2009
  • As shown in Wikipedia, tagging or cross-linking through major keywords in a document collection improves not only the readability of documents but also responsive and adaptive navigation among related documents. In recent years, the Semantic Web has increased the importance of social tagging as a key feature of the Web 2.0 and, as its crucial phenotype, Tag Cloud has emerged to the public. In this paper we provide an efficient method of automated in-text keyword tagging based on large-scale controlled term collection or keyword dictionary, where the computational complexity of O(mN) - if a pattern matching algorithm is used - can be reduced to O(mlogN) - if an Information Retrieval technique is adopted - while m is the length of target document and N is the total number of candidate terms to be tagged. The result shows that automatic in-text tagging with keywords filtered by Information Retrieval speeds up to about 6 $\sim$ 40 times compared with the fastest pattern matching algorithm.

A Study on Pobeckchuck in the History from the Sunjo to the Sunjong Dynasty (순조(純祖)-순종실록(純宗實錄)에 나타난 포백척(布帛尺)에 관한 연구(硏究))

  • Lee, Eun-Kyung
    • Journal of the Korean Society of Costume
    • /
    • v.58 no.3
    • /
    • pp.116-122
    • /
    • 2008
  • This study aims at defining the meaning of Pobeckchuck in the historical view-point, which appeared in the History of Joseon Dynasty, regarding the periods from the ruling period of Sunjo to that of Sunjong as the latter part of history. Pobeckchuck used in King Sejong was redressed in accordance with the measurement in the Kyeonggukdadejeon(code), in which time one Pobeckchuck was 46.80cm long. It is known that Juchuck, Hwangjongchuck, Youngjochuck, Joraegichuck etc. which had been used in the ruling period of Sejong Dynasty, were used till the period of Youngjo. Also, the document shows that in the 12th ruling period of Sunjo, Pobeckchuck was used for measurement, and in the 20th ruling period of Sunjo, newly-made ruler was only used for the measurement of fields, but no more details about how long it was. But according to the document complied at that time, one Pobeckchuck was 46.80cm long, which fact reveals that the same measurement was used as in the ruling period of Sunjo. When all the measurement laws which were established in the 3rd year of Junghee, the 6th year of Kwangmu were abolished, Pobeckchuck was solely banned from its use, which fact offers a glimpse of how confusing at that period was. The comparison and examination among many documents in the latter part of Joseon Dynasty show the differences within about 4cm that one Pobeckchuck ranged from 44.80cm to 48.80cm long. But no other document on measurement appeared in the History of Joseon Dynasty, except for the 46.80cm. Thus, the 46.80cm corrected in the ruling period of Sunjo proves that one chuck in Pobeckchuck adopted by the dynasty was used as the measurement of length till the ruling period of Sunjong.

The Extraction of Table Lines and Data in Document Image (문서영상에서 표 구성 직선과 데이터 추출)

  • Jang, Dae-Geun;Kim, Eui-Jeong
    • Journal of the Korea Institute of Information and Communication Engineering
    • /
    • v.10 no.3
    • /
    • pp.556-563
    • /
    • 2006
  • We should extract lines and data which consist of the table in order to classify the table region and analyze its structure in document image. But it is difficult to extract lines and data exactly because the lines are cut and their lengths are changed, or characters or noises are merged to the table lines. These problems result from the error of image input device or image reduction. In this paper, we propose the better method of extracting lines and data for table region classification and structure analysis than the previous ones including commercial softwares. The prposed method extracts horizontal and vertical lines which consist of the table by the use of one dimensional median filter. This filter not only eliminates the noises which attach to the line and the lines which are orthogonal to the filtering direction, but also connects the cut line of which the gap is shorter than the length of the filter tap in the process of extracting lines to the filtering direction. Furthermore, texts attached to the line are separated in the process of extracting vertical lines. This is an example of ABSTRACT format.

Improving Multinomial Naive Bayes Text Classifier (다항시행접근 단순 베이지안 문서분류기의 개선)

  • 김상범;임해창
    • Journal of KIISE:Software and Applications
    • /
    • v.30 no.3_4
    • /
    • pp.259-267
    • /
    • 2003
  • Though naive Bayes text classifiers are widely used because of its simplicity, the techniques for improving performances of these classifiers have been rarely studied. In this paper, we propose and evaluate some general and effective techniques for improving performance of the naive Bayes text classifier. We suggest document model based parameter estimation and document length normalization to alleviate the Problems in the traditional multinomial approach for text classification. In addition, Mutual-Information-weighted naive Bayes text classifier is proposed to increase the effect of highly informative words. Our techniques are evaluated on the Reuters21578 and 20 Newsgroups collections, and significant improvements are obtained over the existing multinomial naive Bayes approach.

Extensible Node Numbering Scheme for Updating XML Documents (XML문서 갱신을 위한 확장 가능한 노드 넘버링 구조)

  • Park Chung-Hee;Koo Heung-Seo;Lee Sang-Joon
    • Journal of Korea Multimedia Society
    • /
    • v.8 no.5
    • /
    • pp.606-617
    • /
    • 2005
  • There have been many research efforts which find efficiently all occurrences of the structural relationships between the nodes in the XML document tree for processing a XML query. Most of them use the region numbering which is based on positions of the nodes. Hut The position-based node numbering schemes require update of many node numbers in the XML document li the XML nodes are inserted frequently in the same position. In this paper, we propose ENN(extensible node numbering) scheme using the bucket number which is represented with the variable-length string for decreasing the number of the node numbers which is updated. We also present an performance analysis by comparing ENN method with TP(extended preorder) method.

  • PDF

Multi-Level Sequence Alignment : An Adaptive Control Method Between Speed and Accuracy for Document Comparison (계산속도 및 정확도의 적응적 제어가 가능한 다단계 문서 비교 시스템)

  • Seo, Jong-Kyu;Tak, Haesung;Cho, Hwan-Gue
    • Journal of KIISE
    • /
    • v.41 no.9
    • /
    • pp.728-743
    • /
    • 2014
  • Finger printing and sequence alignment are well-known approaches for document similarity comparison. A fingerprinting method is simple and fast, but it can not find particular similar regions. A string alignment method is used for identifying regions of similarity by arranging the sequences of a string. It has an advantage of finding particular similar regions, but it also has a disadvantage of taking more computing time. The Multi-Level Alignment (MLA) is a new method designed for taking the advantages of both methods. The MLA divides input documents into uniform length blocks, and then extracts fingerprints from each block and calculates similarity of block pairs by comparing the fingerprints. A similarity table is created in this process. Finally, sequence alignment is used for specifying longest similar regions in the similarity table. The MLA allows users to change block's size to control proportion of the fingerprint algorithm and the sequence alignment. As a document is divided into several blocks, similar regions are also fragmented into two or more blocks. To solve this fragmentation problem, we proposed a united block method. Experimentally, we show that computing document's similarity with the united block is more accurate than the original MLA method, with minor time loss.

Comparison of Six Tests for Assessing Hamstring Muscle Length (슬괵근 유연성 평가에 관한 연구)

  • Kim, Suhn-Yeop
    • The Journal of Korean Academy of Orthopedic Manual Physical Therapy
    • /
    • v.5 no.1
    • /
    • pp.39-51
    • /
    • 1999
  • Background and Purpose. Objective measurements of hamstring muscle length are needed to quantify baseline limitations and to document the effectiveness of therapeutic interventions. Several indirect clinical tests for measuring hamstring muscle length are available, but influence of their test procedure is not well documented. The purpose of this study were 1) to describe hamstring muscle length as reflected by use of six tests(active straight leg raising(ASLR), passive straight leg raising(PSLR), passive straight leg raising with the lower back flat(PSLRB), active knee extension(AKE), passive knee extension(PKE), hip joint angle(HJA). 2) to examine the correlation among the tests. Subjects, Sixty subjects(30 men. 30 women) ranging in age from 18 to 25 years(mean 20.2 years) and with no limitation hamstring flexibility and no neurological and orthopedical problems. Methods. All subjects performed six tests. A inclinometer was used to determine the end point of range of motion. HJA was measured using an inclinometer placed over the sacrum. PSLRB were tested PSLR with the low back flat and the opposite thigh slightly flexed and support on pillows. Results, A mean ASLR value of 85.9 degrees, PSLR value of 99.9 degrees, PSLRB value of 109.8 degrees, AKE value of 77.2 degrees PKE value of 83.1 degrees and HJA value of 73.0 degrees were obtained for all subjects. A dependent t-test showed significant difference between the angles of ASLR and PSLR(p<0.001). There was a significant difference between the angles of PSLR and PSLRB(p<0.001). There was a significant difference between the angles of AKE and PKE(p<0.001). The highest correlation was between PSLR and PSLRB(r=0.915, p<0.001). All SLR tests were significants related(p<0.001), as well as AKE and PKE(p<0.001). The lowest correlation was between PKE and HJA(r=0.171. p>0.05). Conclusion and Discussion. The results indicated that the hip flexion angles for ASLR, PSLR and PSLRB were a difference, and the knee extension angles for AKE and PKE were a difference.

  • PDF