• Title/Summary/Keyword: "One text

Search Result 1,397, Processing Time 0.025 seconds

An Efficient Machine Learning-based Text Summarization in the Malayalam Language

  • P Haroon, Rosna;Gafur M, Abdul;Nisha U, Barakkath
    • KSII Transactions on Internet and Information Systems (TIIS)
    • /
    • v.16 no.6
    • /
    • pp.1778-1799
    • /
    • 2022
  • Automatic text summarization is a procedure that packs enormous content into a more limited book that incorporates significant data. Malayalam is one of the toughest languages utilized in certain areas of India, most normally in Kerala and in Lakshadweep. Natural language processing in the Malayalam language is relatively low due to the complexity of the language as well as the scarcity of available resources. In this paper, a way is proposed to deal with the text summarization process in Malayalam documents by training a model based on the Support Vector Machine classification algorithm. Different features of the text are taken into account for training the machine so that the system can output the most important data from the input text. The classifier can classify the most important, important, average, and least significant sentences into separate classes and based on this, the machine will be able to create a summary of the input document. The user can select a compression ratio so that the system will output that much fraction of the summary. The model performance is measured by using different genres of Malayalam documents as well as documents from the same domain. The model is evaluated by considering content evaluation measures precision, recall, F score, and relative utility. Obtained precision and recall value shows that the model is trustable and found to be more relevant compared to the other summarizers.

Study on the Space in Works of Mies Van der Rohe in Terms of Text - Focused on Tugendhat, Hubbe House and Barcelona Pavilion - (Text 측면에서 본 Mies Van der Rohe 작품의 공간성 연구 - Tugendhat, Hubbe 주택과 Barcelona Pavilion을 중심으로 -)

  • Yook, Ok-Soo
    • Journal of the Korean housing association
    • /
    • v.25 no.6
    • /
    • pp.101-109
    • /
    • 2014
  • It was early in the $20^{th}$ century when the space was begun to say through the mutual circumstances of form and contents. Adrian Forty explained that the characteristics of space can be divided into three steps by the period: a space of enclosure, a space as continuum and a space as an extension of the body. And there is common condition that all three spaces are accompanied by the form. In the new thinking of architectural form in terms of text in modern society, architecture becomes to more complex to understanding. Saying that there is nothing outside text (Il n'y a rien en dehors du text.) in the world, Jacques Derrida insisted the world to be texted and not to be special centrality, where can be existed by difference and delay its meaning. Text is the structural meaning (sign), not a metaphorical one (symbol). Without the symbol, the architecture can be recognized as text with signing to the form. For that, there is a question how can be explained the space in terms of text extracting the meaning and the symbol. Absolutely not intended by Mies van der Rohe, but in his works of houses and pavilion, its characteristics and traces of text can be seen. If it is possible to analyse his works in the textual view, space of Mies will be found in the same direction of text. And it will be an important opportunity to re-evaluate the space of Mies works standing in the heart of Modern Architecture.

The Analysis of the Film Cooperation Mode of "One text, two productions" in China and South Korea - Take Extreme Job and Lobster cop as Examples - (중국 한국 "하나의 시나리오로 두개를 찍는" 협력방식분석 - <극한직업>과 <랍스타 캅>을 중심으로 -)

  • Li, Hai-Long
    • Journal of Korea Entertainment Industry Association
    • /
    • v.13 no.8
    • /
    • pp.157-165
    • /
    • 2019
  • With the deepening cooperation between China and South Korea in the field of film, the ways of cooperation between China and South Korea are becoming more and more diversified. "One text, two productions" film cooperation mode refers to the filmmakers of two countries sharing script resources, using the same script to create films separately. South Korean Extreme Job(극한직업) and Chinese Lobster cop(龙虾刑警) are the products of this cooperation model. The scripts of the two films originated from the China and South Korea story joint development plan(중한시나리오공통개발프로젝트) in 2015. After winning the award, the script were filmed by Korean and Chinese filmmakers. Filmmakers in two countries adapted the script to different degrees according to their respective cultural backgrounds and aesthetic characteristics. The film cooperation mode of "One text, two productions" promotes the development of transnational cooperation and opens up a new way of film cooperation.

On supporting full-text retrievals in XML query

  • Hong, Dong-Kweon
    • International Journal of Fuzzy Logic and Intelligent Systems
    • /
    • v.7 no.4
    • /
    • pp.274-278
    • /
    • 2007
  • As XML becomes the standard of digital data exchange format we need to manage a lot of XML data effectively. Unlike tables in relational model XML documents are not structural. That makes it difficult to store XML documents as tables in relational model. To solve these problems there have been significant researches in relational database systems. There are two kinds of approaches: 1) One way is to decompose XML documents so that elements of XML match fields of relational tables. 2) The other one stores a whole XML document as a field of relational table. In this paper we adopted the second approach to store XML documents because sometimes it is not easy for us to decompose XML documents and in some cases their element order in documents are very meaningful. We suggest an efficient table schema to store only inverted index as tables to retrieve required data from XML data fields of relational tables and shows SQL translations that correspond to XML full-text retrievals. The functionalities of XML retrieval are based on the W3C XQuery which includes full-text retrievals. In this paper we show the superiority of our method by comparing the performances in terms of a response time and a space to store inverted index. Experiments show our approach uses less space and shows faster response times.

Investigation on the Effect of Multi-Vector Document Embedding for Interdisciplinary Knowledge Representation

  • Park, Jongin;Kim, Namgyu
    • Knowledge Management Research
    • /
    • v.21 no.1
    • /
    • pp.99-116
    • /
    • 2020
  • Text is the most widely used means of exchanging or expressing knowledge and information in the real world. Recently, researches on structuring unstructured text data for text analysis have been actively performed. One of the most representative document embedding method (i.e. doc2Vec) generates a single vector for each document using the whole corpus included in the document. This causes a limitation that the document vector is affected by not only core words but also other miscellaneous words. Additionally, the traditional document embedding algorithms map each document into only one vector. Therefore, it is not easy to represent a complex document with interdisciplinary subjects into a single vector properly by the traditional approach. In this paper, we introduce a multi-vector document embedding method to overcome these limitations of the traditional document embedding methods. After introducing the previous study on multi-vector document embedding, we visually analyze the effects of the multi-vector document embedding method. Firstly, the new method vectorizes the document using only predefined keywords instead of the entire words. Secondly, the new method decomposes various subjects included in the document and generates multiple vectors for each document. The experiments for about three thousands of academic papers revealed that the single vector-based traditional approach cannot properly map complex documents because of interference among subjects in each vector. With the multi-vector based method, we ascertained that the information and knowledge in complex documents can be represented more accurately by eliminating the interference among subjects.

User Authentication Based on Keystroke Dynamics of Free Text and One-Class Classifiers (자유로운 문자열의 키스트로크 다이나믹스와 일범주 분류기를 활용한 사용자 인증)

  • Seo, Dongmin;Kang, Pilsung
    • Journal of Korean Institute of Industrial Engineers
    • /
    • v.42 no.4
    • /
    • pp.280-289
    • /
    • 2016
  • User authentication is an important issue on computer network systems. Most of the current computer network systems use the ID-password string match as the primary user authentication method. However, in password-based authentication, whoever acquires the password of a valid user can access the system without any restrictions. In this paper, we present a keystroke dynamics-based user authentication to resolve limitations of the password-based authentication. Since most previous studies employed a fixed-length text as an input data, we aims at enhancing the authentication performance by combining four different variable creation methods from a variable-length free text as an input data. As authentication algorithms, four one-class classifiers are employed. We verify the proposed approach through an experiment based on actual keystroke data collected from 100 participants who provided more than 17,000 keystrokes for both Korean and English. The experimental results show that our proposed method significantly improve the authentication performance compared to the existing approaches.

A viewpoint of mathematics through the preface of the mathematics text(算學書) (산학서의 서문(序文)에 나타난 산학(算學)에 대한 인식)

  • Lee, Kyung-Eon
    • Communications of Mathematical Education
    • /
    • v.23 no.3
    • /
    • pp.563-581
    • /
    • 2009
  • In this study we review the representations used for emphasizing the significance and requirement of mathematics in Chinese and Korean mathematics text(算學書). Especially, we study four terms; first 六藝之一(육예지일, one of the six arts), second 伏義(복희, Fuxi) 周公(주공, Zhougong) 孔子(공자, Kongzi) 孔門(공문, Kongmen), third 道(도, dao) (색, ze) 微奧(미오, weiai) 精微(정미, jingwei), forth 經世之實用(경세지실용, usefulness in the real life). Through these representations that can be seen in the many mathematics text, we consider the author's efforts to improve the mathematics.

  • PDF

Logistic Regression Ensemble Method for Extracting Significant Information from Social Texts (소셜 텍스트의 주요 정보 추출을 위한 로지스틱 회귀 앙상블 기법)

  • Kim, So Hyeon;Kim, Han Joon
    • KIPS Transactions on Software and Data Engineering
    • /
    • v.6 no.5
    • /
    • pp.279-284
    • /
    • 2017
  • Currenty, in the era of big data, text mining and opinion mining have been used in many domains, and one of their most important research issues is to extract significant information from social media. Thus in this paper, we propose a logistic regression ensemble method of finding the main body text from blog HTML. First, we extract structural features and text features from blog HTML tags. Then we construct a classification model with logistic regression and ensemble that can decide whether any given tags involve main body text or not. One of our important findings is that the main body text can be found through 'depth' features extracted from HTML tags. In our experiment using diverse topics of blog data collected from the web, our tag classification model achieved 99% in terms of accuracy, and it recalled 80.5% of documents that have tags involving the main body text.

A Study on the Textuality Represented in Modern Fashion Photographs (현대 패션사진에 나타난 텍스트성 연구)

  • Park, Mi-Joo;Yang, Sook-Hi
    • The Research Journal of the Costume Culture
    • /
    • v.18 no.5
    • /
    • pp.977-990
    • /
    • 2010
  • Today, as individuals show their social identities and reflect their being as the members of society with a culture, an art style and communication function are stood out in fashion photographs. Accordingly, the meanings of images into text are expanded in its interpretative width through the acceptor's various terms. This researcher looked into four theories of both positions on the textuality of language and image, and considered the point of discussion on image of each theory through modern fashion photographs. First, the theory which divides language and image as auditory and visual recognitions in the textuality of language and image is limited from the view it focuses on only one side without considering the ambivalent elements of each field. For the textuality in modern fashion photographs, the observer attempts to turn it into text to give meaning to it as the recognition through five senses conforming to the acceptor's condition. Second, the theory dividing language and image into the text of time properties and spacial properties has limitation in the text, for acceptor's experience of the object appears as the structured form in time and space rather than being defined as two things like time and space. Third, the theory classifying the language and image text into conventional taste and natural taste has limitation from the view that image text is hardly an object of consistent classification in ease of recognition by the code accepted in society. Thus, this can't be fundamental approach for the understanding of the text of decoding trend represented in modern fashion photographs. Fourth, accordingly, this researcher focussed on contextual and arbitrary text of fashion photographs through the theory of Nelson Goodman which discusses image text through the differences in textuality. Basic mechanism of perceiving and recognizing and distinguish image is closely related to habit and custom like language. So, each acceptor perceives the image as a text through arbitrary interpretation obtained by individual, empirical, historical, and educational viewpoints. The textuality of modern fashion photographs aims to widen the range of diverse knowledge and understanding, transcending the regulations of simple function of existing fashion photographs. Consequently, this researcher puts forward the opinion of consistent and diverse follow-up studies on instilling meaning into fashion photographs for the understanding de-regulatory and de-constructive through various senses by avoiding only one sense-dependent fixed and regulatory properties of it.

Overlay Text Graphic Region Extraction for Video Quality Enhancement Application (비디오 품질 향상 응용을 위한 오버레이 텍스트 그래픽 영역 검출)

  • Lee, Sanghee;Park, Hansung;Ahn, Jungil;On, Youngsang;Jo, Kanghyun
    • Journal of Broadcast Engineering
    • /
    • v.18 no.4
    • /
    • pp.559-571
    • /
    • 2013
  • This paper has presented a few problems when the 2D video superimposed the overlay text was converted to the 3D stereoscopic video. To resolve the problems, it proposes the scenario which the original video is divided into two parts, one is the video only with overlay text graphic region and the other is the video with holes, and then processed respectively. And this paper focuses on research only to detect and extract the overlay text graphic region, which is a first step among the processes in the proposed scenario. To decide whether the overlay text is included or not within a frame, it is used the corner density map based on the Harris corner detector. Following that, the overlay text region is extracted using the hybrid method of color and motion information of the overlay text region. The experiment shows the results of the overlay text region detection and extraction process in a few genre video sequence.