• Title/Summary/Keyword: Textmining Analysis

Search Result 46, Processing Time 0.023 seconds

Analysis of trend in construction using textmining method (텍스트마이닝을 활용한 건설분야 트랜드 분석)

  • Jeong, Cheol-Woo;Kim, Jae-Jun
    • Journal of The Korean Digital Architecture Interior Association
    • /
    • v.12 no.2
    • /
    • pp.53-60
    • /
    • 2012
  • In this paper, we present new methods for identifying keywords for foresight topics that utilize the internet and textmining techniques to draw objective and quantified information that support experts' qualitative opinions and evaluations in foresight. Furthermore, by applying this fabricated procedure, we have derived keywords to analyze priorities in architectural engineering. Not much difference between qualitative methods of experts and quantitative methods such as text mining has been observed from comparison between technologies derived via qualitative method from "The Science Technology Vision" (control group). Therefore, as a quantitative tool useful for drawing keywords for foresight, textmining can supplement quantitative analysis by experts. In addition, depending on the level and type of raw data, text mining can bring better results in deriving foresight keywords. For this reason, research activities accommodating Internet search results and the development of textmining methods for analyzing current trends are in demand.

A Study on the Analysis of Agricultural R&D Keywords Using Textmining Method (텍스트마이닝을 활용한 농업 R&D 키워드 분석)

  • Kim, Ji-Hoon;Kim, Seong-Sup
    • Journal of the Korea Academia-Industrial cooperation Society
    • /
    • v.22 no.2
    • /
    • pp.721-732
    • /
    • 2021
  • This study analyzed keywords for agricultural R&D using the textmining method to examine the trend of agricultural R&D. Data used for the analysis included R&D project information provided by NTIS, and the research and development step by year from 2003 to 2018 were classified and applied. The TF-IDF approach was used as the analysis method, and ranking was derived based on score. Furthermore, we analyzed by grouping for similar keywords. The main analysis results are as follows. First, agricultural R&D trends are changing according to the introduction of new technologies and changes in the external environment. Second, keyword changes appeared with a time lag in the R&D step. The main keywords are changing in the order of basic research - applied research - development research. Third, the main keyword of agricultural R&D was 'rice.' However, the direction and purpose of the research were changing according to changes in the domestic and foreign agricultural environments.

Big Data Analytics Applied to the Construction Site Accident Factor Analysis

  • KIM, Joon-soo;Lee, Ji-su;KIM, Byung-soo
    • International conference on construction engineering and project management
    • /
    • 2015.10a
    • /
    • pp.678-679
    • /
    • 2015
  • Recently, safety accidents in construction sites are increasing. Accordingly, in this study, development of 'Big-Data Analysis Modeling' can collect articles from last 10 years which came from the Internet News and draw the cause of accidents that happening per season. In order to apply this study, Web Crawling Modeling that can collect 98% of desired information from the internet by using 'Xml', 'tm', "Rcurl' from the library of R, a statistical analysis program has been developed, and Datamining Model, which can draw useful information by using 'Principal Component Analysis' on the result of Work Frequency of 'Textmining.' Through Web Crawling Modeling, 7,384 out of 7,534 Internet News articles that have been posted from the past 10 years regarding "safety Accidents in construction sites", and recognized the characteristics of safety accidents that happening per season. The result showed that accidents caused by abnormal temperature and localized heavy rain, occurred frequently in spring and winter, and accidents caused by violation of safety regulations and breakdown of structures occurred frequently in spring and fall. Plus, the fact that accidents happening from collision of heavy equipment happens constantly every season was acknowledgeable. The result, which has been obtained from "Big-Data Analysis Modeling" corresponds with prior studies. Thus, the study is reliable and able to be applied to not only construction sites but also in the overall industry.

  • PDF

On the Development of Risk Factor Map for Accident Analysis using Textmining and Self-Organizing Map(SOM) Algorithms (재해분석을 위한 텍스트마이닝과 SOM 기반 위험요인지도 개발)

  • Kang, Sungsik;Suh, Yongyoon
    • Journal of the Korean Society of Safety
    • /
    • v.33 no.6
    • /
    • pp.77-84
    • /
    • 2018
  • Report documents of industrial and occupational accidents have continuously been accumulated in private and public institutes. Amongst others, information on narrative-texts of accidents such as accident processes and risk factors contained in disaster report documents is gaining the useful value for accident analysis. Despite this increasingly potential value of analysis of text information, scientific and algorithmic text analytics for safety management has not been carried out yet. Thus, this study aims to develop data processing and visualization techniques that provide a systematic and structural view of text information contained in a disaster report document so that safety managers can effectively analyze accident risk factors. To this end, the risk factor map using text mining and self-organizing map is developed. Text mining is firstly used to extract risk keywords from disaster report documents and then, the Self-Organizing Map (SOM) algorithm is conducted to visualize the risk factor map based on the similarity of disaster report documents. As a result, it is expected that fruitful text information buried in a myriad of disaster report documents is analyzed, providing risk factors to safety managers.

Building Modeling for Unstructured Data Analysis Using Big Data Processing Technology (빅데이터 처리 기술을 활용한 비정형데이터 분석 모델링 구축)

  • Kim, Jung-Hoon;Kim, Sung-Jin;Kwon, Gi-Yeol;Ju, Da-Hye;Oh, Jae-Yong;Lee, Jun-Dong
    • Proceedings of the Korean Society of Computer Information Conference
    • /
    • 2020.07a
    • /
    • pp.253-255
    • /
    • 2020
  • 기업 및 기관 데이터는 워드프로세서, 프레젠테이션, 이메일, open api, 엑셀, XML, JSON 등과 같은 텍스트 기반의 비정형 데이터로 구성되어 있습니다. 텍스트 마이닝(Textmining)을 통해서 자연어 처리 및 기계학습 등의 기술을 이용하여 정보의 추출부터 요약·분류·군집·연관도 분석 등의 과정을 수행울 진행한다. 다양한 시각화 데이터를 보여줄 수 있는 다양한 모델 구축을 진행한 후 민원 신청 내용을 분석 및 변환 작업을 진행한다. 본 논문은 AI 기술과 빅데이터를 활용하여 민원을 분석을 하여 알맞은 부서에 민원을 자동으로 할당해 주는 기술을 다룬다.

  • PDF

Monitoring Trends of Safety Technology Development of Industry Fields Using Patent Analysis (특허분석을 활용한 산업별 안전기술개발 동향 모니터링)

  • Choi, Yuri;Suh, Yongyoon
    • Journal of the Korean Society of Safety
    • /
    • v.35 no.4
    • /
    • pp.92-100
    • /
    • 2020
  • Along with the rapid development of industrial technology, the industrial structure has been continuously changed. Accordingly, safety technologies have been gradually developed to be applied into various industrial fields as well, not limited to a specific industry area. As a result, it became important to analyze and predict trends of safety technology development in order to establish technology strategies for industrial safety. In particular, since patents are easily accessible to gather the technology and business information, many studies have highlighted technology forecasting using patent information. Thus, this study proposes the patent analysis of monitoring trends of safety technologies of industry fields, taking into account both static and dynamic aspects through index and text analysis. First, patent documents containing safety-related keywords are collected from the WIPSON database for extracting technology information. Then, the development trends of safety technologies by industry fields are identified and analyzed through the analysis of indicators such as marketability, growth, and activation. The results of various indicator analyses of safety technologies are visualized to compare among industrial safety technologies for businesses and technology developers. Second, textmining algorithm is applied to identify trends of specific technology keywords of major industries extracted from patent index analysis. As a result, it is expected that the safety manager uses the patent analysis of safety technologies to provide safety technology information with safety-related companies and institutes. The extracted safety technologies are applicable to business practice and predict future promising technologies.

Entitymetrics Analysis of the Research Works of Dong-ju Yun using Textmining (텍스트마이닝을 이용한 윤동주 연구의 개체계량학적 분석)

  • Park, Jinkyeun;Kim, Taekyoun;Song, Min
    • Journal of the Korean BIBLIA Society for library and Information Science
    • /
    • v.28 no.1
    • /
    • pp.191-207
    • /
    • 2017
  • This paper employs entitymetrics analysis on the research works of Dong-ju Yun. He was a Korean poet who was studied by many researchers on his works, religion and life. We collected 1,076 papers about Dong-ju Yun and conducted various approaches including co-author citation analysis, topic modeling analysis to identify the topic trend in the study of Dong-ju Yun. Also we extracted entities like person's name and literature's title from abstract to examine the relationship among them. The result of this paper enables us to objectively identify the topic trend and infer implicit relationships between key concept associated with Dong-ju Yun based on text data. Moreover, we observed sub-research topics such as life, poem, aesthetic existence, comparative literature, literary translation, and religious beliefs. This paper shows how entitymetrics can be utilized to study intellectual structures in the humanities.

Structuring of unstructured big data and visual interpretation (부산지역 교통관련 기사를 이용한 비정형 빅데이터의 정형화와 시각적 해석)

  • Lee, Kyeongjun;Noh, Yunhwan;Yoon, Sanggyeong;Cho, Youngseuk
    • Journal of the Korean Data and Information Science Society
    • /
    • v.25 no.6
    • /
    • pp.1431-1438
    • /
    • 2014
  • We analyzed the articles from "Kukje Shinmun" and "Busan Ilbo", which are two local newpapers of Busan Metropolitan City. The articles cover from January 1, 2013 to December 31, 2013. Meaningful pattern inherent in 2889 articles of which the title includes "Busan" and "Traffic" and related data was analyzed. Textmining method, which is a part of datamining, was used for the social network analysis (SNA). HDFS and MapReduce (from Hadoop ecosystem), which is open-source framework based on JAVA, were used with Linux environment (Uubntu-12.04LTS) for the construction of unstructured data and the storage, process and the analysis of big data. We implemented new algorithm that shows better visualization compared with the default one from R package, by providing the color and thickness based on the weight from each node and line connecting the nodes.

CiNet: GUI based Literature analysis tool using citation information

  • Lee, Se-Jun;Lee, Kwang-H.
    • Bioinformatics and Biosystems
    • /
    • v.2 no.1
    • /
    • pp.33-36
    • /
    • 2007
  • Scientific literature is the most reliable and comprehensive source of knowledge for scientific and biomedical information. Citation information in the literature is also reliable source for linking between literatures. We proposed CiNet, a graphic user interface based tool that extracts the trend of the research using citation information. We can navigate related literatures and extract keywords from the linked literature using this tool. These extracted keywords will be helpful to researchers who want to survey the information.

  • PDF

Text Analytics for Classifying Types of Accident Occurrence Using Accident Report Documents (사고보고문서를 이용한 텍스트 기반 사고발생 유형 및 관계 분석)

  • Kim, Beom Soo;Chang, Seongrok;Suh, Yongyoon
    • Journal of the Korean Society of Safety
    • /
    • v.33 no.3
    • /
    • pp.58-64
    • /
    • 2018
  • Recently, a lot of accident report documents have accumulated in almost all of industries, including critical information of accidents. Accordingly, text data contained in accident report documents are considered useful information for understanding accident processes. However, there has been a lack of systematic approaches to analyzing accident report documents. In this respect, this paper aims at proposing text analytics approach to extracting critical information on accident processes. To be specific, major causes of the accident occurrence are classified based on text information contained in accident report documents by using both textmining and latent Dirichlet allocation (LDA) algorithms. The textmining algorithm is used to structure the document-term matrix and the LDA algorithm is applied to extract latent topics included in a lot of accident report documents. We extract ten topics of accidents as accident types and related keywords of accidents with respect to each accident type. The cause-and-effect diagram is then depicted as a tool for navigating processes of the accident occurrence by structuring causes extracted from LDA. Further, the trends of accidents are identified to explore patterns of accident occurrence in each of types. Three patterns of increasing to decreasing, decreasing to increasing, or only increasing are presented in the case of a chemical plant. The proposed approach helps safety managers systematically supervise the causes and processes of accidents through analysis of text information contained in accident report documents.