• Title/Summary/Keyword: 단어 출현빈도

Search Result 132, Processing Time 0.032 seconds

English Bible Text Visualization Using Word Clouds and Dynamic Graphics Technology (단어 구름과 동적 그래픽스 기법을 이용한 영어성경 텍스트 시각화)

  • Jang, Dae-Heung
    • The Korean Journal of Applied Statistics
    • /
    • v.27 no.3
    • /
    • pp.373-386
    • /
    • 2014
  • A word cloud is a visualization of word frequency in a given text. The importance of each word is shown in font size or color. This plot is useful for quickly perceiving the most prominent words and for locating a word alphabetically to determine its relative prominence. With dynamic graphics, we can find the changing pattern of prominent words and their frequencies according to the changing selection of chapters in a given text. We can define the word frequency matrix. In this matrix, rows are chapters in text and columns are ranks corresponding to word frequency about the words in the text. We can draw the word frequency matrix plot with this matrix. Dynamic graphic can indicate the changing pattern of the word frequency matrix according to the changing selection of the range of ranks of words. We execute an English Bible text visualization using word clouds and dynamic graphics technology.

Appearance Frequency of 'Eco-Friendly' Emotion and Sensibility Words and their Changes (친환경 감성 어휘의 종류별 사용빈도 및 변화 양상)

  • Na, Young-Joo
    • Science of Emotion and Sensibility
    • /
    • v.14 no.2
    • /
    • pp.207-220
    • /
    • 2011
  • The purpose of this study is to investigate sensibility words related with eco-friendly in the two media fashion magazines and internet newspapers and to analysis their appearance frequency and changes by the year through 1999~2010. Most frequently used words are 'nature, eco, cotton, natural fiber, health, fresh, clear, preservation, harmony, com fiber, and Lohas'. The words are divided in 4 groups: 'Nature/Environment, Material/Fiber, Human, and Adjectives/Micell'. A point of appearing time is analyzed: 'ecology, memory-shape material, organic, spa' were used before 2000, 'nature environment, eco-friendly, stretch material, wellbeing, substitute, recycling' were in 2000-2001, 'smart material, eco material, green' in 2002-2003, 'coolbiz, Lohas, natural dye' in 2004-2005, 'herb medicine, sustainable, warmbiz' in 2006-2007, 'greensumer, greenlife, solar energy, forest bath' in 2008-2009. Looking into their changes, in early 2000, the words of eco-friendly emotion and sensibility had appeared frequently relatively, but later on they decreased, and again recently increased showing highest appearing frequency. 'Nature/Environment' words have appeared recently very much, while 'Human' sensibility words have not changed much or decreased a little. 'Adjective/Micell' words has increased little bit recently. 'Material/Fiber' words showed decrease at fashion magazine, while they increased at the pages of internet news.

  • PDF

Implementation of the Text Abstraction System using the Statistical Information of Korean Documents (한국어 문서의 통계적 정보를 이용한 문서 요약 시스템 구현)

  • Kang, Sang-Bae;Cho, Hyuk-Kyu;Kwon, Hyuk-Chul;Park, Jae-Deuk;Park, Dong-In
    • Annual Conference on Human and Language Technology
    • /
    • 1997.10a
    • /
    • pp.28-33
    • /
    • 1997
  • 이 논문에서는 문장 유사도 측정 기법과 말뭉치 정보를 이용한 문서요약 시스템을 구현하였다. 문서 요약은 문서에서 문장 단위로 단어를 추출하여 문장을 단어의 벡터로 표현하고, 문서 내 단어의 출현빈도와 말뭉치 내 단어의 사용빈도를 이용하여 각 문장의 중요도를 계산한다. 그리고 중요도가 높은 상위 몇 위의 문장을 요약문장으로 추출한다. 실험 결과, 문서내 단어빈도의 중요도를 낮추고, 말뭉치내 일반 사용빈도를 단어의 가중치에 추가했을 때 가장 좋은 효율을 보였다. 또 요약하고자 하는 문서와 유사한 말뭉치를 사용 했을 때 높은 효율을 보였다.

  • PDF

The Effect of Word Frequency on Noun Definitions (단어빈도가 명사정의하기에 미치는 효과)

  • Lee, Chan-Jong
    • The Journal of the Acoustical Society of Korea
    • /
    • v.27 no.6
    • /
    • pp.303-308
    • /
    • 2008
  • The purpose of the present study is to investigate that word frequency has significant influence on noun definitions in Korean. The experimental group was 80 students from Elementary school, Middle school, High school and University. They rated familiarity and wrote definitions for nouns. Noun definitions were analyzed with semantic categories such as "use/purpose," "description," "association/relation," "partial explanation," "explanation," "error," "partial explanation-attribute," "partial explanation-specific class," "partial explanation-nonspecific class," "explanation-specific class," "explanation-nonspecific class." As a result, they showed familiarity for high-frequency nouns. "EXPL" categories that use class terms or critical attributes were used more frequently in definitions of high-frequency nouns compared with low-frequency nouns. They increased with age and errors decreased with age. Word frequency had a significant influence on noun definitions.

Analysis of Technology Trends from Words in Patent Titles (특허 발명의 명칭에 쓰인 단어를 이용한 기술동향 분석 연구)

  • Kim, Tae-Jung;Lee, Myung-Sun;Choi, Ho-Nam
    • The Journal of the Korea Contents Association
    • /
    • v.10 no.4
    • /
    • pp.433-437
    • /
    • 2010
  • Patent contains meaningful technical achievement. There are many cases explaining technology trends from the analysis of frequency of term. Term sometimes has different meaning on fields. In this paper, words from patent titles of US, Japan, Korea PCT and EPO are collected by the 5 categories of WIPO. Frequency changes rate of each word were calculated and high ranked words of 5 categories were analyzed to find relationship between patent and technology development as well as technology trends.

Automatic Document Classification Based on Word Frequency Weight (단어 빈도 가중치를 이용한 자동 문서 분류)

  • Noh, Hyun-A;Kim, Min-Soo;Kim, Soo-Hyung;Park, Hyuk-Ro
    • Proceedings of the Korea Information Processing Society Conference
    • /
    • 2002.11a
    • /
    • pp.581-584
    • /
    • 2002
  • 본 논문에서는 범주 내의 키워드 빈도에 의해 문서를 자동으로 분류하는 방법을 제안한다. 문서 자동분류 시스템에서는 문서와 문서를 비교하기 위해서 분류 자질(feature)에 적절한 가중치를 부여할 필요가 있다. 본 논문에서는 수작업으로 분류된 신문기사를 이용하여 자질의 가중치를 학습하는 방법을 사용하였다. 기존의 용어가중치 방법은 각 범주별로 가장 많이 등장한 명사부터 순서대로 추출하여 가중치를 주는 방법을 사용한 것에 비해 본 논문에서는 명사의 출현 횟수뿐만 아니라 출현위치를 함께 고려하여 가중치를 계산하는 방법을 제안한다. 또한 단어 빈도 가중치 방법의 변형된 방식을 사용함으로써 기존의 단어 빈도 가중치 방법과 비교하여 분류 정확도 측면에서 9%이상 성능 향상을 있음을 보인다.

  • PDF

Analysis of ICT Education Trends using Keyword Occurrence Frequency Analysis and CONCOR Technique (키워드 출현 빈도 분석과 CONCOR 기법을 이용한 ICT 교육 동향 분석)

  • Youngseok Lee
    • Journal of Industrial Convergence
    • /
    • v.21 no.1
    • /
    • pp.187-192
    • /
    • 2023
  • In this study, trends in ICT education were investigated by analyzing the frequency of appearance of keywords related to machine learning and using conversion of iteration correction(CONCOR) techniques. A total of 304 papers from 2018 to the present published in registered sites were searched on Google Scalar using "ICT education" as the keyword, and 60 papers pertaining to ICT education were selected based on a systematic literature review. Subsequently, keywords were extracted based on the title and summary of the paper. For word frequency and indicator data, 49 keywords with high appearance frequency were extracted by analyzing frequency, via the term frequency-inverse document frequency technique in natural language processing, and words with simultaneous appearance frequency. The relationship degree was verified by analyzing the connection structure and centrality of the connection degree between words, and a cluster composed of words with similarity was derived via CONCOR analysis. First, "education," "research," "result," "utilization," and "analysis" were analyzed as main keywords. Second, by analyzing an N-GRAM network graph with "education" as the keyword, "curriculum" and "utilization" were shown to exhibit the highest correlation level. Third, by conducting a cluster analysis with "education" as the keyword, five groups were formed: "curriculum," "programming," "student," "improvement," and "information." These results indicate that practical research necessary for ICT education can be conducted by analyzing ICT education trends and identifying trends.

Analysis of Real Estate Market Trend Using Text Mining and Big Data (빅데이터와 텍스트마이닝을 이용한 부동산시장 동향분석)

  • Chun, Hae-Jung
    • Journal of Digital Convergence
    • /
    • v.17 no.4
    • /
    • pp.49-55
    • /
    • 2019
  • This study is on the trend of real estate market using text mining and big data. The data were collected through internet news posted on Naver from August 2016 to August 2017. As a result of TF-IDF analysis, the frequency was high in the order of housing, sale, household, real estate market, and region. Many words related to policies such as loan, government, countermeasures, and regulations were extracted, and the region - related words appeared the most frequently in Seoul. The combination of the words related to the region showed that the frequencies of 'Seoul - Gangnam', 'Seoul - Metropolitan area', 'Gangnam - reconstruction' and 'Seoul - reconstruction' appeared frequently. It can be seen that the people's interest and expectation about the reconstruction of Gangnam area is high.

Analysis of Information Education Related Theses Using R Program (R을 활용한 정보교육관련 논문 분석)

  • Park, SunJu
    • Journal of The Korean Association of Information Education
    • /
    • v.21 no.1
    • /
    • pp.57-66
    • /
    • 2017
  • Lately, academic interests in big data analysis and social network has been prominently raised. Various academic fields are involved in this social network based research trend, which is, social network has been actively used as the research topic in social science field as well as in natural science field. Accordingly, this paper focuses on the text analysis and the following social network analysis with the Master's and Doctor's dissertations. The result indicates that certain words had a high frequency throughout the entire period and some words had fluctuating frequencies in different period. In detail, the words with a high frequency had a higher betweenness centrality and each period seems to have a distinctive research flow. Therefore, it was found that the subjects of the Master's and Doctor's dissertations were changed sensitively to the development of IT technology and changes in information curriculum of elementary, middle and high school. It is predicted that researches related to smart, mobile, smartphone, SNS, application, storytelling, multicultural, and STEAM, which had an increased frequency in period 4, would be continuously conducted. Moreover, the topics of robots, programming, coding, algorithms, creativity, interaction, and privacy will also be studied steadily.

A study on the perception of 3D virtual fashion before and after COVID-19 using textmining

  • Cho, Hyun-Jin
    • Journal of the Korea Society of Computer and Information
    • /
    • v.27 no.12
    • /
    • pp.111-119
    • /
    • 2022
  • The purpose of this paper is to examine the change in perception of 3D virtual fashion before and after COVID-19 using big data analysis. The data collection period is from January 1, 2017, before the outbreak of COVID-19, to October 30, 2022, after the outbreak. Big data was collected for key words related to 3D virtual fashion extracted from social media such as Naver, Daum, Google, and YouTube using Textom. After the collected words were refined, word cloud, word frequency, connection centrality, network visualization, and CONCOR analysis were performed. As a result of extracting and analyzing 32,461 words with 3D virtual fashion as a keyword, the frequency and centrality of fashion, virtual, and technology appeared the highest, and the frequency of appearance of digital, design, clothing, utilization, and manufacturing was also high. Through this, it was found that 3D virtual fashion is being used throughout the industry along with the development of technology. In particular, the key words that stand out the most after COVID-19 are metaverse and 3D education, which are in high demand in the fashion industry.