• 제목/요약/키워드: News Data

검색결과 884건 처리시간 0.027초

Stock News Dataset Quality Assessment by Evaluating the Data Distribution and the Sentiment Prediction

  • Alasmari, Eman;Hamdy, Mohamed;Alyoubi, Khaled H.;Alotaibi, Fahd Saleh
    • International Journal of Computer Science & Network Security
    • /
    • 제22권2호
    • /
    • pp.1-8
    • /
    • 2022
  • This work provides a reliable and classified stocks dataset merged with Saudi stock news. This dataset allows researchers to analyze and better understand the realities, impacts, and relationships between stock news and stock fluctuations. The data were collected from the Saudi stock market via the Corporate News (CN) and Historical Data Stocks (HDS) datasets. As their names suggest, CN contains news, and HDS provides information concerning how stock values change over time. Both datasets cover the period from 2011 to 2019, have 30,098 rows, and have 16 variables-four of which they share and 12 of which differ. Therefore, the combined dataset presented here includes 30,098 published news pieces and information about stock fluctuations across nine years. Stock news polarity has been interpreted in various ways by native Arabic speakers associated with the stock domain. Therefore, this polarity was categorized manually based on Arabic semantics. As the Saudi stock market massively contributes to the international economy, this dataset is essential for stock investors and analyzers. The dataset has been prepared for educational and scientific purposes, motivated by the scarcity of data describing the impact of Saudi stock news on stock activities. It will, therefore, be useful across many sectors, including stock market analytics, data mining, statistics, machine learning, and deep learning. The data evaluation is applied by testing the data distribution of the categories and the sentiment prediction-the data distribution over classes and sentiment prediction accuracy. The results show that the data distribution of the polarity over sectors is considered a balanced distribution. The NB model is developed to evaluate the data quality based on sentiment classification, proving the data reliability by achieving 68% accuracy. So, the data evaluation results ensure dataset reliability, readiness, and high quality for any usage.

Design and Adaptation for Internet News Data Extraction Middleware(INDEM) System

  • Sun, Bok-Keun
    • 한국컴퓨터정보학회논문지
    • /
    • 제21권4호
    • /
    • pp.55-62
    • /
    • 2016
  • In this paper, we propose the INDEM(Internet News Data Extraction Middleware) system for the removal of the unnecessary data in internet news. Although data on the internet can be used in various fields such as source of data of IR(Information Retrieval), Data mining and knowledge information service, it contains a lot of unnecessary information. The removal of the unnecessary data is a problem to be solved prior to the study of the knowledge-based information service that is based on the data of the web page. The INDEM system parses html and explores the XPath, and it is to perform the analysis. The user simply utilize INDEM by implementing an abstract class that provides INDEM, and can obtain the analysis information. INDEM System through this process delivers the analysis information including the main contents of news site to the users. In this paper, the INDEM system was adapted in a stand-alone and web service system and it was evaluated on the basis of 16 news site. As a result, performance of the INDEM system is affected in html source data size and complexity of used html grammar than the main news data size.

국내 언론사 보건의료 뉴스의 Linked Open Data 구축 (Linked Open Data Construction for Korean Healthcare News)

  • 장종선;조완섭;이경희
    • 한국빅데이터학회지
    • /
    • 제1권2호
    • /
    • pp.79-89
    • /
    • 2016
  • 언론사들은 링크드 데이터(Linked Data) 기술을 활용하여 누적된 지적자산으로부터 새로운 가치를 찾는 노력을 하고 있다. 최근 들어 세계적인 언론 매체인 BBC에서는 링크드 데이터 모형을 이용해 자사의 뉴스 기사 가치를 지속해서 향상시키고 있다. 국내 인터넷 신문사들도 누적된 기사를 재활용하고, 이들로부터 새로운 가치를 찾아 뉴스 기사의 가치를 지속해서 향상시킬 필요성이 있다. 본 논문에서는 보건의료 관련 뉴스를 대상으로 링크드 데이터를 구축하는 연구를 소개한다. 기사문에서 보건의료와 관련된 개체명을 인식하여 데이터베이스화하고, 이를 공개된 다른 정보들과 연결하며, 구조화하여 링크드 데이터 서비스를 제공한다. 연구의 결과는 무분별하게 쌓여있는 뉴스데이터를 체계적으로 정리하고, 공개된 다른 정보들과 연결함으로써 기존에 발견하지 못했던 새로운 인사이트를 찾는 기회를 제공하고, 뉴스 데이터가 재활용될 수 있는데 기여할 수 있다. 마지막으로 SPARQL 질의 언어를 이용하여 뉴스 데이터를 대화식으로 탐색할 수 있는데 기여할 수 있다.

  • PDF

XML 기반 멀티미디어 뉴스 관리 시스템 (An XML-based Multimedia News Management System)

  • 김현희;박승수
    • 정보처리학회논문지B
    • /
    • 제11B권7호
    • /
    • pp.785-792
    • /
    • 2004
  • 최근 인터넷과 멀티미디어 관련 기술의 발달로, 멀티미디어 정보에 대한 사용자들의 요구가 다양해지고 있으며, 특히 한 종류의 멀티미디어 뿐만이 아니라, 다른 종류의 멀티미디어 컨텐츠와 의미 관계를 기반으로 관련 멀티미디어 정보를 검색할 필요성이 대두하고 있다. 그러나 일반 데이터와는 달리 멀티미디어 컨텐츠와 의미 관계는 암시적으로 데이터 내에 내포되어 있기 때문에, 관련 멀티미디어 정보를 제공하는 것이 어렵다. 따라서, 대표적인 응용 프로그램인 멀티미디어 뉴스 관리 시스템의 경우, 대부분의 뉴스 서비스들이 텍스트 기사에 대한 관련 기사를 제공하고 있으며, 비디오 혹은 이미지와 같은 멀티미디어 뉴스에 대한 검색은 독립적으로 이루어지고 있다. 본 논문에서는 XML을 기반으로 관련 멀티미디어 뉴스를 통합, 검색 및 전송할 수 있는 멀티미디어 뉴스 관리 시스템을 개발하였다. 미디어 객체, 관계 객체, 뷰 객체로 구성된 데이터 모델을 제안하여 다양한 종류의 멀티미디어 뉴스 컨텐츠를 표현하고 의미 관계를 포착하였으며, 뷰 메커니즘을 개발하여, 생성된 뷰 객체의 가공을 통하여 사용자에게 멀티미디어 뉴스를 효율적으로 제공하도록 하였다

Algorithm Design to Judge Fake News based on Bigdata and Artificial Intelligence

  • Kang, Jangmook;Lee, Sangwon
    • International Journal of Internet, Broadcasting and Communication
    • /
    • 제11권2호
    • /
    • pp.50-58
    • /
    • 2019
  • The clear and specific objective of this study is to design a false news discriminator algorithm for news articles transmitted on a text-based basis and an architecture that builds it into a system (H/W configuration with Hadoop-based in-memory technology, Deep Learning S/W design for bigdata and SNS linkage). Based on learning data on actual news, the government will submit advanced "fake news" test data as a result and complete theoretical research based on it. The need for research proposed by this study is social cost paid by rumors (including malicious comments) and rumors (written false news) due to the flood of fake news, false reports, rumors and stabbings, among other social challenges. In addition, fake news can distort normal communication channels, undermine human mutual trust, and reduce social capital at the same time. The final purpose of the study is to upgrade the study to a topic that is difficult to distinguish between false and exaggerated, fake and hypocrisy, sincere and false, fraud and error, truth and false.

Statistical Properties of News Coverage Data

  • Lim, Eunju;Hahn, Kyu S.;Lim, Johan;Kim, Myungsuk;Park, Jeongyeon;Yoon, Jihee
    • Communications for Statistical Applications and Methods
    • /
    • 제19권6호
    • /
    • pp.771-780
    • /
    • 2012
  • In the current analysis, we examine news coverage data widely used in media studies. News coverage data is usually time series data to capture the volume or the tone of the news media's coverage of a topic. We first describe the distributional properties of autoregressive conditionally heteroscadestic(ARCH) effects and compare two major American newspaper's coverage of U.S.-North Korea relations. Subsequently, we propose a change point detection model and apply it to the detection of major change points in the tone of American newspaper coverage of U.S.-North Korea relations.

Social Media Fake News in India

  • Al-Zaman, Md. Sayeed
    • Asian Journal for Public Opinion Research
    • /
    • 제9권1호
    • /
    • pp.25-47
    • /
    • 2021
  • This study analyzes 419 fake news items published in India, a fake-news-prone country, to identify the major themes, content types, and sources of social media fake news. The results show that fake news shared on social media has six major themes: health, religion, politics, crime, entertainment, and miscellaneous; eight types of content: text, photo, audio, and video, text & photo, text & video, photo & video, and text & photo & video; and two main sources: online sources and the mainstream media. Health-related fake news is more common only during a health crisis, whereas fake news related to religion and politics seems more prevalent, emerging from online media. Text & photo and text & video have three-fourths of the total share of fake news, and most of them are from online media: online media is the main source of fake news on social media as well. On the other hand, mainstream media mostly produces political fake news. This study, presenting some novel findings that may help researchers to understand and policymakers to control fake news on social media, invites more academic investigations of religious and political fake news in India. Two important limitations of this study are related to the data source and data collection period, which may have an impact on the results.

신문의 환경 보도 분석과 신문활용교육의 가능성 (Analysis of Environment Coverage in Newspapers and Possibility of Application in NIE(Newspaper In Education))

  • 오강호;고영구
    • 한국환경교육학회지:환경교육
    • /
    • 제17권1호
    • /
    • pp.67-76
    • /
    • 2004
  • This study is considered how to use newspapers to apply education by the way of analyses of environment coverage in newspapers. Data for the study were gathered by content analyses of KINDS(Korean Integrated News Database System) established by Korean Press Institute. The environment coverage is mainly placed in social and regional magazines of newspapers, and the news story are mainly assigned to straight/feature magazines in type. The news of environment coverage is mostly gathered from data by government-informer, and the news is positive/agreement or negative/disagreement in tenor. The news gathering methods of planning/magazine newspaper serial are scientific and objective, and they are of the firsthand data by news reporter, contributions by experts and interviews. The spaces of the news are specially edited. The environment news is often negative/disagreeable in tenor because the news is mostly of straight ones written by non-experts. Applying newspapers in education is a useful learning method which students could develop thinking power and induce concerning and interest by themselves. From the results of the study, the useful suggestions to apply newspapers to learning are as follows. At first, spaces and types of news must be read in detail. Secondly, it is hopeful that indirect news by not writer himself might be possibly avoided in learning. Thirdly, the themes of news would be picked up in relation with learning contents. Lastly, it suggests that the tenor of news is neutral or, in cases, positive and negative together possibly.

  • PDF

부도예측 모형에서 뉴스 분류를 통한 효과적인 감성분석에 관한 연구 (A Study on Effective Sentiment Analysis through News Classification in Bankruptcy Prediction Model)

  • 김찬송;신민수
    • 한국IT서비스학회지
    • /
    • 제18권1호
    • /
    • pp.187-200
    • /
    • 2019
  • Bankruptcy prediction model is an issue that has consistently interested in various fields. Recently, as technology for dealing with unstructured data has been developed, researches applied to business model prediction through text mining have been activated, and studies using this method are also increasing in bankruptcy prediction. Especially, it is actively trying to improve bankruptcy prediction by analyzing news data dealing with the external environment of the corporation. However, there has been a lack of study on which news is effective in bankruptcy prediction in real-time mass-produced news. The purpose of this study was to evaluate the high impact news on bankruptcy prediction. Therefore, we classify news according to type, collection period, and analyzed the impact on bankruptcy prediction based on sentiment analysis. As a result, artificial neural network was most effective among the algorithms used, and commentary news type was most effective in bankruptcy prediction. Column and straight type news were also significant, but photo type news was not significant. In the news by collection period, news for 4 months before the bankruptcy was most effective in bankruptcy prediction. In this study, we propose a news classification methods for sentiment analysis that is effective for bankruptcy prediction model.

Effects of Fake News and Propaganda on Management of Information on Covid-19 Pandemic in Nigeria

  • Odunlade, Racheal Opeyemi;Ojo, Joshua Onaade;Oche, Nathaniel Agbo
    • International Journal of Knowledge Content Development & Technology
    • /
    • 제11권4호
    • /
    • pp.35-51
    • /
    • 2021
  • This study measured the effects of fake news and propaganda on managing information on COVID-19 among the Nigerian citizenry. This study examined sources of information on COVID-19 available to the people, evaluated reasons behind spreading fake news, examined how fake news has affected the spread of COVID-19 pandemic in Nigeria, established the consequences of fake news on managing COVID-19 pandemic and as well identified ways to contain fake news at a time like this in Nigeria.It is a survey with a sample size of 375 participants selected using simple random technique. Instrument of data gathering was questionnaire widely distributed in the six geo-political zones of Nigeria using Survey monkey. Data was analysed using frequencies, counts and percentages, tables and charts. Findings revealed that people rely more on radio, television, and social media for information on COVID-19. Fake news is spread by people mostly for political reasons and intention to cause panic. In Nigeria, fake news has led to disbelief of the existence of the virus thereby leading to violation of precautionary measures among the citizenry and lack of trust in the government. Concerted effort on the part of the government is required to give public enlightenment on the danger of fake news. Also, directorate of anti-fake news should be established to censor and reprimand sources of fake news. People should always check source of information to confirm its credibility and be weary of sharing unconfirmed information especially on the social media.