• 제목/요약/키워드: Data Mining System

검색결과 1,299건 처리시간 0.027초

이기종 분산환경에서 데이터마이닝을 위한 데이터준비 시스템 구현 (Implementation of Data Preparation System for Data Mining on Heterogenious Distributed Environment)

  • 이상희;이원섭
    • 한국컴퓨터정보학회논문지
    • /
    • 제9권3호
    • /
    • pp.109-113
    • /
    • 2004
  • 본 논문에서는 데이터 마이닝을 위한 데이터 준비 과정에 대하여 기존의 데이터 마이닝 도구들의 효율성을 비교하고, 새로운 효율적인 데이터 준비 시스템 설계 기준을 제안하고자 한다. 지역 및 원격 데이터베이스 접근방법 이기종 컴퓨터간의 정보 교환을 기준으로 기존의 데이터마이닝 도구들의 기능을 비교하였다. 본 논문에서는 앤서트리, 클레멘타인, 엔터프라이즈 마이너, 웨카를 비교하였다. 또한, 본 논문에서는 분산 네트워크 상에서 데이터 마이닝을 위한 효율적인 데이터 준비 시스템을 위한 설계기준을 제안한다.

  • PDF

목표 속성을 고려한 연관규칙과 분류 기법 (Directed Association Rules Mining and Classification)

  • 한경록;김재련
    • 산업경영시스템학회지
    • /
    • 제24권63호
    • /
    • pp.23-31
    • /
    • 2001
  • Data mining can be either directed or undirected. One way of thinking about it is that we use undirected data mining to recognize relationship in the data and directed data mining to explain those relationships once they have been found. Several data mining techniques have received considerable research attention. In this paper, we propose an algorithm for discovering association rules as directed data mining and applying them to classification. In the first phase, we find frequent closed itemsets and association rules. After this phase, we construct the decision trees using discovered association rules. The algorithm can be applicable to customer relationship management.

  • PDF

공간 데이터 마이닝 시스템의 설계 및 구현 (Design and Implementation of a Spatial Data Mining System)

  • 배덕호;백지행;오현교;송주원;김상욱;최명회;조현주
    • 한국공간정보시스템학회 논문지
    • /
    • 제11권2호
    • /
    • pp.119-132
    • /
    • 2009
  • GIS 기술의 발달로 많은 양의 공간 데이터가 축적됨에 따라 공간 데이터 마이닝의 중요성이 커지고 있다. 본 논문에서는 새로운 공간 데이터 마이닝 시스템 SD-Miner를 제안한다. SD-Miner는 크게 입력과 출력을 담당하는 사용자 인터페이스, 공간 데이터 마이닝 기능을 처리하는 데이터 마이닝 모듈, DBMS를 이용하여 데이터를 저장하고 관리하는 데이터 저장 모듈의 세 부분으로 구성된다. 특히, 데이터 마이닝 함수 모듈에서는 공간 데이터 마이닝의 주요 기법인 공간 클러스터링, 공간 분류, 공간 특성화, 시공 간 연관규칙 탐사 기능을 제공한다. SD-Miner는 다음과 같은 특징을 가진다. SD-Miner는 사용자로 하여 금 공간 데이터 마이닝뿐만 아니라 비 공간 데이터에 대한 마이닝도 가능하게 하며, 각 마이닝 함수들을 라이브러리 형태로 제공하기 때문에 다른 시스템에서도 쉽게 사용 가능하다. 또한, 마이닝 매개 변수들을 테이블의 형태로 입력받기 때문에 시스템의 범용성이 높다. 개발된 SD-Miner의 실용성을 규명하기 위하여 실제 공간 데이터를 이용한 데이터 마이닝을 수행함으로써 여러 가지 의미있는 결과를 도출한다.

  • PDF

Genetic Algorithm Application to Machine Learning

  • Han, Myung-mook;Lee, Yill-byung
    • 한국지능시스템학회논문지
    • /
    • 제11권7호
    • /
    • pp.633-640
    • /
    • 2001
  • In this paper we examine the machine learning issues raised by the domain of the Intrusion Detection Systems(IDS), which have difficulty successfully classifying intruders. There systems also require a significant amount of computational overhead making it difficult to create robust real-time IDS. Machine learning techniques can reduce the human effort required to build these systems and can improve their performance. Genetic algorithms are used to improve the performance of search problems, while data mining has been used for data analysis. Data Mining is the exploration and analysis of large quantities of data to discover meaningful patterns and rules. Among the tasks for data mining, we concentrate the classification task. Since classification is the basic element of human way of thinking, it is a well-studied problem in a wide variety of application. In this paper, we propose a classifier system based on genetic algorithm, and the proposed system is evaluated by applying it to IDS problem related to classification task in data mining. We report our experiments in using these method on KDD audit data.

  • PDF

데이터마이닝 기법(CHAID)을 이용한 효과적인 데이터베이스 마케팅에 관한 연구 (A Study on the Effective Database Marketing using Data Mining Technique(CHAID))

  • 김신곤
    • 정보기술과데이타베이스저널
    • /
    • 제6권1호
    • /
    • pp.89-101
    • /
    • 1999
  • Increasing number of companies recognize that the understanding of customers and their markets is indispensable for their survival and business success. The companies are rapidly increasing the amount of investments to develop customer databases which is the basis for the database marketing activities. Database marketing is closely related to data mining. Data mining is the non-trivial extraction of implicit, previously unknown and potentially useful knowledge or patterns from large data. Data mining applied to database marketing can make a great contribution to reinforce the company's competitiveness and sustainable competitive advantages. This paper develops the classification model to select the most responsible customers from the customer databases for telemarketing system and evaluates the performance of the developed model using LIFT measure. The model employs the decision tree algorithm, i.e., CHAID which is one of the well-known data mining techniques. This paper also represents the effective database marketing strategy by applying the data mining technique to a credit card company's telemarketing system.

  • PDF

Automated Classification of PubMed Texts for Disambiguated Annotation Using Text and Data Mining

  • Choi, Yun-Jeong;Park, Seung-Soo
    • 한국생물정보학회:학술대회논문집
    • /
    • 한국생물정보시스템생물학회 2005년도 BIOINFO 2005
    • /
    • pp.101-106
    • /
    • 2005
  • Recently, as the size of genetic knowledge grows faster, automated analysis and systemization into high-throughput database has become hot issue. One essential task is to recognize and identify genomic entities and discover their relations. However, ambiguity of name entities is a serious problem because of their multiplicity of meanings and types. So far, many effective techniques have been proposed to analyze documents. Yet, accuracy is high when the data fits the model well. The purpose of this paper is to design and implement a document classification system for identifying entity problems using text/data mining combination, supplemented by rich data mining algorithms to enhance its performance. we propose RTP ost system of different style from any traditional method, which takes fault tolerant system approach and data mining strategy. This feedback cycle can enhance the performance of the text mining in terms of accuracy. We experimented our system for classifying RB-related documents on PubMed abstracts to verify the feasibility.

  • PDF

A Prototyping Framework of the Documentation Retrieval System for Enhancing Software Development Quality

  • Chang, Wen-Kui;Wang, Tzu-Po
    • International Journal of Quality Innovation
    • /
    • 제2권2호
    • /
    • pp.93-100
    • /
    • 2001
  • This paper illustrates a prototyping framework of the documentation-standards retrieval system via the data mining approach for enhancing software development quality. We first present an approach for designing a retrieval algorithm based on data mining, with the three basic technologies of machine learning, statistics and database management, applied to this system to speed up the searching time and increase the fitness. This approach derives from the observation that data mining can discover unsuspected relationships among elements in large databases. This observation suggests that data mining can be used to elicit new knowledge about the design of a subject system and that it can be applied to large legacy systems for efficiency. Finally, software development quality will be improved at the same time when the project managers retrieving for the documentation standards.

  • PDF

데이터마이닝을 이용한 설문조사 및 분석 (Questionnaire Survey and Analysis Using Data Mining)

  • 박만희;채화성;신완선
    • 산업경영시스템학회지
    • /
    • 제25권5호
    • /
    • pp.46-52
    • /
    • 2002
  • Today's database system needs to collect huge amount of questionnaire that results from development of the information technology by the internet, so it has to be administrable. However, there are many difficulties concerned with finding analytic data or useful information in the high capacity-database. Data mining can solve these problems and utilize the database. Questionnaire analysis that uses data mining has drawn relevant patterns that did not look or was tended to overlook before. These patterns can be applied by a new business rule. The purpose of this research is to analyze the questionnaire results and to present the result that can help to make decision easily with data mining. Recognition and analysis about these techniques of data mining show suitable type of questionnaire survey. This research focus on the form of present composition and the model of suitable questionnaire to analyze the type of it. Also, the comparison between the actual questionnaire result and the conventional statistical analysis is examined.

DATA MINING-BASED MULTIDIMENSIONAL EXTRACTION METHOD FOR INDICATORS OF SOCIAL SECURITY SYSTEM FOR PEOPLE WITH DISABILITIES

  • BATYHA, RADWAN M.
    • Journal of applied mathematics & informatics
    • /
    • 제40권1_2호
    • /
    • pp.289-303
    • /
    • 2022
  • This article examines the multidimensional index extraction method of the disability social security system based on data mining. While creating the data warehouse of the social security system for the disabled, we need to know the elements of the social security indicators for the disabled. In this context, a clustering algorithm was used to extract the indicators of the social security system for the disabled by investigating the historical dimension of social security for the disabled. The simulation results show that the index extraction method has high coverage, sensitivity and reliability. In this paper, a multidimensional extraction method is introduced to extract the indicators of the social security system for the disabled based on data mining. The simulation experiments show that the method presented in this paper is more reliable, and the indicators of social security system for the disabled extracted are more effective in practical application.

S-PLUS와 StatServer를 이용한 Data Mining 도구 개발 (Development of Data Mining Tool Using S-PLUS and StatServer)

  • 정인석;이재준
    • 지능정보연구
    • /
    • 제4권2호
    • /
    • pp.129-139
    • /
    • 1998
  • 통계 software에는 data mining에 필요한 다양한 모형과 함수들이 제공되고 있어 이를 이용한 data mining 도구가 소개되고 있다. 본 논문에서는 data mining을 수행하는데 효과적인 환경을 제공하는 S-Plus로 data mining 기법들을 구현하거나 재구성하였으며, StatServer를 이용하여 대용량의 data base를 직접 관리할 수 있게 하고, S-PLUS의 분석기능을 Internet을 통하여 사용할 수 있게 하여 원거리에서 data mining작업을 수행될 수 있도록 구성하였다. 또한 분석자는 찾아낸 모형을 복잡한 프로그래밍 작업 없이 새로운 웹 페이지를 만들 수 있으며, 이를 통해 운영계의 사용자가 최적 모형이 제시하는 결과를 실제 업무에 즉시 이용할 수 있도록 하였다.

  • PDF