• Title/Summary/Keyword: numerals

Search Result 96, Processing Time 0.042 seconds

Combining Multiple Classifiers using Product Approximation based on Third-order Dependency (3차 의존관계에 기반한 곱 근사를 이용한 다수 인식기의 결합)

  • 강희중
    • Journal of KIISE:Software and Applications
    • /
    • v.31 no.5
    • /
    • pp.577-585
    • /
    • 2004
  • Storing and estimating the high order probability distribution of classifiers and class labels is exponentially complex and unmanageable without an assumption or an approximation, so we rely on an approximation scheme using the dependency. In this paper, as an extended study of the second-order dependency-based approximation, the probability distribution is optimally approximated by the third-order dependency. The proposed third-order dependency-based approximation is applied to the combination of multiple classifiers recognizing handwritten numerals from Concordia University and the University of California, Irvine and its usefulness is demonstrated through the experiments.

Automatic Transcription of Three Ambiguous Symbols Used with Arabic Numerals: Period, Colon and Slash. (아라비안 숫자를 동반한 중의적 기호의 자동전사: 온점, 쌍점, 빗금을 중심으로)

  • 윤애선;정영임;권혁철
    • Language and Information
    • /
    • v.8 no.1
    • /
    • pp.117-136
    • /
    • 2004
  • In this paper, we have proposed Auto- TSS, an automatic transcription module of three ambiguous symbols-period (.), colon (:) and slash (/)--using their linguistic contexts. Few previous studies have discussed the problems of ambiguities in reading those symbols into Korean alphabetic letters in order to improve the current Korean TTS (Text-To-Speech) systems. We have classified 9 different reading formulae of the three symbols, analyzed their left and right contexts, and investigated selection rules and distributions between the symbols and their contexts. Based on these linguistic features, 30 stereotyped patterns, 53 rules and 5 heuristics determining the types of reading formulae are investigated for Auto-TSS. This module works modularly in 4 steps. The pilot test was conducted with three test suites, which contain respectively 6,979, 3,491 and 2,450 morpheme clusters containing at least one of three ambiguous symbols and Arabic numeral(s). Encouraging results of 94.3%, 93.0%, 94.2% accuracy were obtained for the test suites. Our next phases are to develop a guessing routine for unknown contexts of the union symbols by using statistical information; to refine the proper nouns and terminology detecting module; and to apply Auto-TSS on a larger scale.

  • PDF

Machine-printed Numeral Recognition using Weighted Template Matching (가중 원형 정합을 이용한 인쇄체 숫자 인식)

  • Jung, Min-Chul
    • Journal of the Korea Academia-Industrial cooperation Society
    • /
    • v.10 no.3
    • /
    • pp.554-559
    • /
    • 2009
  • This paper proposes a new method of weighted template matching fur machine-printed numeral recognition. The proposed weighted template matching, which emphasizes the feature of a pattern using adaptive Hamming distance on local feature areas, improves the recognition rate while template matching processes an input image as one global feature. The experiment compares confusion matrices of the template matching, error back propagation neural network classifier, and the proposed weighted template matching respectively. The result shows that the proposed method improves fairly the recognition rate of the machine-printed numerals.

A Recognition Algorithm of Handwritten Numerals based on Structure Features (구조적 특징기반 자유필기체 숫자인식 알고리즘)

  • Song, Jeong-Young
    • The Journal of the Institute of Internet, Broadcasting and Communication
    • /
    • v.18 no.6
    • /
    • pp.151-156
    • /
    • 2018
  • Because of its large differences in writing style, context-independency and high recognition accuracy requirement, free handwritten digital identification is still a very difficult problem. Analyzing the characteristic of handwritten digits, this paper proposes a new handwritten digital identification method based on combining structural features. Given a handwritten digit, a variety of structural features of the digit including end points, bifurcation points, horizontal lines and so on are identified automatically and robustly by a proposed extended structural features identification algorithm and a decision tree based on those structural features are constructed to support automatic recognition of the handwritten digit. Experimental result demonstrates that the proposed method is superior to other general methods in recognition rate and robustness.

Designing a Classification System for Minhwa DB (민화 DB를 위한 분류체계 설계)

  • Choi, Eunjin;Lee, Young-Suk
    • Journal of Korea Multimedia Society
    • /
    • v.25 no.1
    • /
    • pp.135-143
    • /
    • 2022
  • In order to convert Korean folk paintings called Minhwa, a part of traditional Korean heritage, into DBs, it is necessary to design a classification system suitable for the characteristics of folk paintings. A classification system and the generating of unique codes are required to classify and save them. To realize this, a basic classification system was created by listing objects depicted in folk paintings, and keywords were extracted by reclassifying them for each object. In order to assign a unique code to each piece, we organize the English names of each Minhwa since the English names of the folk painting contain the names of objects. The code name is extracted by applying the order of nouns and consonant priority rules in English names and attaching five Arabic numerals. These codes are later assigned to each image file stored in the database and are input together with the keyword. The Minhwa DB constructed in this way enables storage and search centered on objects and keywords and the intuitive inferring of the type of object from the code name.

Processing of dosage units and design of database schema for formulas in Korean medicine ontology (한의 온톨로지 처방의 용량 단위 가공과 데이터베이스 스키마 설계)

  • Sang-Kyun, Kim;Yong-Taek, Oh;MyungKu, Lee
    • Herbal Formula Science
    • /
    • v.30 no.4
    • /
    • pp.233-240
    • /
    • 2022
  • Objectives : This study aims to propose a processing method for dosage units of medicinal materials and the database schema to manage formula data in Korean medicine ontology. Methods : All dosage units of medicinal materials are collected from the seven textbooks that contain formula data of Korea medicine ontology. Dosages are converted to Arabic numerals and units that are frequently used are converted to representative units. Database schema is designed for processing and managing the formulas and medicinal materials with dosage units. Results : Seven representative units are selected out of 77 units. They will be used in the addition or subtraction of medicinal materials in a formula support system. The remaining units will be made available for references. Conclusions : EMR or chart programs used in clinical hospitals contain formula data that is already standardized. However, the formula data in Korean medicine literature and textbook is not refined, so it is necessary to process the dosages and units of medicinal materials to use in the formula support system. This result is a processing method to utilize the formula data of Korean medicine textbooks and it will be implemented this method in the established formula support system in the future.

Construction of Multiple Classifier Systems based on a Classifiers Pool (인식기 풀 기반의 다수 인식기 시스템 구축방법)

  • Kang, Hee-Joong
    • Journal of KIISE:Software and Applications
    • /
    • v.29 no.8
    • /
    • pp.595-603
    • /
    • 2002
  • Only a few studies have been conducted on how to select multiple classifiers from the pool of available classifiers for showing the good classification performance. Thus, the selection problem if classifiers on how to select or how many to select still remains an important research issue. In this paper, provided that the number of selected classifiers is constrained in advance, a variety of selection criteria are proposed and applied to tile construction of multiple classifier systems, and then these selection criteria will be evaluated by the performance of the constructed multiple classifier systems. All the possible sets of classifiers are trammed by the selection criteria, and some of these sets are selected as the candidates of multiple classifier systems. The multiple classifier system candidates were evaluated by the experiments recognizing unconstrained handwritten numerals obtained both from Concordia university and UCI machine learning repository. Among the selection criteria, particularly the multiple classifier system candidates by the information-theoretic selection criteria based on conditional entropy showed more promising results than those by the other selection criteria.

Development of a Video Caption Recognition System for Sport Event Broadcasting (스포츠 중계를 위한 자막 인식 시스템 개발)

  • Oh, Ju-Hyun
    • 한국HCI학회:학술대회논문집
    • /
    • 2009.02a
    • /
    • pp.94-98
    • /
    • 2009
  • A video caption recognition system has been developed for broadcasting sport events such as major league baseball. The purpose of the system is to translate the information expressed in English units such as miles per hour (MPH) to the international system of units (SI) such as km/h. The system detects the ball speed displayed in the video and recognizes the numerals. The ball speed is then converted to km/h and displayed by the following character generator (CG) system. Although neural-network based methods are widely used for character and numeral recognition, we use template matching to avoid the training process required before the broadcasting. With the proposed template matching method, the operator can cope with the situation when the caption’s appearance changed without any notification. Templates are configured by the operator with a captured screenshot of the first pitch with ball speed. Templates are updated with following correct recognition results. The accuracy of the recognition module is over 97%, which is still not enough for live broadcasting. When the recognition confidence is low, the system asks the operator for the correct recognition result. The operator chooses the right one using hot keys.

  • PDF

Chief causes for the development of the dewey decimal classification (듀이 십진분류법의 발전요인)

  • 이창수
    • Journal of Korean Library and Information Science Society
    • /
    • v.13
    • /
    • pp.85-111
    • /
    • 1986
  • Dewey Decimal Classification (DDC) was first published in 1876. Since its first edition it has been revised, on an average, 6 years, and now it has become the widely used library classification system of which the scheme was translated in various languages. The purpose of this study is to find out the chief causes for the development of the DDC. The results of the study can be summarized as follows: 1. It allows materials to be shelved in a relative location as the collection expands. before the DDC was introduced, libraries used a fixed location for materials in which each item was assigned to a certain location set aside for a subject. 2. It is a practical system. The fact that it has survived many storms in the past hundred years and is still the most widely used classification scheme in the world today attests to its practical value. 3. The pure notation of arabic numerals is universally recognizable. People from any cultural or language background can adapt to the system easily. 4. The use of the decimal system enable infinite expansion and sub-division. And it has adaptability for use in libraries of various size and kinds because of its hierarchically expressive notation which permits varying degrees of inclusiveness and exclusiveness within its decimal structure. 5. The notation is simple and easily understood. The self-evident numerical sequence facilitates filing and shelving. And the mnemonic nature of the notation helps the readers to memorize and recognize the class numbers. 6. The relative index brings together different aspects of the same subject scattered in different disciplines. 7. We can avail of DDC numbers for specific titles easily because of its use by many central bibliographic services. 8. It is being continuously revised by a permanent office established in the library of congress in 1933. This office has been responsible for editing all editions of the DDC since the 16th (1958). And the periodic revision at regular intervals ensures the currentness of the scheme. 9. It has adaptability both for conventional (manual) shelf or classed catalogue analysis and also, through its meaningful nation, for retrieval through mechanization and computerized systems.

  • PDF

A High Order Product Approximation Method based on the Minimization of Upper Bound of a Bayes Error Rate and Its Application to the Combination of Numeral Recognizers (베이스 에러율의 상위 경계 최소화에 기반한 고차 곱 근사 방법과 숫자 인식기 결합에의 적용)

  • Kang, Hee-Joong
    • Journal of KIISE:Software and Applications
    • /
    • v.28 no.9
    • /
    • pp.681-687
    • /
    • 2001
  • In order to raise a class discrimination power by combining multiple classifiers under the Bayesian decision theory, the upper bound of a Bayes error rate bounded by the conditional entropy of a class variable and decision variables obtained from training data samples should be minimized. Wang and Wong proposed a tree dependence first-order approximation scheme of a high order probability distribution composed of the class and multiple feature pattern variables for minimizing the upper bound of the Bayes error rate. This paper presents an extended high order product approximation scheme dealing with higher order dependency more than the first-order tree dependence, based on the minimization of the upper bound of the Bayes error rate. Multiple recognizers for unconstrained handwritten numerals from CENPARMI were combined by the proposed approximation scheme using the Bayesian formalism, and the high recognition rates were obtained by them.

  • PDF