Search | Korea Science

A Study on Implementation of Printed Character Recognition System And Performance Evaluation (인쇄체 문자 인식기의 성능 평가에 관한 연구)

Kim, Min-Soo;Kang, Eun-Young;Kim, Eun-Young;Han, Sun-Hwa;Kim, Jin-Hyung
- The Transactions of the Korea Information Processing Society
- /
- v.7 no.11
- /
- pp.3584-3591
- /
- 2000
In this paper we propose measure for performance evaluationof character recognition, We used three commercial character recognizers and one laboratory character recognizer for test. The characteristics of each recognizer is compared by proposed evaluation standrd, and analyzed characteristrics For the input test data, KT test collection are used. KT test collection is composed of 1000 document images about and complete source text. In this paper we propose method for measuring recognition rage in character unit for evaluation of character recogrition, The recogrition rates are compared and analyzed by single feature characteristic or mixed feature characteristic.
PDF

전자문서 통합 보관할 공인전자문서보관소 눈앞

Jang, Hong-Il
- 프린팅코리아
- /
- s.40
- /
- pp.116-119
- /
- 2005
과거에 비해 단순히 종이 문서를 보관, 열람하는 이들이 점차 줄어들고 있다. 그 이면에는 21세기에 접어들면서 저변 확대의 폭을 넓혀가고 있는 전자 출판의 급상승률이 내제돼 있다. 이제는 손에 놓고 보는 기능을 가진 문서가 아니라 보관이 용이하고 가독률을 높일 수 있는 매개체가 도래했음을 시사하는 것이다. 일례로 지하철 등 대중교통을 이용할 때 휴대폰 문자를 송신하는 것 또한 우리 곁에 가장 가까이 있는 전자 문서의 한 유형이다. 개인별 휴대폰 보유 현황에 있어서도 'OECD 가입 국가 중 최상위권을 유지하고 있다'는 한 연구기관의 발표를 차치하더라도 각 가정에서 쓰고 있는 휴대폰 등 개인이 사용하고 있는 전자 기기만 보아도 피부로 직접 느낄 수 있다. 그렇다면 국내 전자 문서에 대한 인식은 어떻게 변해가고 있을까
PDF

Recognition of Various Printed Hangul Images by using the Boundary Tracing Technique (경계선 기울기 방법을 이용한 다양한 인쇄체 한글의 인식)

Baek, Seung-Bok;Kang, Soon-Dae;Sohn, Young-Sun
- Journal of the Korean Institute of Intelligent Systems
- /
- v.13 no.1
- /
- pp.1-5
- /
- 2003
In this paper, we realized a system that converts the character images of the printed Korean alphabet (Hangul) to the editable text documents by using the black and white CCD camera, We were able to abstract the contours information of the character which is based on the structural character by using the boundary tracing technique that is strong to the noise on the character recognition. By using the contours information, we recognized the horizontal vowels and vertical vowels of the character image and classify the character into the six patterns. After that, the character is divided to the unit of the consonant and vowel. The vowels are recognized by using the maximum length projection. The separated consonants are recognized by comparing the inputted pattern with the standard pattern that has the phase information of the boundary line change. We realized a system that the recognized characters are inputted to the word editor with the editable KS Hangul completion type code.
https://doi.org/10.5391/JKIIS.2003.13.1.001 인용 PDF KSCI

A Study on the Arabic numeral reading rules in Modern Korean (현대 한국어에서 아라비안 숫자의 읽기 규칙 연구)

Jung, Young-Im;Kim, Jeong-Se;Kim, Sang-Hoon;Lee, Young-Jik;Yoon, Ae-Sun
- Annual Conference on Human and Language Technology
- /
- 2002.10e
- /
- pp.16-23
- /
- 2002
본 논문에서는 아라비안 숫자를 포함한 텍스트를 음성으로 합성하기 위하여, 숫자 형태와 분류사 그리고 숫자가 나오는 문맥에 따라 숫자를 자동으로 문자화할 수 있는 전처리 규칙을 설정하는데 목적을 둔다. 먼저 선행연구를 통해 숫자를 포함한 수사 및 수사표현의 읽기 규칙의 적용 범위 및 한계점을 살펴보고, 음성 합성을 위한 아라비안 숫자의 문자화 규칙을 설정하고자 한다. 현대 한국어에서 아라비안 숫자를 읽는 방식은 크게 고유어 방식과 한자어 방식이 있으며 단(單)단위에서는 영어가 사용되기도 한다. 또한 한자어 방식에서도 단위를 붙여 읽는 경우와 모든 수를 단 단위로 읽는 경우가 있으므로, 아라비안 숫자의 문자화를 단순한 규칙을 설정하여 자동화하기에는 중의성이 높다. 본 연구에서는 (1) 숫자 전 전치어(pre-numeral), (2) 기호를 포함한 숫자열의 표현 형식과 크기, (3) 단위 표현, (4) 숫자 후치어(post-numeral), (5) 분류사(classifier) (6) 분류사 후치어(post-classifier), (7) 수사표현 앞뒤 문맥에 따라, 아라비안 숫자표현이 문자화되는 방식을 살펴보았다. 분석 대상 말뭉치는 C 신문의 2000년 1월부터 2000년 4월까지 전체 기사 1,400건에서 숫자가 포함된 숫자표현 약 63,000개론 구성하였다. 패턴화된 구조 및 중의성이 없는 구조를 12가지로 밝히고 중의성이 있는 구조의 유형을 밝혔으며 분류사 후치어와의 결합 관계, 좌우 문맥정보를 통해 중의성 해결의 단서를 제시하고자 하였다.
PDF

Implementation of a Spam Message Filtering System using Sentence Similarity Measurements (문장유사도 측정 기법을 통한 스팸 필터링 시스템 구현)

Ou, SooBin;Lee, Jongwoo
- KIISE Transactions on Computing Practices
- /
- v.23 no.1
- /
- pp.57-64
- /
- 2017
Short message service (SMS) is one of the most important communication methods for people who use mobile phones. However, illegal advertising spam messages exploit people because they can be used without the need for friend registration. Recently, spam message filtering systems that use machine learning have been developed, but they have some disadvantages such as requiring many calculations. In this paper, we implemented a spam message filtering system using the set-based POI search algorithm and sentence similarity without servers. This algorithm can judge whether the input query is a spam message or not using only letter composition without any server computing. Therefore, we can filter the spam message although the input text message has been intentionally modified. We added a specific preprocessing option which aims to enable spam filtering. Based on the experimental results, we observe that our spam message filtering system shows better performance than the original set-based POI search algorithm. We evaluate the proposed system through extensive simulation. According to the simulation results, the proposed system can filter the text message and show high accuracy performance against the text message which cannot be filtered by the 3 major telecom companies.
https://doi.org/10.5626/KTCP.2017.23.1.57 인용 KSCI

Destination Address Block Location on Machine-printed and Handwritten Korean Mail Piece Images (인쇄 및 필기 한글 우편영상에서의 수취인 주소 영역 추출 방법)

정선화;장승익;임길택;남윤석
- Journal of KIISE:Software and Applications
- /
- v.31 no.1
- /
- pp.8-19
- /
- 2004
In this paper, we propose an efficient method for locating destination address block on both of machine-Printed and handwritten Korean mail piece images. The proposed method extracts connected components from the binary mail piece image, generates text lines by merging them, and then groups the text fines into nine clusters. The destination address block is determined by selecting some clusters. Considering the geometric characteristics of address information on Korean mail piece, we split a mail piece image into nine areas with an equal size. The nine clusters are initialized with the center coordinate of each area. A modified Manhattan distance function is used to compute the distance between text lines and clusters. We modified the distance function on which the aspect ratio of mail piece could be reflected. The experiment done with live Korean mail piece images has demonstrated the superiority of the Proposed method. The success rate for 1, 988 testing images was about 93.56%.
PDF KSCI

Developing XML based multilingual language education system (다국어 학습을 위한 XML기반 학습시스템의 설계)

Jeong, Hwi-Woong;Yoon, Ae-Sun
- Annual Conference on Human and Language Technology
- /
- 1999.10e
- /
- pp.407-412
- /
- 1999
XML은 언어정보의 재사용성 및 다른 유형의 정보로 변환이 용이하여 최근 그 사용이 급증하고 있다. 그러나 XML은 아직까지 일부 분야에 국한되어 이용되고 있으며, 국내에서도 XML을 실제 활용하여 개발되고 있는 시스템은 극히 미약하다. 본 연구에서는 XML의 이점을 살려 한글을 포함한 다국어간 언어학습 컨텐트를 쉽게 구성하고 가공할 수 있는 XML 문서 내의 다국어 표현 방법에 대해 연구하였다. 또한 다국어 정보를 웹 환경에서 구현하기 위한 XSL과 유사한 문서 변환 구조 및 이를 처리할 수 있는 XML 처리기의 구조에 대해서도 소개한다. 본 연구에서 소개하는 문서 변환 구조를 이용할 경우 문자로 표현 가능한 매체를 매개로 하여 다양한 멀티미디어 컨텐트를 쉽게 작성할 수 있다.
PDF

Approach to develop a software supporting sign design strategy -Focusing on the letter information- (사인디자인 지원 소프트웨어 개발을 위한 방안 -문자정보를 중심으로-)

Paik, Jin-Kyung;Choi, In-Kyu;Shim, Eun-Mi;Lee, Kyung-Mi
- Archives of design research
- /
- v.17 no.4
- /
- pp.149-158
- /
- 2004
Domestic sign industry has seen a rapid growth in recent years, however, the level of quality is not so high. The main reason for this is due to the lack of well educated specialists. Thus, in this investigation, we will develop sign design experience software for the experienced employees in sign industry. Especially, focusing on the indoor letter information, we will develop the software that can provide the graphic effect associated with letter information. First, current situation and preliminary study in sign system used in public information were conducted, then design elements were analyzed from this. Then, experiments were performed on teh analyzed elements, and evaluation methods for the sign design were determined. Experiments including typography, color, layout, line alignment, letter space, and line space were conducted to analyze the visual perception. From these experiments, effective elements can be extracted and used as a elements for sign design software. As a next step, we will propose the development of experience software in sign design that can help to determine the sign elements. Thus, our investigation will provide sign design experience software that can be used by any inexperienced individuals, and be a great help for the more effective sign planning.
PDF

Variance Recovery in Text Detection using Color Variance Feature (색 분산 특징을 이용한 텍스트 추출에서의 손실된 분산 복원)

Choi, Yeong-Woo;Cho, Eun-Sook
- Journal of the Korea Society of Computer and Information
- /
- v.14 no.10
- /
- pp.73-82
- /
- 2009
This paper proposes a variance recovery method for character strokes that can be missed in applying the previously proposed color variance approach in text detection of natural scene images. The previous method has a shortcoming of missing the color variance due to the fixed length of horizontal and vertical windows of variance detection when the character strokes are thick or long. Thus, this paper proposes a variance recovery method by using geometric information of bounding boxes of connected components and heuristic knowledge. We have tested the proposed method using various kinds of document-style and natural scene images such as billboards, signboards, etc captured by digital cameras and mobile-phone cameras. And we showed the improved text detection accuracy even in the images of containing large characters.
https://doi.org/10.9708/jksci.2009.14.10.073 인용 PDF

Spelling Correction in Korean Using the `Eojeol` generation Dictionary (어절 생성 사전을 이용한 한국어 철자 교정)

Lee, Yeong-Sin;Park, Yeong-Ja;Song, Man-Seok
- The KIPS Transactions:PartB
- /
- v.8B no.1
- /
- pp.98-104
- /
- 2001
본 논문에서는 어절 생성 사전을 이용한 한국어 철자 교정을 제안한다. 어절 생성 사전은 두 문자열 간 음절 특성이 고려된 편집 거리 계산을 기반으로 탐색되어 언어와 오류 유형에 의존적인 정보를 이용하지 않고 오류 어절에 대한 후보 어절을 생성한다. 또한 교정된 어절들의 가능한 형태소 분석들을 산출하여 후보들 간의 순위 계산 시에 재차 형태소 분석을 수행하지 않고 언어 정보를 적용할 수 있다. 본 논문에서 제안하는 철자 교정은 두 단계로 구성된다. 첫째, 오류 어절로부터 가능한 오류 정정 어간들을 계산한다. 둘째, 계산된 어간들로부터 어절 생성 사전을 탐색하여 원형 후보 어절들을 생성한다. 또한 품사 태깅과 공기 정보를 사용하여 오류 수정된 결과의 순위를 매긴다. 본 시스템의 자동 철자 교정 성능을 평가한 결과 3,000개의 어절에서 시험한 결과 단어 수준으로 93%가 옳게 교정되었다.
PDF

Search Result 57, Processing Time 0.023 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)