Search | Korea Science

Implementation of JBIG2 CODEC with Effective Document Segmentation (문서의 효율적 영역 분할과 JBIG2 CODEC의 구현)

백옥규;김현민;고형화
- The Journal of Korean Institute of Communications and Information Sciences
- /
- v.27 no.6A
- /
- pp.575-583
- /
- 2002
JBIG2 is an International Standard fur compression of Bi-level images and documents. JBIG2 supports three encoding modes for high compression according to region features of documents. One of which is generic region coding for bitmap coding. The basic bitmap coder is either MMR or arithmetic coding. Pattern matching coding method is used for text region, and halftone pattern coding is used for halftone region. In this paper, a document is segmented into line-art, halftone and text region for JBIG2 encoding and JBIG2 CODEC is implemented. For efficient region segmentation of documents, region segmentation method using wavelet coefficient is applied with existing boundary extraction technique. In case of facsimile test image(IEEE-167a), there is improvement in compression ratio of about 2% and enhancement of subjective quality. Also, we propose arbitrary shape halftone region coding, which improves subjective quality in talc neighboring text of halftone region.
PDF KSCI

The Character Area Extraction and the Character Segmentation on the Color Document (칼라 문서에서 문자 영역 추출믹 문자분리)

김의정
- Journal of the Korean Institute of Intelligent Systems
- /
- v.9 no.4
- /
- pp.444-450
- /
- 1999
This paper deals with several methods: the clustering method that uses k-means algorithm to abstract the area of characters on the image document and the distance function that suits for the HIS coordinate system to cluster the image. For the prepossessing step to recognize this, or the method of characters segmentate, the algorithm to abstract a discrete character is also proposed, using the linking picture element. This algorithm provides the feature that separates any character such as the touching or overlapped character. The methods of projecting and tracking the edge have so far been used to segment them. However, with the new method proposed here, the picture element extracts a discrete character with only one-time projection after abstracting the character string. it is possible to pull out it. dividing the area into the character and the rest (non-character). This has great significance in terms of processing color documents, not the simple binary image, and already received verification that it is more advanced than the previous document processing system.
PDF

A System for the Decomposition of Text Block into Words (텍스트 영역에 대한 단어 단위 분할 시스템)

Jeong, Chang-Boo;Kwag, Hee-Kue;Jeong, Seon-Hwa;Kim, Soo-Hyung
- Proceedings of the Korea Information Processing Society Conference
- /
- 2000.10a
- /
- pp.293-296
- /
- 2000
본 논문에서는 주제어 인식에 기반한 문서영상의 검색 및 색인 시스템에 적용하기 위한 단어 단위 분한 시스템을 제안한다. 제안 시스템은 영상 전처리, 문서 구조 분석을 통해 추출된 텍스트 영역을 입력으로 단어 단위 분할을 수행하는데, 텍스트 영역에 대해 텍스트 라인을 분할하고 분할된 텍스트 라인을 단어 단위로 분할하는 계층적 접근 방법을 사용한다. 텍스트라인 분할은 수평 방향 투영 프로파일을 적용하여 분할 지점을 구한다. 그리고 단어 분할은 연결요소들을 추출한 후 연결요소간의 gap 정보를 구하고, gap 군집화 기법을 사용하여 단어 단위 분한 지점을 구한다. 이때 단어 단위 분할의 성능을 저하시키는 특수기호에 대해서는 휴리스틱 정보를 이용하여 검출한다. 제안 시스템의 성능 평가는 50개의 텍스트 영역에 적용하여 99.83%의 정확도를 얻을 수 있었다.
PDF

Structure Analysis of Low Contrast Fax Cover Pages (저해상도 팩스 표지 영상의 구조 분석)

임영규;이성환
- Proceedings of the Korean Information Science Society Conference
- /
- 1998.10c
- /
- pp.387-389
- /
- 1998
팩스가 보편적인 정보 전달 매체로 자리잡게 됨에 따라 기업체나 관공서 뿐만 아니라 가정에서도 많은 작업이 팩스를 통해 이루어지게 되었다. 이에 따라 팩스 문서의 분석 및 인식의 필요성이 증가하게 되었다. 팩스 문서는 표지와 내용이 두 부분으로 이루어지는데 팩스 문서의 처리를 위해서는 성명, 주소등을 포함하는 팩스 표지의 분석이 중요하다. 따라서 본 논문에서는 팩스 표지 영상의 구조 분석 방법을 제안한다. 제안한 팩스 표지 구조 분석 방법은 팩스 표지가 헤드, 송/수신 정보, 메시지로 구성된다는 점에 착안하여 위치 정보를 이용한 영역 분리에 중점을 두었으며, 팩스 표지의 종류를 몇 가지로 분류하여 도표 형태의 팩스 표지도 분석이 가능하도록 하였다. 분자 인식에서는 팩스 문자 인식에 우수한 성능을 보이고 있는 자소 기반 한글 문자 인식기를 사용하였다. 또한 한글의 자소 모델에 기반한 후처리 방법을 개발하여 인식 오류를 교정하였다.
PDF

한글 문서 영상에서의 문자와 비문자의 분리 추출

Lee, Jong-Guk;Kim, Hang-Jun
- Annual Conference on Human and Language Technology
- /
- 1990.11a
- /
- pp.219-219
- /
- 1990
본 논문에서는 국제 컴퓨터 망을 통하여 한글 정보를 전송할 수 있는 한 방안을 제안하였다. 한글 문서를 서구 문자로 바꾸어 서구문만의 전송이 가능한 컴퓨터 망에서도 전송이 가능하도록 하였고 한글과 영문이 혼용된 문서를 서구 문자로 전자하는데 한/영 구분기호를 사용하였으며 전자된 한글이 포함되어 있음을 표현하는 통신문 서식을 만들어 사용하였다. 또한 한글을 서구문자로 전자하고 복원하는 소프트웨어를 작성하였다.
PDF

Digital Watermarking for Document Image (문서 이미지를 위한 디지털 워터마킹)

Li, De;Choi, Jonguk
- Proceedings of the Korean Information Science Society Conference
- /
- 2003.04a
- /
- pp.464-466
- /
- 2003
전자정부, 전자상거래의 활성화로 디지털 문서가 빠른 속도로 유통되고 있기에 이에 대한 효과적인 보호대척이 필요한 실정이다. 본 논문에서는 Window Pattern을 이용하여 문서 이미지에 저작권 정보를 삽입하는 방안을 제안한다. 삽입대상 Window Patte을 결정하고 이러한 Pattern으로 원본 영상을 Scan하면서 픽셀 값에 변화를 주게 된다. 이렇게 되어 하나의 Pattern에 1 bit의 정보의 삽입이 가능하게 되고 추출 시 원본은 필요로 하지 않으며 실용성이 높고 적용분야도 넓다.
PDF

Halftone Noise Removal in Scanned Images using HOG based Adaptive Smoothing Filter (HOG 기반의 적응적 평활화를 이용한 스캔된 영상의 하프톤 잡음 제거)

Hur, Kyu-Sung;Baek, Yeul-Min;Kim, Whoi-Yul
- Journal of Broadcast Engineering
- /
- v.17 no.2
- /
- pp.316-324
- /
- 2012
In this paper, a novel descreening method using HOG(histogram of gradient)-based adaptive smoothing filter is proposed. Conventional edge-oriented smoothing methods does not provide enough smoothing to the halftone image due to the edge-like characteristic of the halftone noise. Moreover, clustered-dot halftoning method, which is commonly used in printing tends to create Moire pattern because of the intereference in color channels. Therefore, the proposed method uses HOG to distinguish edges and the amount of smoothing to be performed on the halftone image is then calculated according to the magnitude of the HOG in the edge and edge normal orientation. The proposed method was tested on various scanned halftone materials, and the results show that it effectively removes halftone noises as well as Moire pattern while preserving image details.
https://doi.org/10.5909/JEB.2012.17.2.316 인용 PDF KSCI

Composite Document Object Retrieval and Searching System-[IN2] DOR (복합문서 개체 검색 시스템- [IN2] DOR)

Ahn, Tae-Sung;Yim, Joong-Su;Kim, Myung-Hoon;Ahn, Woo-Ram;Lee, Kyung-Il
- Annual Conference on Human and Language Technology
- /
- 2003.10d
- /
- pp.113-118
- /
- 2003
기존 문서 검색 시스템의 경우 단순히 문서 내에서 텍스트를 추출한 후 그 텍스트를 색인, 검색하는 형태를 가지고 있었다. 본 논문에서는 MS Word, Excel, HWP 등 다양한 형태의 문서에서 텍스트, 표, 이미지, 차트, 동영상 등의 문서 개체를 분석, 색인하고 이를 검색하는 시스템의 개발 방법을 제외하였다. 제안된 시스템은 문서의 내부 자료 구조를 CDML(Composite Document Markup Language)로 변환하고, 이를 색인, 저장함으로 기존의 전문 검색 시스템의 한계를 효과적으로 극복했으며, 문서 내의 검색 대상 개체로 자동 이동하고 하일라이팅 시키는 기술을 구현함으로 사용자 편익성을 높였다. 개발된 시스템의 성능을 평가한 결과, 다양한 문서 형식에 대해 평균 97% 이상의 CDML변환 성공률과 개체 검색 성공률을 보였으며, 이진 파일에서 직접 개체를 추출함으로 매우 높은 분석 및 색인 속도가 달성되었음을 확인할 수 있었다. 본 논문에서 소개된 새로운 패러다임의 문서 검색 솔루션을 통해 다양한 기술적 상업적 파급 효과가 기대되고 있다.
PDF

The Effect of Orthography on Electronic Character Reading and Comprehending Ability in Japanese Education using ICT (ICT를 활용한 일본어 교육에서 문장 표기 형식이 영상문자 낭독 및 내용 파악에 미치는 효과)

Kang, Shin-Cheol;Kim, Min-Ki
- The Journal of Korean Association of Computer Education
- /
- v.7 no.6
- /
- pp.85-93
- /
- 2004
We investigated the proper display environment for japanese electronic character reading lessons through the experiment with a projection TV and a computer. For the purpose of finding out the effect of prior learning activities at the context of authentic Japanese text orthography, which includes dual notation, words spacing, etc., we also made an experiment on comprehending the web documents which are extracted from japanese web sites. From the experimental results, we acquired a conclusion that two approaches are needed to enhance the ability of comprehending Japanese web documents which is newly added to the 7th curriculum revision. For short-term approach, we need to utilize Japanese web documents as learning materials. For long-term approach, we have to reconsider whether the orthography of the current Japanese textbooks is suitable or not.
PDF

Text line extraction based on filtering and peak detection (필터링 및 피크검출을 이용한 텍스트 추출)

Jin, Bora;Cho, Nam-Ik
- Proceedings of the Korean Society of Broadcast Engineers Conference
- /
- 2013.11a
- /
- pp.41-42
- /
- 2013
본 논문에서는 문서 영상 처리의 중요한 전처리 과정인 텍스트 라인 추출을 위하여 가우시안 필터링 및 피크 검출을 이용하는 방법을 제안한다. 이는 문서 영상 내의 글자 영역의 픽셀 강도와 텍스트 라인 사이의 간격에 해당하는 강도의 차이로 인해 문서 영상의 각 열마다 높은 피크와 낮은 피크가 번갈아 가며 나타나는 것에 기반으로, 제안하는 알고리즘은 필터 스케일 추정, 필터량 및 피크 검출, 라인 성분 그룹화의 세 단계로 구성된다. 필터 스케일 추정 단계에서는 여러 초기 값으로 필터링하여 피크 차이 간의 히스토그램을 만듦으로써 글자 크기를 대략적으로 예축하며, 필터링 및 피크 검출 단계에서 앞서 예측된 스케일의 가우시안 필터를 이용하여 필터링 한 후, 각각의 열마다 피크를 검출한다. 마지막으로 라인 성분 그룹화를 통하여 검출된 피크를 서로 연결하여 하나의 텍스트 라인을 구성하는 성분들로 그룹화시켜 텍스트 라인을 추출한다. 실험 결과를 통하여, 제안하는 알고리즘은 이진화 과정을 거치지 않음으로써 균일하지 못한 조명환경 등으로 이진화 성능이 좋지 못할 경우에도 텍스트 라인을 추출할 수 있으며, 텍스트 라인 간격이 인정하지 않고 휘어진 라인을 포함하는 경우에도 적용할 수 있음을 확인 할 수 있다.
PDF

Search Result 381, Processing Time 0.026 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)