• Title/Summary/Keyword: Captions

Search Result 64, Processing Time 0.027 seconds

Using similarity based image caption to aid visual question answering (유사도 기반 이미지 캡션을 이용한 시각질의응답 연구)

  • Kang, Joonseo;Lim, Changwon
    • The Korean Journal of Applied Statistics
    • /
    • v.34 no.2
    • /
    • pp.191-204
    • /
    • 2021
  • Visual Question Answering (VQA) and image captioning are tasks that require understanding of the features of images and linguistic features of text. Therefore, co-attention may be the key to both tasks, which can connect image and text. In this paper, we propose a model to achieve high performance for VQA by image caption generated using a pretrained standard transformer model based on MSCOCO dataset. Captions unrelated to the question can rather interfere with answering, so some captions similar to the question were selected to use based on a similarity to the question. In addition, stopwords in the caption could not affect or interfere with answering, so the experiment was conducted after removing stopwords. Experiments were conducted on VQA-v2 data to compare the proposed model with the deep modular co-attention network (MCAN) model, which showed good performance by using co-attention between images and text. As a result, the proposed model outperformed the MCAN model.

Generate Korean image captions using LSTM (LSTM을 이용한 한국어 이미지 캡션 생성)

  • Park, Seong-Jae;Cha, Jeong-Won
    • Annual Conference on Human and Language Technology
    • /
    • 2017.10a
    • /
    • pp.82-84
    • /
    • 2017
  • 본 논문에서는 한국어 이미지 캡션을 학습하기 위한 데이터를 작성하고 딥러닝을 통해 예측하는 모델을 제안한다. 한국어 데이터 생성을 위해 MS COCO 영어 캡션을 번역하여 한국어로 변환하고 수정하였다. 이미지 캡션 생성을 위한 모델은 CNN을 이용하여 이미지를 512차원의 자질로 인코딩한다. 인코딩된 자질을 LSTM의 입력으로 사용하여 캡션을 생성하였다. 생성된 한국어 MS COCO 데이터에 대해 어절 단위, 형태소 단위, 의미형태소 단위 실험을 진행하였고 그 중 가장 높은 성능을 보인 형태소 단위 모델을 영어 모델과 비교하여 영어 모델과 비슷한 성능을 얻음을 증명하였다.

  • PDF

Development of Video Caption Editor with Kinetic Typography (글자가 움직이는 동영상 자막 편집 어플리케이션 개발)

  • Ha, Yea-Young;Kim, So-Yeon;Park, In-Sun;Lim, Soon-Bum
    • Journal of Korea Multimedia Society
    • /
    • v.17 no.3
    • /
    • pp.385-392
    • /
    • 2014
  • In this paper, we developed an Android application named VIVID where users can edit the moving captions easily on smartphone videos. This makes it convenient to set the time range, text, location and motion of caption text on the video. The editing result is uploaded to web server in html and can be shared with other users.

Knowledge-Based Numeric Open Caption Recognition for Live Sportscast

  • Sung, Si-Hun
    • Proceedings of the IEEK Conference
    • /
    • 2003.07e
    • /
    • pp.1871-1874
    • /
    • 2003
  • Knowledge-based numeric open caption recognition is proposed that can recognize numeric captions generated by character generator (CG) and automatically superimpose a modified caption using the recognized text only when a valid numeric caption appears in the aimed specific region of a live sportscast scene produced by other broadcasting stations. in the proposed method, mesh features are extracted from an enhanced binary image as feature vectors, then a valuable information is recovered from a numeric image by perceiving the character using a multiplayer perceptron (MLP) network. The result is verified using knowledge-based hie set designed for a more stable and reliable output and then the modified information is displayed on a screen by CG. MLB Eye Caption based on the proposed algorithm has already been used for regular Major League Base-ball (MLB) programs broadcast five over a Korean nationwide TV network and has produced a favorable response from Korean viewer.

  • PDF

Expected Matching Score Based Document Expansion for Fast Spoken Document Retrieval (고속 음성 문서 검색을 위한 Expected Matching Score 기반의 문서 확장 기법)

  • Seo, Min-Koo;Jung, Gue-Jun;Oh, Yung-Hwan
    • Proceedings of the KSPS conference
    • /
    • 2006.11a
    • /
    • pp.71-74
    • /
    • 2006
  • Many works have been done in the field of retrieving audio segments that contain human speeches without captions. To retrieve newly coined words and proper nouns, subwords were commonly used as indexing units in conjunction with query or document expansion. Among them, document expansion with subwords has serious drawback of large computation overhead. Therefore, in this paper, we propose Expected Matching Score based document expansion that effectively reduces computational overhead without much loss in retrieval precisions. Experiments have shown 13.9 times of speed up at the loss of 0.2% in the retrieval precision.

  • PDF

Implement closed captioning systems for the deaf (청각 장애인을 위한 자막방송 시스템 구현)

  • Kim, Minho;Kang, Hyosoon
    • Journal of Korea Game Society
    • /
    • v.16 no.1
    • /
    • pp.103-110
    • /
    • 2016
  • The hearing impaired have a substantially lower comprehension of audiovisual television programs due to their inability to hear sounds. Therefore, research was needed to improve their understanding by making these programs more accessible. This thesis is based on a proposal to find a solution for automating closed captions.

Synchronization of VOD Content and Captions Using Speech Recognition and Modified Dynamic Programming (음성인식과 변경된 동적계획법을 이용한 VOD 콘텐트와 자막의 동기화)

  • Oh, Juhyun
    • Proceedings of the Korean Society of Broadcast Engineers Conference
    • /
    • 2021.06a
    • /
    • pp.131-134
    • /
    • 2021
  • 지상파 방송에서는 청각장애인을 위해 폐쇄자막(closed caption) 서비스가 제공되고 있지만, 이를 저장하여 VOD 서비스 등에 제공하고자 할 때는 영상과의 비동기화(desynchronization) 문제로 인해 활용할 수 없는 문제가 있다. 본 논문에서는 이를 해결하기 위해 자동 음성인식(automatic speech recognition)과, 자막 동기화 문제에 맞게 변경된 동적계획법(modified dynamic programming)을 이용하는 방법을 제안한다. 문자열 정렬에서 삽입과 삭제 등 간격(gap)의 발생을 제어하는 제약조건과 그에 따른 점수 구조를 적용함으로써 문자열 정렬 성능을 개선한다. 또한 정렬된 폐쇄자막과 음성인식 문자열로부터 시간 동기정보를 복원하고 동기화된 자막을 생성하는 방법을 제안한다. 실제 TV 프로그램과 자막에 적용하여 기존 방법에 비해 성능의 향상이 있음을 확인하였다.

  • PDF

Villard de Honnecourt: the Characteristics and Authors of the Sketchbook (Villard de Honnecourt: 스케치북의 저자와 특성)

  • Hong, Seong-Woo
    • Journal of architectural history
    • /
    • v.7 no.3 s.16
    • /
    • pp.107-120
    • /
    • 1998
  • Even though Gothic architecture, one of the most technologically complex sophisticated structural systems, has been interpreted by art and architectural historians since the nineteenth century, we still cannot entirely comprehend either the medieval builder's constructional technique and structural knowledge or the meaning of Gothic architectural elements. The major reason is that contemporaneous written documentation concerning design methods and constructional techniques of medieval architecture is lacking. In 1955, the Bibliotheque Nationale in Paris exhibited the sketchbook of the thirteenth century architect Villard do Honnecourt. After the exhibition, analysis on the architectural drawings of Villard's sketchbook had reported widely. Most of analysis on Villard, however, has been on his drawing and artistic style, and there has been very little published analysis of his profession and question on the author of the sketchbook. Thus, the purpose of this study is to investigate the characteristics of the sketchbook and identify the artist who drew it. The sketchbook poses a number of unsolved questions. There is no doubt that several hands have contributed some drawing with appropriate captions, particularly in the section devoted to the application of practical geometry to problems of masonry and carpentry. Scholars have assumed and revealed that it was not made by only one person, and it dealt too many different fields and styles. Through this study, the sketchbook drawings consist of five different styles and person (original painter, master1, master2, master3, and the last owner), and they, not Villard, just redrew the original drawings and bound the sketchbook. Therefore, Villard de Honnecourt was just a mentor of the sketchbook and he did not participate any writing and drawing in the sketchbook.

  • PDF

Extraction Analysis for Crossmodal Association Information using Hypernetwork Models (하이퍼네트워크 모델을 이용한 비전-언어 크로스모달 연관정보 추출)

  • Heo, Min-Oh;Ha, Jung-Woo;Zhang, Byoung-Tak
    • 한국HCI학회:학술대회논문집
    • /
    • 2009.02a
    • /
    • pp.278-284
    • /
    • 2009
  • Multimodal data to have several modalities such as videos, images, sounds and texts for one contents is increasing. Since this type of data has ill-defined format, it is not easy to represent the crossmodal information for them explicitly. So, we proposed new method to extract and analyze vision-language crossmodal association information using the documentaries video data about the nature. We collected pairs of images and captions from 3 genres of documentaries such as jungle, ocean and universe, and extracted a set of visual words and that of text words from them. We found out that two modal data have semantic association on crossmodal association information from this analysis.

  • PDF

Perceptions and Image Analysis of Elementary Students on Scientists studying Small Organisms (작은 생물을 연구하는 과학자에 대한 초등학생들의 인식 및 이미지 분석)

  • Choi, Youngmi;Hong, Seung-Ho
    • Journal of Korean Elementary Science Education
    • /
    • v.33 no.4
    • /
    • pp.655-673
    • /
    • 2014
  • We investigated perceptions and image analysis on scientists studying small organisms reflected in elementary student's drawing using a modified version of the Drawing-A-Scientist-Test. The participants were 530 of fifth and sixth graders consisted of 449 ordinary students and 81 science gifted students. The data were collected from associated words, images and explanatory notes depicted by students engaged in questionnaires. The results indicated that a larger number of students reminded small sized animals and/or plants as words associated with small organisms. In addition, some students depicted anthropomorphic or abstract microorganisms. In this study, more stereotypes of scientists' appearance were exhibited at sixth graders and city region group. Most of the students depicted indicators such as lab coat, glasses, scientific instruments for observing, indoor, male and young, whereas only a few students depicted collaborative work. There was statistically significant difference between girls and boys, because boys perceived male scientists only, while half of girls depicted female. More frequent research instruments and scientific captions were used when science gifted students depicted scientists studying small organisms. These results could be contributed to education on microorganisms in elementary science.