Search | Korea Science

Korean Image Caption Generator Based on Show, Attend and Tell Model (Show, Attend and Tell 모델을 이용한 한국어 캡션 생성)

Kim, Dasol;Lee, Gyemin
- Proceedings of the Korean Society of Broadcast Engineers Conference
- /
- 2022.11a
- /
- pp.258-261
- /
- 2022
최근 딥러닝 기술이 발전하면서 이미지를 설명하는 캡션을 생성하는 모델 또한 발전하였다. 하지만 기존 이미지 캡션 모델은 대다수 영어로 구현되어있어 영어로 캡션을 생성하게 된다. 따라서 한국어 캡션을 생성하기 위해서는 영어 이미지 캡션 결과를 한국어로 번역하는 과정이 필요하다는 문제가 있다. 이에 본 연구에서는 기존의 이미지 캡션 모델을 이용하여 한국어 캡션을 직접 생성하는 모델을 만들고자 한다. 이를 위해 이미지 캡션 모델 중 잘 알려진 Show, Attend and Tell 모델을 이용하였다. 학습에는 MS-COCO 데이터의 한국어 캡션 데이터셋을 이용하였다. 한국어 형태소 분석기를 이용하여 토큰을 만들고 캡션 모델을 재학습하여 한국어 캡션을 생성할 수 있었다. 만들어진 한국어 이미지 캡션 모델은 BLEU 스코어를 사용하여 평가하였다. 이때 BLEU 스코어를 사용하여 생성된 한국어 캡션과 영어 캡션의 성능을 평가함에 있어서 언어의 차이에 인한 결과 차이가 발생할 수 있으므로, 영어 이미지 캡션 생성 모델의 출력을 한국어로 번역하여 같은 언어로 모델을 평가한 후 최종 성능을 비교하였다. 평가 결과 한국어 이미지 캡션 생성 모델이 영어 이미지 캡션 생성 모델을 한국어로 번역한 결과보다 좋은 BLEU 스코어를 갖는 것을 확인할 수 있었다.
PDF

GAN-based Dance Performance Visual Background Generation Method using Emotion Analysis on Lyrics (가사의 감정 분석을 이용한 GAN 기반 댄스 공연 배경 생성 방법)

Yoon, Hyewon;Kwak, Jeonghoon;Sung, Yunsick
- Annual Conference of KIPS
- /
- 2020.05a
- /
- pp.530-531
- /
- 2020
최근 인공지능을 활용하여 예술 작품에 몰입할 수 있도록 무대 효과를 디자인하는 연구가 진행되고 있다. 무대 효과 중에서 무대 배경은 공연의 분위기를 형성한다. 춤의 장르별로 무대 배경에 사용되는 이미지를 생성하기 위해 소셜 미디어 기반 무대 배경 생성 시스템이 있다. 하지만 같은 장르 춤은 동일한 무대 배경 이미지가 제공되는 문제가 있다. 같은 장르의 춤이지만 노래의 분위기를 반영하여 차별된 무대 배경 이미지를 제공하는 것이 필요하다. 본 논문은 노래 가사의 감정을 활용하여 Generative Adversarial Network(GAN)을 통해 각 노래의 분위기를 고려한 무대 배경 이미지를 생성하는 방법을 제안한다. GAN은 노래에 포함된 단락별 감정 단어를 추출하여 스타일을 생성하도록 학습된다. 학습된 GAN은 노래 가사에 포함된 감정 단어를 활용하여 곡의 분위기를 반영한 무대 배경 이미지를 생성한다. 노래 가사를 고려하여 무대 배경 이미지를 생성함으로써 곡의 분위기가 고려된 무대 배경 이미지 생성이 가능하다.
https://doi.org/10.3745/PKIPS.y2020m05a.530 인용 PDF

An Image Management System for Fast Virtual Desktop Creation (빠른 가상 데스크탑 생성을 위한 이미지 관리 시스템)

Oh, Soo-Cheol;Cho, JungHyun;Kim, DaeWon;Kim, Seon-Uk;Kim, SeongWoon;Kim, HakYoung
- Annual Conference of KIPS
- /
- 2014.11a
- /
- pp.14-16
- /
- 2014
가상 데스크탑 시스템은 단일 물리적 서버상에 가상화 기술을 사용하여 다수의 가상 머신을 실행하고, 이를 네트워크로 연결된 클라이언트에서 사용하는 기술이다. 가상 데스크탑에는 저장장치의 역할을 하는 가상 데스크탑 이미지가 연결되며, 본 이미지에는 운영체제 및 필요한 응용 프로그램이 설치되어 배포된다. 따라서, 새로운 가상 데스크탑을 생성할 때 가상 데스크탑 이미지를 함께 생성해야 하며, 이는 저장장치를 사용한 작업으로 시간이 많이 소요되는 작업이다. 본 논문에서는 이미지 풀을 사용한 빠른 가상 데스크탑 생성 방안을 제안한다. 이미지 풀은 일정 수의 가상 데스크탑 이미지를 포함하고 있으며, 이미지 준비기는 이미지 풀에 있는 이미지의 개수가 일정하게 유지되도록 골든 이미지에서 복사해오는 역할을 담당한다. 가상 데스크탑 생성시, 이미지 풀에서 필요한 이미지를 지연시간 없이 바로 가져옴으로써, 가상 데스크탑 생성에 소요되는 시간을 감소시킬 수 있다.
https://doi.org/10.3745/PKIPS.y2014m11a.14 인용 PDF

An Edge Detection Technique for Performance Improvement of eGAN (eGAN 모델의 성능개선을 위한 에지 검출 기법)

Lee, Cho Youn;Park, Ji Su;Shon, Jin Gon
- KIPS Transactions on Software and Data Engineering
- /
- v.10 no.3
- /
- pp.109-114
- /
- 2021
GAN(Generative Adversarial Network) is an image generation model, which is composed of a generator network and a discriminator network, and generates an image similar to a real image. Since the image generated by the GAN should be similar to the actual image, a loss function is used to minimize the loss error of the generated image. However, there is a problem that the loss function of GAN degrades the quality of the image by making the learning to generate the image unstable. To solve this problem, this paper analyzes GAN-related studies and proposes an edge GAN(eGAN) using edge detection. As a result of the experiment, the eGAN model has improved performance over the existing GAN model.
https://doi.org/10.3745/KTSDE.2021.10.3.109 인용 PDF KSCI

Generate Korean image captions using LSTM (LSTM을 이용한 한국어 이미지 캡션 생성)

Park, Seong-Jae;Cha, Jeong-Won
- 한국어정보학회:학술대회논문집
- /
- 2017.10a
- /
- pp.82-84
- /
- 2017
본 논문에서는 한국어 이미지 캡션을 학습하기 위한 데이터를 작성하고 딥러닝을 통해 예측하는 모델을 제안한다. 한국어 데이터 생성을 위해 MS COCO 영어 캡션을 번역하여 한국어로 변환하고 수정하였다. 이미지 캡션 생성을 위한 모델은 CNN을 이용하여 이미지를 512차원의 자질로 인코딩한다. 인코딩된 자질을 LSTM의 입력으로 사용하여 캡션을 생성하였다. 생성된 한국어 MS COCO 데이터에 대해 어절 단위, 형태소 단위, 의미형태소 단위 실험을 진행하였고 그 중 가장 높은 성능을 보인 형태소 단위 모델을 영어 모델과 비교하여 영어 모델과 비슷한 성능을 얻음을 증명하였다.
PDF

A Study of Brush Stroke Generation Using Color Transfer (칼라변환을 이용한 브러쉬 스트로크의 생성에 관한 연구)

Park, Young-Sup;Yoon, Kyung-Hyun
- Journal of the Korea Computer Graphics Society
- /
- v.9 no.1
- /
- pp.11-18
- /
- 2003
본 논문에서는 회화적 렌더링에서 칼라변환을 이용한 브러쉬 스트로크의 생성에 관한 새로운 알고리즘을 제안한다. 본 논문의 브러쉬 스트로크 생성을 위한 전체적인 구성은 다음과 같다. 첫째, 두 장의 사진(한 장의 소스 이미지와 한 장의 참조 이미지)을 입력으로 하여 칼라 변환 이론을 적용하여 색상 테이블이 바뀐 새로운 이미지를 생성한다. 이 방법은 소스 이미지의 칼라 분포 형태를 창조 이미지의 칼라 분포 형태로 변환하기 위해, 선형 히스토그램 매칭이라 불리는, 간단한 통계학적 방법을 이용한다. 둘째, 가우시안 블러링과 소벨 필터를 이용하여 에지를 검출한다. 검출된 에지는 브러쉬 스트로크 렌더링 시 에지 부분에서 스트로크를 클리핑 함으로써 이미지의 윤곽선 보존을 위해 사용된다. 셋째, 브러쉬 스트로크의 방향을 결정하기 위한 방향맵을 생성한다. 방향맵은 입력 영상에 대한 영역 분할 및 병합을 토대로 만들어진다. 영역별 각 픽셀들에 대해 이미지 그래디언트에 기초한 일정한 방향을 부여함으로써 방향맵을 구성한다. 넷째, 구성된 방향맵을 참조하여 브러쉬 스트로크 생성의 기초가 되는 베지어 곡선(Bezier Curve)의 제어점(Control point)을 설정한다. 실제 회화작품에서 사용되는 브러쉬 스트로크는 일반적으로 곡선의 형태를 이루므로 곡선 표현이 가능한 베지어 곡선을 이용하여 브러쉬 스트로크를 표현하였다. 마지막으로, 생성된 브러쉬 스트로크를 에지부문에서 클리핑하고 배경색을 참조하여 블렌딩하거나 퐁 조명 모델을 이용하여 이미지에 적용하게 된다.
PDF

Image generation and classification using GAN-based Semi Supervised Learning (GAN기반의 Semi Supervised Learning을 활용한 이미지 생성 및 분류)

Doyoon Jung;Gwangmi Choi;NamHo Kim
- Smart Media Journal
- /
- v.13 no.3
- /
- pp.27-35
- /
- 2024
This study deals with a method of combining image generation using Semi Supervised Learning based on GAN (Generative Adversarial Network) and image classification using ResNet50. Through this, a new approach was proposed to obtain more accurate and diverse results by integrating image generation and classification. The generator and discriminator are trained to distinguish generated images from actual images, and image classification is performed using ResNet50. In the experimental results, it was confirmed that the quality of the generated images changes depending on the epoch, and through this, we aim to improve the accuracy of industrial accident prediction. In addition, we would like to present an efficient method to improve the quality of image generation and increase the accuracy of image classification through the combination of GAN and ResNet50.
https://doi.org/10.30693/SMJ.2024.13.3.27 인용 PDF

A Study on the Color of AI-Generated Images for Fashion Design -Focused on the Use of Midjourney (패션디자인을 위한 AI 생성 이미지 색상 비교 연구 -미드저니의 활용을 중심으로-)

Park, Keunsoo
- The Journal of the Convergence on Culture Technology
- /
- v.10 no.2
- /
- pp.343-348
- /
- 2024
Today, AI image creation programs are optimized for various and specialized purposes such as fashion product advertising, customized fashion style suggestions, and design development, and are actively utilized in the fashion industry. Meanwhile, color is a powerful formative element and plays an important role in expressing images for suggesting products or fashion styles. This study seeks to expand understanding of the use of Midjourney by identifying the characteristics of color combinations that appear in clothing images created using Midjourney among AI image creation tools. The results of this study are as follows. First, the initial image created in Midjourney reflects the existing image color used to create the image more than the color specified in the command. Second, the color combinations that appear in the clothes of the images created in Midjourney are divided into separate and mixed colors. The ratio of colors expressed in a separate color scheme is affected by the color order specified in the command. The number of colors combined in a mixed color scheme appears as a combination of fewer colors than the total number of colors of clothing in the existing image used to create the image in Midjourney and the number of colors specified in the command. Third, caution is needed because changes in background color can affect the user's color perception of the clothes in the image and the formation of the costume image. It is hoped that the results of this study will be helpful in fashion design education and practice.
https://doi.org/10.17703/JCCT.2024.10.2.343 인용 PDF

Image Caption Generation using Recurrent Neural Network (Recurrent Neural Network를 이용한 이미지 캡션 생성)

Lee, Changki
- Journal of KIISE
- /
- v.43 no.8
- /
- pp.878-882
- /
- 2016
Automatic generation of captions for an image is a very difficult task, due to the necessity of computer vision and natural language processing technologies. However, this task has many important applications, such as early childhood education, image retrieval, and navigation for blind. In this paper, we describe a Recurrent Neural Network (RNN) model for generating image captions, which takes image features extracted from a Convolutional Neural Network (CNN). We demonstrate that our models produce state of the art results in image caption generation experiments on the Flickr 8K, Flickr 30K, and MS COCO datasets.
https://doi.org/10.5626/JOK.2016.43.8.878 인용 KSCI

Constructing Panorama Image using Synthesized Homography (혼합된 호모그래피를 이용한 파노라마 이미지 생성)

Kim, Seong-Do;Uh, Young-Jung;Byun, Hye-Ran
- Proceedings of the Korean Information Science Society Conference
- /
- 2012.06b
- /
- pp.459-461
- /
- 2012
일반적으로 같은 장면을 찍은 여러 장의 이미지를 이용하여 파노라마를 생성하려는 경우에도 각 이미지 사이에는 많은 기하학적 제약이 존재하기 때문에 이미지들간의 관계를 단 하나의 호모그래피로 나타낼 수 없다. 하지만 현존하는 대부분의 파노라마 생성 알고리즘은 하나의 호모그래피를 이용하여 파노라마 이미지를 생성하는 방법을 이용하기 때문에 여러 가지 기하학적 제약을 제대로 나타낼 수 없다. 따라서 이러한 알고리즘을 이용한 파노라마의 결과 이미지는 많은 왜곡과 부정합을 포함하게 된다. 본 논문에서 우리는 이러한 문제를 해결하기 위하여 여러 개의 호모그래피를 생성하고 합성하여 파노라마 이미지를 생성하는 방법을 제안한다. 제안하는 방법을 통하여 기존 파노라마 생성 알고리즘에서 나타난 많은 왜곡과 부정합을 줄일 수 있으며 호모그래피 개수도 자동으로 판별하여 주기 때문에 사용자의 입력을 필요로 하지 않는다.

Search Result 1,543, Processing Time 0.029 seconds

이메일무단수집거부

이용약관

제 1 장 총칙

제 2 장 이용계약의 체결

제 3 장 계약 당사자의 의무

제 4 장 서비스의 이용

제 5 장 계약 해지 및 이용 제한

제 6 장 손해배상 및 기타사항

Detail Search

Image Search (β)