• Title/Summary/Keyword: vocabulary translation

Search Result 34, Processing Time 0.023 seconds

Sentence Translation and Vocabulary Retention in an EFL Reading Class

  • Kim, Boram
    • English Language & Literature Teaching
    • /
    • v.18 no.2
    • /
    • pp.67-84
    • /
    • 2012
  • The present study investigated the effect of sentence translation as a production task on short-term and long-term retention of foreign vocabulary. 87 EFL university students at a beginning level, enrolled in reading class participated in the study. The study compared the performance of three groups on vocabulary recall: (1) Control group, (2) Translation group, and (3) Copy group. During the treatment sessions, translation group translated L1 sentences into English, while copy group simply copied given English sentences with each target word. Results of the immediate test were collected each week from week 2 to week 5 and analyzed by one-way ANOVA. Results revealed that regarding short-term vocabulary retention, participants in rote-copy condition outperformed those in translation group. Four weeks later a delayed test was administered to measure long-term vocabulary retention. In contrast, the results of two-way repeated measures ANOVA showed that long-term vocabulary retention of translation group was significantly greater than copy group. The findings suggest that although sentence translation is rather challenging to low-level learners, it may facilitate long-term retention of new vocabulary given the more elaborate and deeper processing the task entails.

  • PDF

O-JMeSH: creating a bilingual English-Japanese controlled vocabulary of MeSH UIDs through machine translation and mutual information

  • Soares, Felipe;Tateisi, Yuka;Takatsuki, Terue;Yamaguchi, Atsuko
    • Genomics & Informatics
    • /
    • v.19 no.3
    • /
    • pp.26.1-26.3
    • /
    • 2021
  • Previous approaches to create a controlled vocabulary for Japanese have resorted to existing bilingual dictionary and transformation rules to allow such mappings. However, given the possible new terms introduced due to coronavirus disease 2019 (COVID-19) and the emphasis on respiratory and infection-related terms, coverage might not be guaranteed. We propose creating a Japanese bilingual controlled vocabulary based on MeSH terms assigned to COVID-19 related publications in this work. For such, we resorted to manual curation of several bilingual dictionaries and a computational approach based on machine translation of sentences containing such terms and the ranking of possible translations for the individual terms by mutual information. Our results show that we achieved nearly 99% occurrence coverage in LitCovid, while our computational approach presented average accuracy of 63.33% for all terms, and 84.51% for drugs and chemicals.

A Comparative Study of Chinese Translations of 『Who ate all the Shinga?』 - Focusing on the Translation strategy of 4 types of Translations (『그 많던 싱아는 누가 다 먹었을까』의 중국어 번역본 비교 연구 - 4종 번역본의 번역전략을 중심으로)

  • YANG, LEI;MOON, DAE IL
    • The Journal of the Convergence on Culture Technology
    • /
    • v.8 no.1
    • /
    • pp.403-408
    • /
    • 2022
  • This study analyzed the translation strategies of four Chinese translations of 『Who ate all the Sing a?』. As is well known, Park Wan-seo's works contain many psychological descriptions, abstract vocabulary, idioms, proverbs, dialects, etc., so when translating into Chinese, various translation strategies such as translation, interpretation, and creative translation are required. Although all four types studied in this paper are somewhat different depending on the translator, all translation strategies were used in a comprehensive way. As a result of the study, all four translation strategies used a strategy of direct translation of Chinese characters when translating geographical namesand names of people. The interpretational translation strategy was used for the translation of vocabulary that requires historical, social, cultural, and geography background interpretation. was utilized. The creative translation strategy was used when translating overlapping issues, political and historically sensitive issues, and issues related to Korean pronunciation and grammar. Based on the results of this study, it is expected that translation strategy research on various Chinese translations of Korean modern literature as well as various Chinese translations of Park Wan-seo will be expanded.

Character-Level Neural Machine Translation (문자 단위의 Neural Machine Translation)

  • Lee, Changki;Kim, Junseok;Lee, Hyoung-Gyu;Lee, Jaesong
    • Annual Conference on Human and Language Technology
    • /
    • 2015.10a
    • /
    • pp.115-118
    • /
    • 2015
  • Neural Machine Translation (NMT) 모델은 단일 신경망 구조만을 사용하는 End-to-end 방식의 기계번역 모델로, 기존의 Statistical Machine Translation (SMT) 모델에 비해서 높은 성능을 보이고, Feature Engineering이 필요 없으며, 번역 모델 및 언어 모델의 역할을 단일 신경망에서 수행하여 디코더의 구조가 간단하다는 장점이 있다. 그러나 NMT 모델은 출력 언어 사전(Target Vocabulary)의 크기에 비례해서 학습 및 디코딩의 속도가 느려지기 때문에 출력 언어 사전의 크기에 제한을 갖는다는 단점이 있다. 본 논문에서는 NMT 모델의 출력 언어 사전의 크기 제한 문제를 해결하기 위해서, 입력 언어는 단어 단위로 읽고(Encoding) 출력 언어를 문자(Character) 단위로 생성(Decoding)하는 방법을 제안한다. 출력 언어를 문자 단위로 생성하게 되면 NMT 모델의 출력 언어 사전에 모든 문자를 포함할 수 있게 되어 출력 언어의 Out-of-vocabulary(OOV) 문제가 사라지고 출력 언어의 사전 크기가 줄어들어 학습 및 디코딩 속도가 빨라지게 된다. 실험 결과, 본 논문에서 제안한 방법이 영어-일본어 및 한국어-일본어 기계번역에서 기존의 단어 단위의 NMT 모델보다 우수한 성능을 보였다.

  • PDF

An Analysis on the Vocabulary in the English-Translation Version of Donguibogam Using the Corpus-based Analysis (코퍼스 분석방법을 이용한 『동의보감(東醫寶鑑)』 영역본의 어휘 분석)

  • Jung, Ji-Hun;Kim, Dong-Ryul;Kim, Do-Hoon
    • The Journal of Korean Medical History
    • /
    • v.28 no.2
    • /
    • pp.37-45
    • /
    • 2015
  • Objectives : A quantitative analysis on the vocabulary in the English translation version of Donguibogam. Methods : This study quantitatively analyzed the English-translated texts of Donguibogam with the Corpus-based analysis, and compared the quantitative results analyzing the texts of original Donguibogam. Results : As the results from conducting the corpus analysis on the English-translation version of Donguibogam, it was found that the number of total words (Token) was about 1,207,376, and the all types of used words were about 20.495 and the TTR (Type/Token Rate) was 1.69. The accumulation rate reaching to the high-ranking 1000 words was 83.54%, and the accumulation rate reaching to the high-ranking 2000 words was 90.82%. As the words having the high-ranking frequency, the function words like 'the, and of, is' mainly appeared, and for the content words, the words like 'randix, qi, rhizoma and water' were appeared in multi frequencies. As the results from comparing them with the corpus analysis results of original version of Donguibogam, it was found that the TTR was higher in the English translation version than that of original version. The compositions of function words and contents words having high-ranking frequencies were similar between the English translation version and the original version of Donguibogam. The both versions were also similar in that their statements in the parts of 'Remedies' and 'Acupuncture' showed higher composition rate of contents words than the rate of function words. Conclusions : The vocabulary in the English translation version of Donguibogam showed that this book was a book keeping the complete form of sentence and an Korean medical book at the same time. Meanwhile, the English translation version of Donguibogam had some problems like the unification of vocabulary due to several translators, and the incomplete delivery of word's meanings from the Chinese character-culture area to the English-culture area, and these problems are considered as the matters to be considered in a work translating Korean old medical books in English.

The effects of corpus-based vocabulary tasks on high school students' English vocabulary learning and attitude (코퍼스를 기반으로 한 어휘 과제가 고등학생의 영어 어휘 학습과 태도에 미치는 영향)

  • Lee, Hyun Jin;Lee, Eun-Joo
    • English Language & Literature Teaching
    • /
    • v.16 no.4
    • /
    • pp.239-265
    • /
    • 2010
  • This study investigates the effects of corpus-based vocabulary tasks on the acquisition of English vocabulary in an attempt to explore the influence of corpus use on EFL pedagogy. For this to be realized, a total of 40 Korean high school students participated in the study over a 4-week period. An experimental group used a set of corpus-based tasks for vocabulary learning, whereas a control group carried out a traditional task (i.e., the L1-L2 translation) for vocabulary learning. To assess learning gains, the students were asked to complete the pre- and post-treatment tests measuring the word form, meaning, and use aspects of target lexical items. Results of the study indicate that in the experimental group the corpus-based vocabulary tasks were beneficial for the learning of word forms and use. In particular, corpus-based benefits were greatest in the low-proficiency EFL learners' collocational aspects of vocabulary use. On the other hand, in the control group, the traditional vocabulary tasks benefited the meaning aspects of target vocabulary items the most. In addition, survey results revealed that most students were positive about the corpus-based learning experience although some expressed reservations about the heavy cognitive load and the time-consuming nature of the analysis of corpus data primarily due to learners' lack of language proficiency.

  • PDF

English-Korean Transfer Dictionary Extension Tool in English-Korean Machine Translation System (영한 기계번역 시스템의 영한 변환사전 확장 도구)

  • Kim, Sung-Dong
    • KIPS Transactions on Software and Data Engineering
    • /
    • v.2 no.1
    • /
    • pp.35-42
    • /
    • 2013
  • Developing English-Korean machine translation system requires the construction of information about the languages, and the amount of information in English-Korean transfer dictionary is especially critical to the translation quality. Newly created words are out-of-vocabulary words and they appear as they are in the translated sentence, which decreases the translation quality. Also, compound nouns make lexical and syntactic analysis complex and it is difficult to accurately translate compound nouns due to the lack of information in the transfer dictionary. In order to improve the translation quality of English-Korean machine translation, we must continuously expand the information of the English-Korean transfer dictionary by collecting the out-of-vocabulary words and the compound nouns frequently used. This paper proposes a method for expanding of the transfer dictionary, which consists of constructing corpus from internet newspapers, extracting the words which are not in the existing dictionary and the frequently used compound nouns, attaching meaning to the extracted words, and integrating with the transfer dictionary. We also develop the tool supporting the expansion of the transfer dictionary. The expansion of the dictionary information is critical to improving the machine translation system but requires much human efforts. The developed tool can be useful for continuously expanding the transfer dictionary, and so it is expected to contribute to enhancing the translation quality.

Research on Subword Tokenization of Korean Neural Machine Translation and Proposal for Tokenization Method to Separate Jongsung from Syllables (한국어 인공신경망 기계번역의 서브 워드 분절 연구 및 음절 기반 종성 분리 토큰화 제안)

  • Eo, Sugyeong;Park, Chanjun;Moon, Hyeonseok;Lim, Heuiseok
    • Journal of the Korea Convergence Society
    • /
    • v.12 no.3
    • /
    • pp.1-7
    • /
    • 2021
  • Since Neural Machine Translation (NMT) uses only a limited number of words, there is a possibility that words that are not registered in the dictionary will be entered as input. The proposed method to alleviate this Out of Vocabulary (OOV) problem is Subword Tokenization, which is a methodology for constructing words by dividing sentences into subword units smaller than words. In this paper, we deal with general subword tokenization algorithms. Furthermore, in order to create a vocabulary that can handle the infinite conjugation of Korean adjectives and verbs, we propose a new methodology for subword tokenization training by separating the Jongsung(coda) from Korean syllables (consisting of Chosung-onset, Jungsung-neucleus and Jongsung-coda). As a result of the experiment, the methodology proposed in this paper outperforms the existing subword tokenization methodology.

A Study on 『Korean Translation of ·』 -Focused on declared characteristics and characteristics in different versions- (『국역본 <>·<>』 고찰 -표기적 특징과 이본적 성격을 중심으로-)

  • Kan, Ho-yun
    • Journal of Korean Classical Literature and Education
    • /
    • no.15
    • /
    • pp.355-387
    • /
    • 2008
  • The purpose of the study was to decide Korean translation and the copying period of "Korean Translation of " and to look all around their characteristics in different versions carefully until now. The "Korean Translation" is a collection of Korean-translated romance and love stories excavated by a professor Kim,Il Geun, and there is not a little meaning in the context of novel history in the point of view of 'Korean translation of a court possession'. Arranging conclusion of the study generally, it is as follows. (1) Considering phonological phenomena, grammar and vocabulary in the study of Korean language, it is presumed that they would be translated into Korean and copied between the regime period of the King Sukjong and the regime period of the King Yungjo in the Joseon Dynasty. For, they were composed of a middle declaration of copied 'Myeoknambon "Korean Translation of Taepyeonggwanggi(태평광기)"' and 'NakseonJaebon(낙선재본)' between the middle of the 17th century and the middle of the 18th century and the regime period of the King Jeongjo in the Joseon Dynasty appointed as the background period of the novels should be excepted. Consequently, through the Korean Translation, we can confirm that the novel scope between the 17th century and the 18th century in Korean novel history was widened until 'The Royal Court' and 'Women'. (2) In the side of vocabulary, the "Korean Translation" also has not a little meaning in the side of a collection translated in the Royal Court. It doesn't have new vocabularies, but partial vocabularies as '(Traces:痕)' '(Clean eyes:明眸)', ' (Sail:帆)', '(Get up:起)', '글이플(Weak grass:弱草)', '쇼록(Owl:? 梟 or 鴉?)', '이 사라심(This life:此生)', and '노혀오매(Look for:訪)' are good data in the study of Korean language. (3) The "Korean Translation" is a valuable data about translation and copying of a court novel and we can discover intentionally changed parts and partially omitted sentences rather in the than in the . There are differences between a translation book and a copying book and we can catch sight of intention of translation and unsettledness of copying in the second work. Therefore, we can know that the "Korean Translation" has a double context which one work is translated and a work in different version is derived, compared to a simple copy. (4) The "Korean Translation" has a close relation with "Hangoldong(閒汨董)", but it doesn't regard the same copy as a foundation. The basic copy of translation of the "Korean Translation" is a different version of the same line as "Hangoldong" and "Jeochobon(저초본:정명기 소장본)" and is more similar line to "Hangoldong", but it is also not the same basic copy. (5) Considering that the "Korean Translation" doesn't has a distinct relation with the "Hangoldong", there is no correlation between the "Korean Translation" and and the "Hangoldong" and . In addition, we could not discover a writer's identity between the two.

Disease-Related Vocubulary and its translingual practice in Late 19th to Early 20th century (19세기 말 20세기 초 질병 어휘와 언어횡단적 실천)

  • Lee, Eunryoung
    • Journal of Sasang Constitutional Medicine
    • /
    • v.31 no.1
    • /
    • pp.65-78
    • /
    • 2019
  • Objectives This study aims to investigate how the Korean disease-related vocabulary is established or changed when it is translated into French or English. Through this, we examine changes in the meaning of diseases and the ecosystem of disease-related vocabulary in transition period of $19^{th}$ to $20^{th}$ century. Methods Korean disease-related vocabulary are extracted from a total of 148,000 Korean headwords included in our corpus of three bilingual dictionaries. Among them, the scope of analyisis is limited to group of vocabularies that include a high frequency words, disease(病) and symptom(症). Results The first type of change is the emergence of a neologism. In this case, coexistence of existing vocabulary and new words is observed. The second change is the appearance of loan words written in Hangul. The third is the case where the interpretation of meaning is changed while maintaining the word form. Finally, the fourth change is that the orthographic variants are displayed while maintaining the meaning of the existing vocabulary. Discussion Disease-related vocabulary increased greatly between 1897 and 1931. The increasing factor of vocabulary was the emergence of coined words, compound words and the influx of foreign words. The Korean language and the Western language made a new lexical form in order to introduce a new unknown concept to the Korean. We could also confirm that the way in which English word expanded its semantic field by modifying the way of representing the meaning of Korean Disease-related vocabulary.