• 제목/요약/키워드: Corpus-based Analysis

검색결과 200건 처리시간 0.02초

코퍼스 분석방법을 이용한 『동의보감(東醫寶鑑)』의 어휘 분석 (Corpus-based Analysis on Vocabulary Found in 『Donguibogam』)

  • 정지훈;김동율
    • 한국의사학회지
    • /
    • 제28권1호
    • /
    • pp.135-141
    • /
    • 2015
  • The purpose of this study is to analyze vocabulary found in "Donguibogam", one of the medical books in mid-Chosun, through Corpus-based analysis, one of the text analysis methods. According to it, Donguibogam has total 871,000 words in it, and Chinese characters used in it are total 5,130. Among them, 2,430 characters form 99% of the entire text. The most frequently appearing 20 Chinese characters are mainly function words, and with this, we can see that "Donguibogam" is a book equipped with complete forms of sentences just like other books. Examining the chapters of "Donguibogam" by comparison, Remedies and Acupuncture indicated lower frequencies of function words than Internal Medicine, External Medicine, and Miscellaneous Diseases. "Yixuerumen (Introduction to Medicine)" which influenced "Donguibogam" very much has lower frequencies of function words than "Donguibogam" in its most frequently appearing words. This may be because "Yixuerumen" maintains the form of Chileonjeolgu (a quatrain with seven Chinese characters in each line with seven-word lines) and adds footnotes below it. Corpus-based analysis helps us to see the words mainly used by measuring their frequencies in the book of medicine. Therefore, this researcher suggests that the results of this analysis can be used for education of Chinese characters at the college of Korean Medicine.

코퍼스를 통한 고등학교 영어교과서의 어휘 분석 (Usage analysis of vocabulary in Korean high school English textbooks using multiple corpora)

  • 김영미;서진희
    • 영어어문교육
    • /
    • 제12권4호
    • /
    • pp.139-157
    • /
    • 2006
  • As the Communicative Approach has become the norm in foreign language teaching, the objectives of teaching English in school have changed radically in Korea. The focus in high school English textbooks has shifted from mere mastery of structures to communicative proficiency. This paper will study five polysemous words which appear in twelve high school English textbooks used in Korea. The twelve text books are incorporated into a single corpus and analyzed to classify the usage of the selected words. Then the usage of each word was compared with that of three other corpora based sources: the BNC(British National Corpus) Sampler, ICE Singapore(International Corpus of English for Singapore) and Collins COBUILD learner's dictionary which is based on the corpus, "The Bank of English". The comparisons carried out as part of this study will demonstrate that Korean text books do not always supply the full range of meanings of polysemous words.

  • PDF

한국어교육학에서의 담화 연구 분석 (Issues of Discourse Studies in Korean Language Education)

  • 강현화
    • 한국어교육
    • /
    • 제23권1호
    • /
    • pp.219-256
    • /
    • 2012
  • The aim of this study is to observe the trend of discourse study in language education and analyze the main issues by investigating the literatures related to discourse in Korean language education in the last ten years. This study observed the discourse study conducted in Korean language education from the perspectives of study subject, study method and study data. Moreover, based on the results, it estimated the achievements and effectiveness of the discourse study conducted in Korean language education. The subject of discourse study was mainly dealt with discourse function, discourse pattern, discourse marker, discourse structure. In the study methods, analysis of corpus and survey were mainly used as the study methods, and spoken corpus, written corpus and semi-spoken corpus were used as study materials. In particular, the semi-spoken corpus was used at a very high rate among them. This showed that discourse study in Korean language education was mainly focused on spoken corpus study. This study divided the detailed field of Korean language education into four fields of linguistic knowledge, communication function, teaching activities and learning activities, and observed the trends of discourse study in each field. Overall, it was recognized that relatively many studies were focused on linguistic knowledge, particularly in pragmatic perspective. It can be said that the study based on discourse has a language educational effectiveness in that it is based on actual data and improves practical communication skills in the environment of various languages.

Using Corpora for Studying English Grammar

  • Kwon, Heok-Seung
    • 한국영어학회지:영어학
    • /
    • 제4권1호
    • /
    • pp.61-81
    • /
    • 2004
  • This paper will look at some grammatical phenomena which will illustrate some of the questions that can be addressed with a corpus-based approach. We will use this approach to investigate the following subjects in English grammar: number ambiguity, subject-verb concord, concord with measure expressions, and (reflexive) pronoun choice in coordinated noun phrases. We will emphasize the distinctive features of the corpus-based approach, particularly its strengths in investigating language use, as opposed to traditional descriptions or prescriptions of structure in English grammar. This paper will show that a corpus-based approach has made it possible to conduct new kinds of investigations into grammar in use and to expand the scope of earlier investigations. Native speakers rarely have accurate information about frequency of use. A large representative corpus (i.e., The British National Corpus) is one of the most reliable sources of frequency information. It is important to base an analysis of language on real data rather than intuition. Any description of grammar is more complete and accurate if it is based on a body of real data.

  • PDF

Metadiscourse in the Bank Negara Malaysia Governor's Speech Texts

  • Aziz, Roslina Abdul;Baharum, Norzie Diana
    • 아시아태평양코퍼스연구
    • /
    • 제2권2호
    • /
    • pp.1-15
    • /
    • 2021
  • The study aims to explore the use of metadiscourse in the Bank Negara Malaysia Governor's speeches based on Hyland's Interpersonal Model of Metadiscourse. The corpus data consist of 343 speech texts, which were extracted from the Malaysian Corpus of Financial English (MacFE), amounting to 688,778 tokens. Adopting both quantitative and qualitative approaches to data analysis the study investigates (1) the overall use of metadiscourse in the Bank Negara Governor's speech texts and (2) the functions of the most prominent metadiscourse resources used and their functions in the speech texts. The findings reveal that the Governor's speech texts to be interactional rather than interactive, revealing a rich distribution of interactional metadiscourse resources, namely engagement markers, self-mention, hedges, boosters and attitude markers throughout the texts. The interactional metadiscourse resources function to establish speaker-audience engagement and alignment of views, as well as to express degree of uncertainty and certainty and attitudes. The study concludes that the speech texts are not merely informational or propositional, but rather interpersonal.

Citation Practices in Academic Corpora: Implications for EAP Writing

  • Min, Su-Jung
    • 영어어문교육
    • /
    • 제10권3호
    • /
    • pp.113-126
    • /
    • 2004
  • Explicit reference to the work of other authors is an essential feature of most academic research writings. Corpus analysis of academic text can reveal much about what writers actually do and why they do so. Application of corpus tools in language education has been well documented by many scholars (Pedersen, 1995, Swales, 1990, Thompson, 2000). They demonstrate how computer technology can assist in the effective analysis of corpus based data. For teaching purposes, tills recent research provides insights in the areas of English for Academe Purposes (EAP). The need for such support is evident when students have to use appropriate citations in their writings. Using Swales' (1990) division of citation forms into integral and non-integral and Thompson and Tnbble's (2001) classification scheme, this paper codifies academic texts in a corpus. The texts are academic research articles from different disciplines. The results lead into a comparison of the citation practices m different disciplines. Finally, it is argued that the information obtained in this study is useful for EAP writing courses in EFL countries.

  • PDF

An Attempt to Measure the Familiarity of Specialized Japanese in the Nursing Care Field

  • Haihong Huang;Hiroyuki Muto;Toshiyuki Kanamaru
    • 아시아태평양코퍼스연구
    • /
    • 제4권2호
    • /
    • pp.57-74
    • /
    • 2023
  • Having a firm grasp of technical terms is essential for learners of Japanese for Specific Purposes (JSP). This research aims to analyze Japanese nursing care vocabulary based on objective corpus-based frequency and subjectively rated word familiarity. For this purpose, we constructed a text corpus centered on the National Examination for Certified Care Workers to extract nursing care keywords. The Log-Likelihood Ratio (LLR) was used as the statistical criterion for keyword identification, giving a list of 300 keywords as target words for a further word recognition survey. The survey involved 115 participants of whom 51 were certified care workers (CW group) and 64 were individuals from the general public (GP group). These participants rated the familiarity of the target keywords through crowdsourcing. Given the limited sample size, Bayesian linear mixed models were utilized to determine word familiarity rates. Our study conducted a comparative analysis of word familiarity between the CW group and the GP group, revealing key terms that are crucial for professionals but potentially unfamiliar to the general public. By focusing on these terms, instructors can bridge the knowledge gap more efficiently.

On the Analysis of Natural Language Processing Morphology for the Specialized Corpus in the Railway Domain

  • Won, Jong Un;Jeon, Hong Kyu;Kim, Min Joong;Kim, Beak Hyun;Kim, Young Min
    • International Journal of Internet, Broadcasting and Communication
    • /
    • 제14권4호
    • /
    • pp.189-197
    • /
    • 2022
  • Today, we are exposed to various text-based media such as newspapers, Internet articles, and SNS, and the amount of text data we encounter has increased exponentially due to the recent availability of Internet access using mobile devices such as smartphones. Collecting useful information from a lot of text information is called text analysis, and in order to extract information, it is performed using technologies such as Natural Language Processing (NLP) for processing natural language with the recent development of artificial intelligence. For this purpose, a morpheme analyzer based on everyday language has been disclosed and is being used. Pre-learning language models, which can acquire natural language knowledge through unsupervised learning based on large numbers of corpus, are a very common factor in natural language processing recently, but conventional morpheme analysts are limited in their use in specialized fields. In this paper, as a preliminary work to develop a natural language analysis language model specialized in the railway field, the procedure for construction a corpus specialized in the railway field is presented.

In My Opinion: Modality in Japanese EFL Learners' Argumentative Essays

  • Pemberton, Christine
    • 아시아태평양코퍼스연구
    • /
    • 제1권2호
    • /
    • pp.57-72
    • /
    • 2020
  • This study seeks to add to the current understanding of learners' use of modality in argumentative writing. A learner corpus of argumentative essays on four topics was created and compared to native English speaker data from the International Corpus Network of Asian Learners of English (ICNALE). The relationship between learners' use of modal devices (MDs) and the devices' appearance in the school's curriculum was also examined. The results showed that learners relied on a very narrow range of MDs compared to those in previous studies. The frequency of use of MDs varied based on the topic and did not seem to be driven by cultural factors as has been previously suggested. Learners used more hedges than boosters on all topics, contradicting most previous studies. Curriculum was determined to have a direct correlation with MD use, and other important factors may include perception of topic and overreliance on certain MDs over others (the One-to-One principal). This research implies that learners' perception of topic should be explored further as a variable affecting MD use. Curricula should be designed based on frequency of MD use by English native speakers, and learners should receive instruction that teaches the norms of MD use in academic writing. The methodology used in the study to determine correlations between MD use and the curriculum has a wide range of potential applications in the field of Contrastive Interlanguage Analysis.

초음파 진단장치를 이용한 축우의 번식효율증진에 관한 연구 II. 무발정 젖소에서 초음파검사 및 progesterone 농도측정에 의한 난소 구조물의 비교평가 (Use of ultrasonography for improving reproductive efficiency in cows II. Comparative evaluation of ovarian structures using ultrasonography and plasma progesterone analysis in subestrous dairy cows)

  • 손창호;강병규;최한선;강현구;백인석;서국현
    • 대한수의학회지
    • /
    • 제38권3호
    • /
    • pp.642-651
    • /
    • 1998
  • The accuracy of ultrasonography for determining the presence of a functional corpus luteum in subestrous dairy cows was investigated, using a radioimmunoassay for progesterone in plasma. Luteal status (high or low progesterone concentrations) was diagnosed in 534 cows, using B-mode transrectal ultrasonography. Accuracy of ultrasonography was 96.3% and 88.8% in the cows with and without functional corpus luteum, respectively. In 362 cows diagnosed with functional corpus luteum by ultrasonographic examination, 20 cows were diagnosed with the non-functional corpus luteum by analysis of plasma progesterone concentrations (false positive). In 172 cows with non-functional corpus luteum by ultrasonographic examination, 13 cows were diagnosed with the functional corpus luteum based on plasma progesterone assay (false negative). Most of the corpus luteum with well-defined border and homogeneous echotexture were diagnosed with functional corpus luteum. All cows that were not detected a corpus luteum by ultrasonographic examination were diagnosed as non-functional corpus luteum. The corpus luteum of cows that were diagnosed with false positive appeared homogeneous echotexture and above 15 mm in diameter, but the corpus luteum was the non-functional corpus luteum within Day 5 (Day 0 is ovulation day) or after Day 19. The corpus luteum of cows that were diagnosed with false negative appeared heterogeneous echogenicity and below 15 mm in diameter, but the corpus luteum was the functional corpus luteum after Day 5 or around Day 17. It was concluded that accuracy of ultrasonography was excellent for determining the presence of a functional corpus luteum in subestrous dairy cows. The corpus luteum that was diagnosed with false positive or false negative was the developing or regressing states. Thus, ultrasonography was required a serial examination of two or three times accurately diagnosing these corpus luteum.

  • PDF