• Title/Summary/Keyword: 렉

Search Result 17, Processing Time 0.022 seconds

Enhancement of a language model using two separate corpora of distinct characteristics

  • Cho, Sehyeong;Chung, Tae-Sun
    • Journal of the Korean Institute of Intelligent Systems
    • /
    • v.14 no.3
    • /
    • pp.357-362
    • /
    • 2004
  • Language models are essential in predicting the next word in a spoken sentence, thereby enhancing the speech recognition accuracy, among other things. However, spoken language domains are too numerous, and therefore developers suffer from the lack of corpora with sufficient sizes. This paper proposes a method of combining two n-gram language models, one constructed from a very small corpus of the right domain of interest, the other constructed from a large but less adequate corpus, resulting in a significantly enhanced language model. This method is based on the observation that a small corpus from the right domain has high quality n-grams but has serious sparseness problem, while a large corpus from a different domain has more n-gram statistics but incorrectly biased. With our approach, two n-gram statistics are combined by extending the idea of Katz's backoff and therefore is called a dual-source backoff. We ran experiments with 3-gram language models constructed from newspaper corpora of several million to tens of million words together with models from smaller broadcast news corpora. The target domain was broadcast news. We obtained significant improvement (30%) by incorporating a small corpus around one thirtieth size of the newspaper corpus.

Part-Of-Speech Tagging using multiple sources of statistical data (이종의 통계정보를 이용한 품사 부착 기법)

  • Cho, Seh-Yeong
    • Journal of the Korean Institute of Intelligent Systems
    • /
    • v.18 no.4
    • /
    • pp.501-506
    • /
    • 2008
  • Statistical POS tagging is prone to error, because of the inherent limitations of statistical data, especially single source of data. Therefore it is widely agreed that the possibility of further enhancement lies in exploiting various knowledge sources. However these data sources are bound to be inconsistent to each other. This paper shows the possibility of using maximum entropy model to Korean language POS tagging. We use as the knowledge sources n-gram data and trigger pair data. We show how perplexity measure varies when two knowledge sources are combined using maximum entropy method. The experiment used a trigram model which produced 94.9% accuracy using Hidden Markov Model, and showed increase to 95.6% when combined with trigger pair data using Maximum Entropy method. This clearly shows possibility of further enhancement when various knowledge sources are developed and combined using ME method.

Exploration on Tokenization Method of Language Model for Korean Machine Reading Comprehension (한국어 기계 독해를 위한 언어 모델의 효과적 토큰화 방법 탐구)

  • Lee, Kangwook;Lee, Haejun;Kim, Jaewon;Yun, Huiwon;Ryu, Wonho
    • Annual Conference on Human and Language Technology
    • /
    • 2019.10a
    • /
    • pp.197-202
    • /
    • 2019
  • 토큰화는 입력 텍스트를 더 작은 단위의 텍스트로 분절하는 과정으로 주로 기계 학습 과정의 효율화를 위해 수행되는 전처리 작업이다. 현재까지 자연어 처리 분야 과업에 적용하기 위해 다양한 토큰화 방법이 제안되어 왔으나, 주로 텍스트를 효율적으로 분절하는데 초점을 맞춘 연구만이 이루어져 왔을 뿐, 한국어 데이터를 대상으로 최신 기계 학습 기법을 적용하고자 할 때 적합한 토큰화 방법이 무엇일지 탐구 해보기 위한 연구는 거의 이루어지지 않았다. 본 논문에서는 한국어 데이터를 대상으로 최신 기계 학습 기법인 전이 학습 기반의 자연어 처리 방법론을 적용하는데 있어 가장 적합한 토큰화 방법이 무엇인지 알아보기 위한 탐구 연구를 진행했다. 실험을 위해서는 대표적인 전이 학습 모형이면서 가장 좋은 성능을 보이고 있는 모형인 BERT를 이용했으며, 최종 성능 비교를 위해 토큰화 방법에 따라 성능이 크게 좌우되는 과업 중 하나인 기계 독해 과업을 채택했다. 비교 실험을 위한 토큰화 방법으로는 통상적으로 사용되는 음절, 어절, 형태소 단위뿐만 아니라 최근 각광을 받고 있는 토큰화 방식인 Byte Pair Encoding (BPE)를 채택했으며, 이와 더불어 새로운 토큰화 방법인 형태소 분절 단위 위에 BPE를 적용하는 혼합 토큰화 방법을 제안 한 뒤 성능 비교를 실시했다. 실험 결과, 어휘집 축소 효과 및 언어 모델의 퍼플렉시티 관점에서는 음절 단위 토큰화가 우수한 성능을 보였으나, 토큰 자체의 의미 내포 능력이 중요한 기계 독해 과업의 경우 형태소 단위의 토큰화가 우수한 성능을 보임을 확인할 수 있었다. 또한, BPE 토큰화가 종합적으로 우수한 성능을 보이는 가운데, 본 연구에서 새로이 제안한 형태소 분절과 BPE를 동시에 이용하는 혼합 토큰화 방법이 가장 우수한 성능을 보임을 확인할 수 있었다.

  • PDF

Assessment of Wicking and Fast Dry Properties According to Moisture Transport Measurement Method of Knit and Woven Fabrics for Garment (의류소재용 직·편물의 수분이동 특성 측정 방법에 따른 흡한속건성 평가)

  • Kim, Hyun-ah;Kim, Seung-jin
    • Science of Emotion and Sensibility
    • /
    • v.20 no.2
    • /
    • pp.117-126
    • /
    • 2017
  • In this study, moisture transport characteristics for the woven and knitted fabrics made of 8 kinds of fiber materials using MMT (moisture management tester) were measured and discussed with the Bireck bt MMT and water evaporating rate (WER) measuring methods, which are vertical moisture transport methods. In addition, the drying property by MMT of the eight kinds of specimens was compared and discussed with the results measured by the vertical drying measurement. MMT experimental result which is horizental moisture transport appeared to be similar to the result of the Bireck method, which is the vertical moisture transport experiment. Absortion time measured from drip method of the fabrics made of the bamboo, linen, and cotton/nylon composite fabrics was short and thus they showed best wicking property, which was attributed to the low contact angle on the fabric surface and high porosity of the fabrics due to the staple yarn structure composed of the hydrophilic staple fibers. In drying property of the fabric specimens by MMT, maximum absorption radius of the dry-zone knit and bamboo woven fabrics were the highest and they showed the best drying property, which was a little different result compared with vertical drying measurement method. Half time of the drying rate in the MMT method was highly correlated with the fabric thickness and saturated moisture absortion rate and their regression coefficients were 0.9 and 0.88, respectively. This means that the knitted and woven fabric design technology for retaining good wicking and drying properties of the fabrics with thin fabric thickness is very important for obtaining high functional wear comfort fabrics. In addition, wicking and drying properties of the fabrics made of different fiber materials and with different yarns and fabric structures showed different results according to the measuring methods.

Economics and Ground Cover Growth Characteristics of a New Method of Shallow Soil Artificial Foundation Planting (저토심 인공지반 녹화공법의 경제성 및 도입 가능한 지피식물의 생육특성)

  • Choi, Jin-Woo;Kim, Hag-Kee;Lee, Kyong-Jae;Kang, Hyun-Kyung
    • Journal of the Korean Institute of Landscape Architecture
    • /
    • v.37 no.5
    • /
    • pp.98-108
    • /
    • 2009
  • The purpose of this study is to analyze the characteristics of limited methods, economics and breeding appropriateness of native and imported ground cover plants in the methodology of a shallow soil rooftop garden. The new shallow soil rooftop gardening method uses a total of 13cm in soil thickness, including 4.5cm of top soil on a 7.5cm rock-wool-mat stacked onto a 1cm roll-type-draining plate. The total construction cost for each method of soil level within the design price standard for SEDUM BLOCK is 89,433won/$m^2$, and for DAKU is 92,550won/$m^2$. By comparing those two methods, the construction cost of the shallow soil artificial foundation methodology is 45,000won/$m^2$; this shows the new method is 50% less expensive than the existing method of shallow soil rooftop gardening. The experiment was executed on the rooftop of the Korean National Housing Corporation to ensure validity of the shallow soil artificial foundation planting, and the sample plants which were imported and grown now in native covering. A list investigating the growing plants was made of the cover rate in each plant class, both while alive and the dry plant weight. The native ground cover plants, Sedum kamtschaticum, Sedum middendorffianum, Allium senescens, Sedum sarmentosum, Aquilegia buergariana, and Caryopteris incana increased the cover rate, live weight and dry weight in the shallow soil artificial foundation method. Among the imported cover plants, Sedum sprium and Sedum reflexum, the cover rate increased and growth conditions improved. However, some species needed weed maintenance. After examination with the less expensive shallow soil artificial foundation method and growth analysis, it was found that rooftop gardens are a low-cost option and the growth of plants is great. This result shows the new method can contribute to the proliferation of rooftop gardens in urban settings.

Effect of Korean Mistletoe Extract and Lectin on the Preneoplastic Hepatic Lesion and Apoptosis in Experimental Hepatocarcinogenesis (실험적 간암모델에서 한국산 겨우살이 추출물 및 렉틴 투여가 전암성 병변의 생성 및 Apoptosis에 미치는 영향)

  • 김미정;김정희;이미숙
    • Journal of the Korean Society of Food Science and Nutrition
    • /
    • v.31 no.5
    • /
    • pp.782-787
    • /
    • 2002
  • This study was done to investigate the effects of Korean mistletoe water extract and lectin on the apoptosis and preneoplastic lesion in chemically induced rat hepatocarcinogenesis. To attain the above objectives, weanling Sprague-Dawley male rats were fed modified AIN-76 diets containing 10% corn oil for 9 weeks. One week after feeding starts, rats were intraperitoneally injected twice with a dose of diethylnitrosamine (DEN, 50 mg/kg body weight (BW). Rats were provided with 0.05% phenobarbital (PB) in drinking water from one week after DEN treatment until the end of experiment. During the period of PB treatment, rats were injected with mistletoe extract (100 $\mu\textrm{g}$/kg BW) and lectin (10 $\mu\textrm{g}$/kg BW) twice a week. At the end of 9th week, rats were sacrificed and the formation of hepatic glutathione S-transferase placental form positive (GST-P$^{+}$) foci, apoptosis, DNA fragmentation and apoptosis related proteins were determined respectively. The formation of GST-P$^{+}$foci was significantly decreased by mistletoe extract or lectin treatment. Although there was no effect on apoptosis and DNA fragmentation in hepatic tissue by mistletoe extract or lectin treatment, caspase-9 and fas-L were increased. These results suggest that Korean mistletoe extract and lectin have a potential to inhibit hepatocarcinogenesis by increasing apoptosis.sis.

Association of Genetic Variations with Pemetrexed-Induced Cytotoxicity in Non-Small Cell Lung Cancer Cells (비소세포폐암 세포주에서 pemetrexed의 세포독성과 유전학적 다형성과의 상관성 조사)

  • Yoon, Seong-Ae;Choi, Jung-Ran;Kim, Jeong-Oh;Shin, Jung-Young;Zhang, XiangHua;Kang, Jin-Hyoung
    • Journal of Life Science
    • /
    • v.20 no.1
    • /
    • pp.103-112
    • /
    • 2010
  • Pemetrexed has demonstrated clinical activity in non-small cell lung cancer (NSCLC) as well as other solid tumors. It transports into the cells via reduced folate carrier (RFC) and is polyglutamated by folypolyglutamate synthetase (FPGS). Pemetrexed directly inhibits several folate-dependent enzymes such as thymidylate synthase (TS), dihydrofolate reductase (DHFR), and glycinamide ribonucleotide formyltransferase (GARFT). We investigated the effects of genetic variations and the expression of RFC, FPGS, TS and DHFR enzymes on drug sensitivity to pemetrexed in NSCLC cells. Polymorphisms in RFC, FPGS, and DHFR were genotyped in four NSCLC cells - A549, PC14, HCC-1588, and H226. Real-time RT-PCR and Western blot was performed to evaluate mRNA transcripts and protein of these genes. The cytotoxicity of pemetrexed was measured by SRB assay. In PC14 and H226 cells, increased mRNA expressions of RFC and FPGS were associated with higher cytotoxicity to pemetrexed. 2R/2R genotype of TS and its increased mRNA expression were associated with drug resistance to pemetrexed in A549 cells, whereas 3R/3R genotype in TS with decreased mRNA expression was associated with higher sensitivity in H226 cells. After pemetrexed treatment, an inverse change of DHFR mRNA and protein expression was found. The strongest linkage disequilibrium (LD) was discovered between-1726C>T and -1188A>C SNP of DHFR gene. Our findings suggest the cytotoxic effect of pemetrexed may be associated with genetic polymorphisms and the expression level of genes involved in pemetrexed metabolisms in NSCLC cells.