• Title/Summary/Keyword: syntactic

Search Result 717, Processing Time 0.022 seconds

Improved Character-Based Neural Network for POS Tagging on Morphologically Rich Languages

  • Samat Ali;Alim Murat
    • Journal of Information Processing Systems
    • /
    • v.19 no.3
    • /
    • pp.355-369
    • /
    • 2023
  • Since the widespread adoption of deep-learning and related distributed representation, there have been substantial advancements in part-of-speech (POS) tagging for many languages. When training word representations, morphology and shape are typically ignored, as these representations rely primarily on collecting syntactic and semantic aspects of words. However, for tasks like POS tagging, notably in morphologically rich and resource-limited language environments, the intra-word information is essential. In this study, we introduce a deep neural network (DNN) for POS tagging that learns character-level word representations and combines them with general word representations. Using the proposed approach and omitting hand-crafted features, we achieve 90.47%, 80.16%, and 79.32% accuracy on our own dataset for three morphologically rich languages: Uyghur, Uzbek, and Kyrgyz. The experimental results reveal that the presented character-based strategy greatly improves POS tagging performance for several morphologically rich languages (MRL) where character information is significant. Furthermore, when compared to the previously reported state-of-the-art POS tagging results for Turkish on the METU Turkish Treebank dataset, the proposed approach improved on the prior work slightly. As a result, the experimental results indicate that character-based representations outperform word-level representations for MRL performance. Our technique is also robust towards the-out-of-vocabulary issues and performs better on manually edited text.

Information Structure and the Use of the English Existential Construction in Korean Learner English

  • Lee, Hanjung
    • Journal of English Language & Literature
    • /
    • v.57 no.6
    • /
    • pp.1017-1041
    • /
    • 2011
  • This study investigates Korean EFL learners' awareness and use of the English existential there-construction by examining data collected from 54 Korean EFL learners of English by means of a pragmalinguistic judgment task and a controlled discourse completion task. The results of the judgment task reveal that lower proficiency learners rated canonical sentences and existentials with a preposed locative best in the communicative situations where the use of existentials would have been most appropriate. A comparison of the ratings by more proficient learners and native speakers shows that existentials received highest ratings by both groups where they are the most natural option, while canonical sentences received significantly higher ratings by the learners. With regard to the production data, learners tended to avoid existentials, but rather relied on canonical sentences. Existentials were rarely used by lower proficiency learners and not used productively even by more proficient learners in the situations where existentials would have been the most natural option. These results suggest that Korean learners' difficulty with the use of existentials is not merely a product of performance limitations, but attributable to limited knowledge about existentials and their syntactic alternatives in terms of contextual appropriateness. Lower proficiency learners lack such knowledge, and more proficient learners, while showing better awareness of the use of existentials, have problems as to the placement of new information when engaging in writing tasks that place lower level of demands on attention to the information status of noun phrases compared to communicative, oral tasks.

Vocabulary Analyzer Based on CEFR-J Wordlist for Self-Reflection (VACSR) Version 2

  • Yukiko Ohashi;Noriaki Katagiri;Takao Oshikiri
    • Asia Pacific Journal of Corpus Research
    • /
    • v.4 no.2
    • /
    • pp.75-87
    • /
    • 2023
  • This paper presents a revised version of the vocabulary analyzer for self-reflection (VACSR), called VACSR v.2.0. The initial version of the VACSR automatically analyzes the occurrences and the level of vocabulary items in the transcribed texts, indicating the frequency, the unused vocabulary items, and those not belonging to either scale. However, it overlooked words with multiple parts of speech due to their identical headword representations. It also needed to provide more explanatory result tables from different corpora. VACSR v.2.0 overcomes the limitations of its predecessor. First, unlike VACSR v.1, VACSR v.2.0 distinguishes words that are different parts of speech by syntactic parsing using Stanza, an open-source Python library. It enables the categorization of the same lexical items with multiple parts of speech. Second, VACSR v.2.0 overcomes the limited clarity of VACSR v.1 by providing precise result output tables. The updated software compares the occurrence of vocabulary items included in classroom corpora for each level of the Common European Framework of Reference-Japan (CEFR-J) wordlist. A pilot study utilizing VACSR v.2.0 showed that, after converting two English classes taught by a preservice English teacher into corpora, the headwords used mostly corresponded to CEFR-J level A1. In practice, VACSR v.2.0 will promote users' reflection on their vocabulary usage and can be applied to teacher training.

Using Syntax and Shallow Semantic Analysis for Vietnamese Question Generation

  • Phuoc Tran;Duy Khanh Nguyen;Tram Tran;Bay Vo
    • KSII Transactions on Internet and Information Systems (TIIS)
    • /
    • v.17 no.10
    • /
    • pp.2718-2731
    • /
    • 2023
  • This paper presents a method of using syntax and shallow semantic analysis for Vietnamese question generation (QG). Specifically, our proposed technique concentrates on investigating both the syntactic and shallow semantic structure of each sentence. The main goal of our method is to generate questions from a single sentence. These generated questions are known as factoid questions which require short, fact-based answers. In general, syntax-based analysis is one of the most popular approaches within the QG field, but it requires linguistic expert knowledge as well as a deep understanding of syntax rules in the Vietnamese language. It is thus considered a high-cost and inefficient solution due to the requirement of significant human effort to achieve qualified syntax rules. To deal with this problem, we collected the syntax rules in Vietnamese from a Vietnamese language textbook. Moreover, we also used different natural language processing (NLP) techniques to analyze Vietnamese shallow syntax and semantics for the QG task. These techniques include: sentence segmentation, word segmentation, part of speech, chunking, dependency parsing, and named entity recognition. We used human evaluation to assess the credibility of our model, which means we manually generated questions from the corpus, and then compared them with the generated questions. The empirical evidence demonstrates that our proposed technique has significant performance, in which the generated questions are very similar to those which are created by humans.

Specification-based Program Slicing and Its Applications (명세 기반 프로그램 슬라이싱 기법과 응용)

  • Chung, In-Sang;Yoon, Gwang-Sik;Lee, Wan-Kwon;Kwon, Yong-Rae
    • Journal of KIISE:Software and Applications
    • /
    • v.29 no.8
    • /
    • pp.529-542
    • /
    • 2002
  • More precise program slices could be obtained by considering the semantic relations between variables of interest, compared to the existing slicing techniques considering only the syntactic relations. In this paper, we present specification-based slicing that allows a better decomposition of the program by taking a specification as its slicing criterion. A specification-based slice consists of a subset of program statements which preserve the behavior and the correctness of the original Program with respect to a specification given by a pre-postcondition pair. Because specification-based slicing enables one to focus attention on only those program statements which realize the functional abstraction specified by the given specification, it can be widely used in many software engineering areas. Of its possible applications, we show how specification-based slicing can improve the Process for extracting reusable parts from existing programs and restructuring complex programs for better maintainability.

A Word Sense Disambiguation Method with a Semantic Network (의미네트워크를 이용한 단어의미의 모호성 해결방법)

  • DingyulRa
    • Korean Journal of Cognitive Science
    • /
    • v.3 no.2
    • /
    • pp.225-248
    • /
    • 1992
  • In this paper, word sense disambiguation methods utilizing a knowledge base based on a semantic network are introduced. The basic idea is to keep track of a set of paths in the knowledge base which correspond to the inctemental semantic interpretation of a input sentence. These paths are called the semantic paths. when the parser reads a word, the senses of this word which are not involved in any of the semantic paths are removed. Then the removal operation is propagated through the knowledge base to invoke the removal of the senses of other words that have been read before. This removal operation is called recusively as long as senses can be removed. This is called the recursive word sense removal. Concretion of a vague word's concept is one of the important word sense disambiguation methods. We introduce a method called the path adjustment that extends the conctetion operation. How to use semantic association or syntactic processing in coorporation with the above methods is also considered.

Automatic Generation of Voice Web Pages Based on SALT (SALT 기반 음성 웹 페이지의 자동 생성)

  • Ko, You-Jung;Kim, Yoon-Joong
    • Journal of KIISE:Software and Applications
    • /
    • v.37 no.3
    • /
    • pp.177-184
    • /
    • 2010
  • As a voice browser is introduced, voice dialog application becomes available on the Web environment. The voice dialog application consists of voice Web pages that need to translate the dialog scripts into SALT(Speech Application Language Tags). The current Web pages have been designed for visual. They, however, are potentially capable of using voice dialog. This paper, therefore, proposes an automated voice Web generation method that finds the elements for voice dialog from Web pages based HTML and converts them into SALT. The automatic generation system of a voice Web page consists of a lexical analyzer and a syntactic analyzer that converts a Web page which is described in HTML to voice Web page which is described in HTML+SALT. The converted voice Web page is designed to be able to handle not only the current mouse and keyboard input but also voice dialog.

Decision Tree based Disambiguation of Semantic Roles for Korean Adverbial Postpositions in Korean-English Machine Translation (한영 기계번역에서 결정 트리 학습에 의한 한국어 부사격 조사의 의미 중의성 해소)

  • Park, Seong-Bae;Zhang, Byoung-Tak;Kim, Yung-Taek
    • Journal of KIISE:Software and Applications
    • /
    • v.27 no.6
    • /
    • pp.668-677
    • /
    • 2000
  • Korean has the characteristics that case postpositions determine the syntactic roles of phrases and a postposition may have more than one meanings. In particular, the adverbial postpositions make translation from Korean to English difficult, because they can have various meanings. In this paper, we describe a method for resolving such semantic ambiguities of Korean adverbial postpositions using decision trees. The training examples for decision tree induction are extracted from a corpus consisting of 0.5 million words, and the semantic roles for adverbial postpositions are classified into 25 classes. The lack of training examples in decision tree induction is overcome by clustering words into classes using a greedy clustering algorithm. The cross validation results show that the presented method achieved 76.2% of precision on the average, which means 26.0% improvement over the method determining the semantic role of an adverbial postposition as the most frequently appearing role.

  • PDF

A Study on the Effect of the Changes in Temporary Exhibition Spaces of Korea's National and Public Museums on the Overall Space Structure of Museum - With Reference to Syntactic Relationship between the Most Integrated Space and Exhibition Space - (국내 국.공립 박물관 기획전시공간의 변화가 전체공간구조에 미치는 영향에 관한 연구 - 뮤지엄내 위상 중심공간과 기획전시실공간의 관계를 중심으로 -)

  • Kang, Hyun-Ji;Moon, Jung-Mook
    • Korean Institute of Interior Design Journal
    • /
    • v.21 no.1
    • /
    • pp.203-210
    • /
    • 2012
  • Since a private museums started in Europe 17C, many private museums established for high-class people like aristocrats to collect and to keep art works and to appreciate for limited members. After the French Revolution in 18C, the publicity became an important social issue through all European regions, and the museum gradually changed into public ones. Like that, as the concept of museum changed, its social role as well as its function was also changed. The concept of collection and display or preservation changed into the concept of exhibition and appreciation featuring the publicity. With the year-round exhibition, a classical concept, the planned-exhibition, a new active concept set as an important factor for a museum's projects. The latter concept embraces new social issues. Therefore as the space for planned-exhibitions reflecting social issues every season was needed, a museum sets its planned-exhibition space with the changeability, and gradually expands this kind of space in size. It is expected that planned-exhibition spaces characterized as the changeability may give some changes on the flow of a museum's overall space, and may have substantial influences on the flow. To recognize the changes in a planned-exhibition space's influence on the museum, this study selected some national, public museums having the planned-exhibition space, and investigated their influences on each museum's overall space structure through the analysis on space syntax. This study assumed the change of planned-exhibition space as the changes in the number of convex spaces, and measured it. And to understand the planned-exhibition's changes on a museum's overall spaces, such changed assumed as the numeric changes in convex spaces and measured them. In addition, the numeric changes's influence on the overall space structure was analyzed by measuring the overall space's average integration level. Through the above two factors, the 3 research methodologies and analyzed results were drawn out.

  • PDF

Prospective Changes of English Digital Textbook Based on the Universal Design for Learning (보편적 학습 설계에 근거한 영어과 디지털 교과서 개선 방안)

  • Kim, Jeong-ryeol
    • The Journal of the Korea Contents Association
    • /
    • v.15 no.7
    • /
    • pp.674-683
    • /
    • 2015
  • One of the issues with the textbooks pertinent to the current study is whether or not the Universal Design for Learning (UDL) factors have been dealt to satisfy students with different aptitudes in learning the core objectives of the lessons. This study develops a modified version of the UDL analysis criteria from the cross curricular criteria to language teaching and learning and uses it to analyze the sequence of digital English textbooks to investigate the descriptive statistics of the UDL factors in the new textbooks. The result shows that the textbook is designed most favorably to the students with the talent of linguistic aptitude and less favorably to the students with other types of aptitudes. The sequence analysis shows that sentence/word length and appearance of new words are incrementally sequenced as students advance upper grades. However, the syntactic complexity of middle school curves up steeply which is different from the elementary school textbooks. The UDL analysis will provide learning factors to consider when designing digital English textbooks to cover different aptitudinal groups.