• Title/Summary/Keyword: holistic scoring

Search Result 14, Processing Time 0.027 seconds

Dual-scale BERT using multi-trait representations for holistic and trait-specific essay grading

  • Minsoo Cho;Jin-Xia Huang;Oh-Woog Kwon
    • ETRI Journal
    • /
    • v.46 no.1
    • /
    • pp.82-95
    • /
    • 2024
  • As automated essay scoring (AES) has progressed from handcrafted techniques to deep learning, holistic scoring capabilities have merged. However, specific trait assessment remains a challenge because of the limited depth of earlier methods in modeling dual assessments for holistic and multi-trait tasks. To overcome this challenge, we explore providing comprehensive feedback while modeling the interconnections between holistic and trait representations. We introduce the DualBERT-Trans-CNN model, which combines transformer-based representations with a novel dual-scale bidirectional encoder representations from transformers (BERT) encoding approach at the document-level. By explicitly leveraging multi-trait representations in a multi-task learning (MTL) framework, our DualBERT-Trans-CNN emphasizes the interrelation between holistic and trait-based score predictions, aiming for improved accuracy. For validation, we conducted extensive tests on the ASAP++ and TOEFL11 datasets. Against models of the same MTL setting, ours showed a 2.0% increase in its holistic score. Additionally, compared with single-task learning (STL) models, ours demonstrated a 3.6% enhancement in average multi-trait performance on the ASAP++ dataset.

An Analysis on Reliabilities of Scoring Methods and Rubric Ratings Number for Performance Assessments of Middle School Students' Science Investigation Activities (중학생 과학탐구활동 수행평가 시 채점 방식 및 척도의 수에 따른 신뢰도 분석)

  • Kim, Hyung-Jun;Yoo, June-Hee
    • Journal of The Korean Association For Science Education
    • /
    • v.30 no.2
    • /
    • pp.275-290
    • /
    • 2010
  • In this study, reliabilities of holistic scoring method and analytic scoring method were analyzed in performance assessments of middle school students' science investigation activity. Reliabilities of 2, 3, and 4~7-level rubric ratings for analytic scoring methods were compared to figure out optimized numbers of rubric ratings. Two trained raters rated four activity sheets of 60 students by two rating methods and three kinds of rubric ratings. Internal consistency reliabilities of holistic scoring methods were higher than those of analytic scoring methods, while intrarater reliabilities of analytic scoring were higher than those of holistic scoring methods. Internal consistency reliabilities and intra-rater reliabilities of 3-level rubric rating showed similar patterns of 4~7-level rubric ratings. But students' discriminations, item difficulties and item-response curves showed that the 3-level rubric ratings was reliable. These results suggest that holistic scoring method could be adapted to increase internal consistency reliabilities with improvement in intra-rater reliabilities by rater's conferences. Also, the 3-level rubric rating would be enough for good reliability in case of adapting analytic scoring methods.

A Study on Development of Balanced Performance Assessment Tasks for Primary School Mathematics -Focused on 1, 2 Stage in the Primary School- (균형 있는 초등수학과 수행평가 과제 개발에 대한 연구 - 1, 2단계를 중심으로 -)

  • 정영옥
    • School Mathematics
    • /
    • v.3 no.2
    • /
    • pp.325-354
    • /
    • 2001
  • The study aims to develop balanced performance assessment tasks for primary school mathematics which can be implemented in the primary school easily. In order to these purposes, I suggest the types of performance assessment tasks and the framework of assessment standards for the balanced performance assessment with describing the procedures of developing tasks and rubrics. The types of task are journal writing, problem posing, constructed task, and descriptive task. In the framework of assessment standards, I suggest holistic scoring which are classified as four levels according to the degree of excellence which students perform totally concerning about the criterion of implication, reasoning, accuracy, and communication. Also I analyse the responses of children to the task “make a beautiful pattern” and suggest its assessment rubric and anchor papers for each level for illustrating the process of developing a rubric in holistic scoring. In order to reflect the viewpoints of children and their Parents concerning about the tasks, the responses in self assessment and parent assessment are analysed. Finally, methods of implementing the assessment tasks and considerations are discussed.

  • PDF

Alternative Assessment in Mathematics Education (대안적인 평가를 통한 수학교육)

  • 최승현
    • Journal of Educational Research in Mathematics
    • /
    • v.8 no.1
    • /
    • pp.217-235
    • /
    • 1998
  • The purpose of this study is to define the altenative assessment and to suggest the method of scoring system. Alternative assessment includes any type of assessment in which student create reponses to a question rather than choosing a responses form given list( as for multiple choice, true/false, or matching). Alternative assessment can includes short answer questions, essay, performances, oral presentation, demonstrations, exhibitions, portfolios, and etc. To evaluate the each type of assessment, we can apply the method of holistic scoring and analytic scoring system. Also we have to concern the type of scoring mechanism directly relate to what we want to assess, our purpose for assessment fitting into the educational enterprise. Before applying the alternative assessment in our classroom, we need to step back and reconsider all our design features and teachers' responsibility.

  • PDF

An Application of Generalizability Theory to Self-introduction Letter and Teacher's Recommendation Letter Used in Identification of Mathematical Gifted Students by Observations and Nominations (관찰.추천에 의한 수학영재 선발 시 사용되는 자기소개서와 교사추천서 평가에 대한 일반화가능도 이론의 활용)

  • Kim, Sung-Chan;Kim, Sung-Yeun;Han, Ki-Soon
    • Communications of Mathematical Education
    • /
    • v.26 no.3
    • /
    • pp.251-271
    • /
    • 2012
  • The purpose of this study is: 1) to determine error sources and the effects of each error source, 2) to investigate optimal measuring conditions from holistic and analytic scoring methods, and 3) to compare the value of reliability between Cronbach's alpha and the generalizability coefficient in self-introduction letter and teacher's recommendation letter based on the generalizability theory in identification of mathematical gifted students by observations and nominations. Data of this study were collected from the science education institute for the gifted attached to the university located within in a capital city for the 2011 academic year. Scores form two raters using holistic and analytic scoring methods in both assessment types were used. The results of this study were as follows. First, as to both assessment types, error sources for people were relatively large regardless of scoring methods. However, error sources for raters in holistic scoring methods had a more significant impact than those of analytic scoring methods. Second, to set optimal measuring conditions in the self-introduction letter and teacher's recommendation letter, if we fixed the number of raters into 2 based on holistic scoring methods, at least 5 and 10 content domains were needed, respectively. In addition, the number of items in teacher's recommendation letter should be more than 3 when we fixed the number of content domains into 4, and the number of items in self-introduction letter should be more than 8 when we fixed the number of content domains into 6 using analytic scoring methods. Third, Cronbach's alpha having only a single source of errors was higher than the generalizability coefficient regardless of assessment types and scoring methods. Hence we recommend that generalizability coefficient based on various error sources such as raters, content domains, and items should be considered to keep a satisfactory level of reliability in both assessment types.

Developing Scoring Rubric and the Reliability of Elementary Science Portfolio Assessment (초등 과학과 포트폴리오의 채점기준 개발과 신뢰도 검증)

  • Kim, Chan-Jong;Choi, Mi-Aee
    • Journal of The Korean Association For Science Education
    • /
    • v.22 no.1
    • /
    • pp.176-189
    • /
    • 2002
  • The purpose of the study is to develop major types of scoring rubrics of portfolio system, and estimate the reliability of the rubrics developed. The portfolio system was developed by Science Education Laboratory, Chongju National University of Education in summer, 2000. The portfolio is based on the Unit 2, The Layer and Fossil, and Unit 4, Heat and Change of Objects at fourth-grade level. Four types of scoring rubrics, holistic-general, holistic-specific, analytical-general, and analytical-specific, were developed. Students' portfolios were scored and inter-rater and intra-rater reliability were calculated. To estimate inter-rater reliability, 3 elementary teachers per each rubric(total 12) scored 12 students' portfolios. Teachers who used analytical-specific rubric scored only six portfolios because it took much more time than other rubrics. To estimate intra-rater reliability, second scoring was administered by two raters per rubric in two and half month. The results show that holistic-general rubric has high inter-rater and moderate intra-rater reliability. Holistic-specific rubric shows moderate inter- and intra-rater reliability. Analytical-general rubric has high inter-rater and moderate intra-rater reliability. Analytical-specific rubric shows high inter- and intra-rater reliability. The raters feel that general rubrics seems to be practical but not clear. Specific rubrics provide more clear guidelines for scoring but require more time and effort to develop the rubrics. Analytical-specific rubric requires more than two times of time to score each portfolio and is proved to be highly reliable but less practical.

A Chaos of Understanding on Performance Assessment in Mathematics Education (수학과 수행평가에 관한 이해의 혼돈 -최근 국내 논문 분석을 중심으로-)

  • 황혜정
    • The Mathematical Education
    • /
    • v.42 no.2
    • /
    • pp.159-176
    • /
    • 2003
  • From the mid-1990s in Korea, performance assessment has been continuously emphasized in school mathematics and thus many researchers and teachers have been steadily studying this topic. But the concepts relevant to performance assessment and its purposes are very confusing because the mathematics educators' different views and voices are vary. As a result most mathematics teachers experience trouble in executing performance assessment properly and effectively in their math class. This unability for proper execution of performance assessment was once again revealed in this study which dealt with 15 articles on performance assessment. These 15 articles includes almost every article written on the topic of performance assessment that have been published in 4 domestic journals since December 1997. By examining this inability, it is required that its concepts and purposes should be organized with a common view and newly defined in the near future. Therefore, to successfully accomplish this, this paper outlines the basic problems on the understanding of performance assessment as follows: ㆍWhat is the relationship between performance assessment and alternative assessment\ulcorner ㆍWhat is the proper types(methods) of performance assessment\ulcorner ㆍIs the subject test a type of performance assessment\ulcorner ㆍWhat is the difference between subject test and essay test\ulcorner ㆍWhat is the relationship between performance assessment and performance tasks\ulcorner ㆍWhat is the relationship between performance tests and project method\ulcorner ㆍWhat is a project method\ulcorner ㆍIs it assessment standard or scoring standard to score a test result\ulcorner ㆍWhat is the difference between analytic scoring method and holistic scoring method?

  • PDF

Measuring plagiarism in the second language essay writing context (영작문 상황에서의 표절 측정의 신뢰성 연구)

  • Lee, Ho
    • English Language & Literature Teaching
    • /
    • v.12 no.1
    • /
    • pp.221-238
    • /
    • 2006
  • This study investigates the reliability of plagiarism measurement in the ESL essay writing context. The current study aims to address the answers to the following research questions: 1) How does plagiarism measurement affect test reliability in a psychometric view? and 2) how do raters conceive the plagiarism in their analytic scoring? This study uses the mixed-methodology that crosses quantitative-qualitative techniques. Thirty eight international students took an ESL placement writing test offered by the University of Illinois. Two native expert raters rated students' essays in terms of 5 analytic features (organization, content, language use, source use, plagiarism) and made a holistic score using a scoring benchmark. For research question 1, the current study, using G-theory and Multi-facet Rasch model, found that plagiarism measurement threatened test reliability. For research question 2, two native raters and one non-native rater in their email correspondences responded that plagiarism was not a valid analytic area to be measured in a large-scale writing test. They viewed the plagiarism as a difficult measurement are. In conclusion, this study proposes that a systematic training program for avoiding plagiarism should be given to students. In addition, this study suggested that plagiarism is measured reliably in the small-scale classroom test.

  • PDF

A Study on the Relation among Mathematical - Spatial - Verbal Abilities and Gender Differences of Engineering Students (공과대학생들의 수리 - 공간 - 언어 능력 사이의 관계 및 성별 차이에 관한 연구)

  • Kim, Yeon Mi
    • Journal of Engineering Education Research
    • /
    • v.18 no.4
    • /
    • pp.34-44
    • /
    • 2015
  • Mathematical, spatial, and verbal abilities are important for future engineers to succeed in the STEM disciplines. The purpose of the study is to assess engineering students' spatial abilities and analyse the relationship with mathematical achievement, verbal achievement, and gender. On the mental rotation tests, 65% of male students demonstrated a substantial level of spatial abilities. But only 30% of female students exhibited spatial skills at the same level as their male colleagues. The correlations between mathematical - spatial - verbal abilities are found to be negligible. When spatial visualization ability was plotted according to the mathematical achievement level, there was no difference in the mean spatial abilities score. But when mathematical achievement score was plotted according to the spatial abilities, there was a noticeable difference. Regression analysis confirmed that female students' mathematical achievement increased as spatial abilities improved. This phenomenon was not observed for male students. It's because male students' spatial ability already contributed to their mathematics achievement. So spatial ability can be regarded as one factor for the gender differences in mathematics achievement. The gender gap on spatial abilities and math achievement is large among high achieving students. For example, there was a 4.3 to 1 male - female ratio and 3.4 to 1 male - female ratio among students scoring 99th percentile in spatial visualization test and scholastic aptitude test-math.

Development and Application of Nursing Service Quality Indicators in Nursing Homes (노인요양시설의 간호서비스 질 평가 지표 개발 및 적용)

  • Chung, Ja-Ne
    • Journal of Korean Academy of Nursing
    • /
    • v.37 no.3
    • /
    • pp.401-413
    • /
    • 2007
  • Purpose: This study was designed to develop Nursing Service Quality Indicators(NSQIs) in nursing homes that would lead to an appropriate evaluation and improvement of nursing service quality. Methods: The preliminary NSQIs were developed through literature reviews and analysis of existing quality indicators. A content validity testing was done twice by using a panel of experts who were from academia and the clinical areas. The final NSQIs were confirmed and applied in three nursing homes to test feasibility. Results: The preliminary NSQIs had 4 domains and 31 indicators. Two content validity testings were performed. The indicators scoring over.80 CVI for each testing were selected and modified by experts' opinions. The final NSQIs consisted of 7 domains and 33 indicators. They were applied in three nursing homes and it was revealed that all the indicators were applicable. Conclusion: In this study, it is shown that this new 'Nursing Service Quality Indicators in Nursing Homes' is suitable for a holistic evaluation of nursing service quality of elderly patients in nursing homes. This NSQIs will be able to provide a basis for establishing nursing care standards and improving the nursing care quality in nursing homes.