• 제목/요약/키워드: holistic scoring

검색결과 14건 처리시간 0.022초

Dual-scale BERT using multi-trait representations for holistic and trait-specific essay grading

  • Minsoo Cho;Jin-Xia Huang;Oh-Woog Kwon
    • ETRI Journal
    • /
    • 제46권1호
    • /
    • pp.82-95
    • /
    • 2024
  • As automated essay scoring (AES) has progressed from handcrafted techniques to deep learning, holistic scoring capabilities have merged. However, specific trait assessment remains a challenge because of the limited depth of earlier methods in modeling dual assessments for holistic and multi-trait tasks. To overcome this challenge, we explore providing comprehensive feedback while modeling the interconnections between holistic and trait representations. We introduce the DualBERT-Trans-CNN model, which combines transformer-based representations with a novel dual-scale bidirectional encoder representations from transformers (BERT) encoding approach at the document-level. By explicitly leveraging multi-trait representations in a multi-task learning (MTL) framework, our DualBERT-Trans-CNN emphasizes the interrelation between holistic and trait-based score predictions, aiming for improved accuracy. For validation, we conducted extensive tests on the ASAP++ and TOEFL11 datasets. Against models of the same MTL setting, ours showed a 2.0% increase in its holistic score. Additionally, compared with single-task learning (STL) models, ours demonstrated a 3.6% enhancement in average multi-trait performance on the ASAP++ dataset.

중학생 과학탐구활동 수행평가 시 채점 방식 및 척도의 수에 따른 신뢰도 분석 (An Analysis on Reliabilities of Scoring Methods and Rubric Ratings Number for Performance Assessments of Middle School Students' Science Investigation Activities)

  • 김형준;유준희
    • 한국과학교육학회지
    • /
    • 제30권2호
    • /
    • pp.275-290
    • /
    • 2010
  • 중학생의 과학탐구활동 수행평가 시 총체적 채점과 분석적 채점의 신뢰도를 비교 분석하였으며, 분석적 채점을 하는 경우에는 신뢰도 확보를 위하여 채점척도의 수준을 어느 정도로 분석적으로 해야 하는지를 조사하였다. 중학생들이 작성한 4개의 과학탐구과제에 대한 활동지를 두 명의 채점자가 총체적 채점 방식, 분석적 채점 방식, 분석적 채점 중 채점척도를 2, 3, 4~7수준으로 다르게 하여 채점하였다. 총체적 채점 방식은 과제 간 내적 일치도가 높게 나타났으며, 분석적 채점 방식은 채점자간 신뢰도가 높게 나타났다. 또한 채점척도 3수준의 경우는 4~7수준의 경우와 활동간 내적 일치도와 채점자간의 신뢰도가 유사하게 나타났으나, 능력추정치별 학생의 분포, 문항곤란도 및 문항특성곡선의 경우 채점척도 3수준의 경우가 적절한 것으로 나타났다. 이러한 연구 결과는 과학탐구활동 수행평가 시 총체적 채점 방식을 선택하는 경우는 과제 간 내적일치도를 높일 수 있으며 분석적 채점 방식에 비해 낮게 나타나는 채점자 간 일치도를 높이기 위한 채점자간 협의등 방안이 필요하다는 것을 시사한다. 또한 분석적 채점 방식을 선택하는 경우는 채점척도 3수준으로 충분히 신뢰도를 확보할 수 있다는 점을 시사한다.

균형 있는 초등수학과 수행평가 과제 개발에 대한 연구 - 1, 2단계를 중심으로 - (A Study on Development of Balanced Performance Assessment Tasks for Primary School Mathematics -Focused on 1, 2 Stage in the Primary School-)

  • 정영옥
    • 대한수학교육학회지:학교수학
    • /
    • 제3권2호
    • /
    • pp.325-354
    • /
    • 2001
  • The study aims to develop balanced performance assessment tasks for primary school mathematics which can be implemented in the primary school easily. In order to these purposes, I suggest the types of performance assessment tasks and the framework of assessment standards for the balanced performance assessment with describing the procedures of developing tasks and rubrics. The types of task are journal writing, problem posing, constructed task, and descriptive task. In the framework of assessment standards, I suggest holistic scoring which are classified as four levels according to the degree of excellence which students perform totally concerning about the criterion of implication, reasoning, accuracy, and communication. Also I analyse the responses of children to the task “make a beautiful pattern” and suggest its assessment rubric and anchor papers for each level for illustrating the process of developing a rubric in holistic scoring. In order to reflect the viewpoints of children and their Parents concerning about the tasks, the responses in self assessment and parent assessment are analysed. Finally, methods of implementing the assessment tasks and considerations are discussed.

  • PDF

대안적인 평가를 통한 수학교육 (Alternative Assessment in Mathematics Education)

  • 최승현
    • 대한수학교육학회지:수학교육학연구
    • /
    • 제8권1호
    • /
    • pp.217-235
    • /
    • 1998
  • The purpose of this study is to define the altenative assessment and to suggest the method of scoring system. Alternative assessment includes any type of assessment in which student create reponses to a question rather than choosing a responses form given list( as for multiple choice, true/false, or matching). Alternative assessment can includes short answer questions, essay, performances, oral presentation, demonstrations, exhibitions, portfolios, and etc. To evaluate the each type of assessment, we can apply the method of holistic scoring and analytic scoring system. Also we have to concern the type of scoring mechanism directly relate to what we want to assess, our purpose for assessment fitting into the educational enterprise. Before applying the alternative assessment in our classroom, we need to step back and reconsider all our design features and teachers' responsibility.

  • PDF

관찰.추천에 의한 수학영재 선발 시 사용되는 자기소개서와 교사추천서 평가에 대한 일반화가능도 이론의 활용 (An Application of Generalizability Theory to Self-introduction Letter and Teacher's Recommendation Letter Used in Identification of Mathematical Gifted Students by Observations and Nominations)

  • 김성찬;김성연;한기순
    • 한국수학교육학회지시리즈E:수학교육논문집
    • /
    • 제26권3호
    • /
    • pp.251-271
    • /
    • 2012
  • 이 연구는 관찰 추천 수학영재선발 시 사용되는 자기소개서와 교사추천서 평가에서 발생하는 오차요인들의 상대적인 영향력을 살펴보고, 교사추천서와 자기소개서를 총체적 채점과 분석적 채점으로 실시했을 때 채점 방법에 따른 일반화가능도계수의 최적화 측정 조건을 탐색하고, 이를 전통적인 신뢰도 추정방법과 비교하였다. 2011학년도 수도권에 소재하고 있는 대학부설 과학영재교육원에서 관찰-추천 영재 선발에 지원한 90명의 자기소개서와 교사추천서에 대해 총체적 채점과 분석적 채점으로 2명의 교사가 각각 점수를 부여하였다. 연구결과는 다음과 같다. 첫째, 교사추천서와 자기소개서의 평가에 있어 채점방법에 따른 공통점은 피험자 관련 분산이 크게 나타났으며, 차이점은 총체적 채점이 분석적 채점보다 채점자의 영향이 더 큰 것으로 나타났다. 둘째, 적정수준의 일반화가능도계수를 얻기 위해서 채점자를 2명으로 고정하는 경우 교사추천서와 자기소개서에서 총체적 채점은 각각 내용영역이 5개, 10개 이상이 요구되어졌으며, 분석적 채점은 각각 내용영역을 4개로 고정한 경우 문항이 3개 이상, 내용영역을 6개로 고정한 경우 문항이 8개 이상이 요구되어졌다. 셋째, 교사추천서와 자기소개서 모두 채점 방법과 상관없이 문항만을 오차요인으로 보는 Cronbach ${\alpha}$가 신뢰도를 과대 추정하는 것으로 나타났다. 따라서 적정수준의 신뢰도를 확보하기 위해서는 채점자, 내용영역, 문항수와 같이 다양한 오차요인을 반영하는 일반화가능도 계수를 고려하는 것이 바람직할 것이다.

초등 과학과 포트폴리오의 채점기준 개발과 신뢰도 검증 (Developing Scoring Rubric and the Reliability of Elementary Science Portfolio Assessment)

  • 김찬종;최미애
    • 한국과학교육학회지
    • /
    • 제22권1호
    • /
    • pp.176-189
    • /
    • 2002
  • 본 연구의 목적은 초등학교 과학과 포트폴리오를 채점할 수 있는 다양한 채점기준을 개발하고, 개발된 각 채점기준의 신뢰도를 검증해 보고자 하는 것이다. 채점기준을 개발하기 위한 포트폴리오는 4학년 2학기 '단원 2. 지층과 화석', '단원 4. 열과 물체의 변화' 를 중심으로 청주교대 과학교육 연구실에서 2000년 여름에 개발한 체제를 같은 해 가을, 경기도 중도시의 한 초등학교 4학년 한 학급에 적용하여 얻은 것이다. 총괄-일반, 총괄-특수, 분석-일반, 분석-특수의 4가지 채점기준을 개발하고, 각 채점기준에 근거하여 학생들이 작성한 포트폴리오 증거물을 채점하여 각 채점 기준별 채점자간 신뢰도와, 채점자내 신뢰도를 구하였다. 1차 채점에서는 총 12명의 채점자들이 각 채점기준별로 3명씩 그룹을 나누어 그룹당 12권의 포트폴리오 증거물을 채점하였다. 단, 분석-특수 채점기준의 경우 6권의 포트폴리오 증거물만을 채점하였다. 채점자내 신뢰도를 알아보기 위해 실시한 채점시기별 신뢰도에서는 l차 채점에 참가한 채점자 중 각 채점기준별로 2명씩 총 8명이 2차 채점에 참가하여 l차 채점과 동일한 방식으로 채점을 실시하였다. 채점결과를 SPSS 통계 프로그램에 입력하여 상관계수를 구한 결과, 총괄-일반 채점기준은 채점자간 신뢰도가 높고 채점자내 신뢰도가 있는 것으로 나타났고 총괄-특수 채점기준은 채점자간 신뢰도와 채점자내 신뢰도가 있는 것으로 나타났다. 분석-일반 채점기준은 채정자간 신뢰도가 높고 채점자내 신뢰도는 있는 것으로 나타났으며, 분석-특수 채점기준은 채점자간 신뢰도와 채점자내 신뢰도가 모두 높은 것으로 나타났다. 일반적인 채점기준들(총괄-일반, 분석-일반)의 경우, 하나의 채점 기준으로 모든 포트폴리오 목표를 채점할 수 있으므로 매우 경제적이고 실용적이나, 채점자들은 채점시 모호함을 느낀다고 하였다. 반면에, 특수적인 채점기준들(총괄-특수, 분석-특수)의 경우, 채점은 더 명확하게 할 수 있으나, 목표별로 채점기준을 개발해야 하므로 많은 시간과 노력이 필요하게 된다. 채점기준의 실용도 측면에서는 분석-특수 채점기준이 다른 기준보다 2배 이상의 시간이 결려 실용도는 낮은 것으로 나타났다.

수학과 수행평가에 관한 이해의 혼돈 -최근 국내 논문 분석을 중심으로- (A Chaos of Understanding on Performance Assessment in Mathematics Education)

  • 황혜정
    • 한국수학교육학회지시리즈A:수학교육
    • /
    • 제42권2호
    • /
    • pp.159-176
    • /
    • 2003
  • From the mid-1990s in Korea, performance assessment has been continuously emphasized in school mathematics and thus many researchers and teachers have been steadily studying this topic. But the concepts relevant to performance assessment and its purposes are very confusing because the mathematics educators' different views and voices are vary. As a result most mathematics teachers experience trouble in executing performance assessment properly and effectively in their math class. This unability for proper execution of performance assessment was once again revealed in this study which dealt with 15 articles on performance assessment. These 15 articles includes almost every article written on the topic of performance assessment that have been published in 4 domestic journals since December 1997. By examining this inability, it is required that its concepts and purposes should be organized with a common view and newly defined in the near future. Therefore, to successfully accomplish this, this paper outlines the basic problems on the understanding of performance assessment as follows: ㆍWhat is the relationship between performance assessment and alternative assessment\ulcorner ㆍWhat is the proper types(methods) of performance assessment\ulcorner ㆍIs the subject test a type of performance assessment\ulcorner ㆍWhat is the difference between subject test and essay test\ulcorner ㆍWhat is the relationship between performance assessment and performance tasks\ulcorner ㆍWhat is the relationship between performance tests and project method\ulcorner ㆍWhat is a project method\ulcorner ㆍIs it assessment standard or scoring standard to score a test result\ulcorner ㆍWhat is the difference between analytic scoring method and holistic scoring method?

  • PDF

영작문 상황에서의 표절 측정의 신뢰성 연구 (Measuring plagiarism in the second language essay writing context)

  • 이호
    • 영어어문교육
    • /
    • 제12권1호
    • /
    • pp.221-238
    • /
    • 2006
  • This study investigates the reliability of plagiarism measurement in the ESL essay writing context. The current study aims to address the answers to the following research questions: 1) How does plagiarism measurement affect test reliability in a psychometric view? and 2) how do raters conceive the plagiarism in their analytic scoring? This study uses the mixed-methodology that crosses quantitative-qualitative techniques. Thirty eight international students took an ESL placement writing test offered by the University of Illinois. Two native expert raters rated students' essays in terms of 5 analytic features (organization, content, language use, source use, plagiarism) and made a holistic score using a scoring benchmark. For research question 1, the current study, using G-theory and Multi-facet Rasch model, found that plagiarism measurement threatened test reliability. For research question 2, two native raters and one non-native rater in their email correspondences responded that plagiarism was not a valid analytic area to be measured in a large-scale writing test. They viewed the plagiarism as a difficult measurement are. In conclusion, this study proposes that a systematic training program for avoiding plagiarism should be given to students. In addition, this study suggested that plagiarism is measured reliably in the small-scale classroom test.

  • PDF

공과대학생들의 수리 - 공간 - 언어 능력 사이의 관계 및 성별 차이에 관한 연구 (A Study on the Relation among Mathematical - Spatial - Verbal Abilities and Gender Differences of Engineering Students)

  • 김연미
    • 공학교육연구
    • /
    • 제18권4호
    • /
    • pp.34-44
    • /
    • 2015
  • Mathematical, spatial, and verbal abilities are important for future engineers to succeed in the STEM disciplines. The purpose of the study is to assess engineering students' spatial abilities and analyse the relationship with mathematical achievement, verbal achievement, and gender. On the mental rotation tests, 65% of male students demonstrated a substantial level of spatial abilities. But only 30% of female students exhibited spatial skills at the same level as their male colleagues. The correlations between mathematical - spatial - verbal abilities are found to be negligible. When spatial visualization ability was plotted according to the mathematical achievement level, there was no difference in the mean spatial abilities score. But when mathematical achievement score was plotted according to the spatial abilities, there was a noticeable difference. Regression analysis confirmed that female students' mathematical achievement increased as spatial abilities improved. This phenomenon was not observed for male students. It's because male students' spatial ability already contributed to their mathematics achievement. So spatial ability can be regarded as one factor for the gender differences in mathematics achievement. The gender gap on spatial abilities and math achievement is large among high achieving students. For example, there was a 4.3 to 1 male - female ratio and 3.4 to 1 male - female ratio among students scoring 99th percentile in spatial visualization test and scholastic aptitude test-math.

노인요양시설의 간호서비스 질 평가 지표 개발 및 적용 (Development and Application of Nursing Service Quality Indicators in Nursing Homes)

  • 정제인
    • 대한간호학회지
    • /
    • 제37권3호
    • /
    • pp.401-413
    • /
    • 2007
  • Purpose: This study was designed to develop Nursing Service Quality Indicators(NSQIs) in nursing homes that would lead to an appropriate evaluation and improvement of nursing service quality. Methods: The preliminary NSQIs were developed through literature reviews and analysis of existing quality indicators. A content validity testing was done twice by using a panel of experts who were from academia and the clinical areas. The final NSQIs were confirmed and applied in three nursing homes to test feasibility. Results: The preliminary NSQIs had 4 domains and 31 indicators. Two content validity testings were performed. The indicators scoring over.80 CVI for each testing were selected and modified by experts' opinions. The final NSQIs consisted of 7 domains and 33 indicators. They were applied in three nursing homes and it was revealed that all the indicators were applicable. Conclusion: In this study, it is shown that this new 'Nursing Service Quality Indicators in Nursing Homes' is suitable for a holistic evaluation of nursing service quality of elderly patients in nursing homes. This NSQIs will be able to provide a basis for establishing nursing care standards and improving the nursing care quality in nursing homes.