Rater Bias in Assessing Iranian EFL Learners’ Writing Performance

Authors
Abstract
Evidence suggests that variability in the ratings of students’ essays results not only from their differences in their writing ability, but also from certain extraneous sources. In other words, the outcome of the rating of essays can be biased by factors which relate to the rater, task, and situation, or an interaction of all or any of these factors which make the inferences and decisions made about students’ writing ability undependable. The purpose of this study, therefore, was to examine the issue of variability in rater judgments as a source of measurement error this was done in relation to EFL learners’ essay writing assessment. Thirty two Iranian sophomore students majoring in English language participated in this study. The learners’ narrative essays were rated by six different raters and the results were analyzed using many-facet Rasch measurement as implemented in the computer program FACETS. The findings suggest that there are significant differences among raters concerning their harshness as well as several cases of bias due to the rater-examinee interaction. This study provides a valuable understanding of how effective and reliable rating can be realized, and how the fairness and accuracy of subjective performance can be assessed. 
Keywords

Article Title Persian

تورش مصحح در ارزیابی نگارش فر اگیران زبان انگلیسی در ایران

Authors Persian

مهناز سعیدی
ماندانا یوسفی
پوریا بقائی
Abstract Persian

شواهد بیانگر این می باشد که تغییر پذیری در ارزیابی نگارش دانشجویان فقط نتیجه تفاوت مهارت نوشتاری آنها نیست بلکه عوامل بیرونی خاصی در این امر دخیلند و عوامل مرتبط با مصحح، فعالییت، موقعییت یا تعامل هر یک از اینها میتوانند در تصمیم گیریها و استنباطها در مورد توانایی نگارش فراگیران تاثیر بگذارند. هدف این تحقیق این است که مساله تغییرپذیری را در داوری مصحح به عنوان منبع خطای سنجش در ارزیابی نگارش فراگیران زبان انگلیسی به عنوان زبان خارجه مورد بررسی قرار دهد. به این منظور مقالات 32 دانشجوی ایرانی زبان انگلیسی بوسیله شش مصحح ارزشیابی و نتایج با استفاده از مدل رش و برنامه فست مورد تحلیل قرار گرفتند. یافته ها نشان دادند که تفاوتهای معنی داری در میزان سختگیری مصحح ها وجود دارد و تعامل مصحح با فراگیر باعث ایجاد تورش در ارزشیابی میشود. از آنجایی که ارزیابی مهارت نوشتن امری ذهنی میباشد و بر اساس داوری مصحح می تواند متغیر باشد، بررسی و شناخت عواملی که باعث عینی تر شدن این ارزیابی میشود امری ضروری به شمارمی رود. ذهنی بودن روند ارزیابی مهارت نوشتن تهدیدی برای روایی آزمون است و سبب می شود که نمره فراگیر نشان دهنده مهارت واقعی وی نباشد. این تحقیق نشان می دهد چطور ارزیابی می تواند کارامد و سودمند باشد و چگونه می توان انصاف و صحت عملکرد ذهنی را سنجید.

Keywords Persian

تورش مصحح
توانایی نگارش
سنجش چند وجهی رش
پایایی بین مصحح