Rater Effects in Assessing Oral Production of L2 Chinese Pragmatic Performance: The Influence of Teaching Experience and Teaching Context
Feng,Yali
Citations
Abstract
Previous research has shown variations in rater behavior and cognitive processes in relation to raters’ background characteristics (e.g., rater native language status, accent familiarity, and rater training) in L2 performance assessment. However, few studies have examined rater effects in L2 pragmatic performance assessment, particularly in the context of Chinese as a second language. This study examined rater effects in the assessment of L2 Chinese oral pragmatic performance, focusing on two rater background characteristics: TCSOL teaching experience and teaching context (China vs. the United States). This study used a mixed-methods design to examine raters’ scoring behavior and reported aspects of rater cognition. It included 72 native Chinese-speaking raters: 24 experienced TCSOL teachers, 24 novice TCSOL teachers, and 24 non-TCSOL teachers. Each rater group included equal numbers of China-based and U.S.-based raters. For the rating task, all raters scored audio-recorded oral responses to speech act and pragmatic routine items produced by American college-level learners of Chinese. Each rater provided a holistic score and a written rationale for each rating. Many-facet Rasch measurement (MFRM) was used to examine rater severity, consistency, and bias. The written rationales were analyzed to compare what raters attended to and how they explained their scoring decisions across groups. Teaching experience and teaching context were both significantly related to raters’ scoring behavior. Non-TCSOL teachers scored significantly less consistently than experienced TCSOL teachers, while novice TCSOL teachers were significantly more severe than experienced TCSOL teachers. China-based teachers were also significantly more severe than U.S.-based teachers, but no interaction was found between either rater background variable and item type. The qualitative analysis also showed differences related to teaching experience. Compared with the other groups, novice TCSOL teachers commented more on linguistic accuracy, used more interpretation monitoring, and made fewer holistic quality-of-language comments. No significant differences were found between the two teaching-context groups in any coding category. Across groups, raters tended to comment more on appropriateness when assessing speech acts and more on communicative function when assessing pragmatic routines. Overall, the findings show that rater background was related to both scoring behavior and raters’ reported attention.
