This study examines the Test of Indonesian Proficiency (
Research article
Evaluating the logic of a policy-driven national language test in Indonesia: A critical discursive investigation of the Test of Indonesian Proficiency (UKBI)
Abstract
Select search scope: search across all journals or within the current journal
This study examines the Test of Indonesian Proficiency (
This study explores the integration of generative artificial intelligence (GenAI) with human experts to improve the quality of distractors in multiple-choice questions (MCQs) for second language (L2) listening tests. A psychometric analysis of responses from 2267 EFL Chinese undergraduates, using the two-parameter logistic nested logit model (2PLNLM), identified problematic items and distractors. Guided by established distractor design principles, GenAI was applied iteratively to refine these distractors, and GenAI was iteratively used to revise these distractors, with human experts providing ongoing feedback throughout the process. The revised versions were then evaluated by expert judgment and NLP-based cosine similarity analysis. The results indicate that GenAI effectively enhanced distractor quality by maintaining content and structural alignment and ensuring semantic independence. However, it struggled to fully capture listening miscomprehension patterns and contextualized language use. These preliminary findings suggest that GenAI revisions, guided by principle-based prompts and supervised by humans, tend to effectively improve the quality of distractors. This study offers practical insights into the potential and limitations of GenAI in improving L2 listening tests.
This three-level meta-analysis investigated the human–machine correlation of automated speech evaluation (ASE) systems. Sixty-seven studies representing 392 effect sizes were included. The results indicated a positive overall correlation (
Assessment rubrics have increasingly been developed and deployed to evaluate the quality of language interpreting, yet understanding of rubric-based interpreting assessment remains limited. This systematic review aims to: (a) catalog rubrics, (b) examine rubric design features, (c) understand rubric use, and (d) evaluate rubric utility. A rigorous review process, involving database searching, citation tracking, and targeted review of core literature, identified 80 unique rubrics comprising a total of 265 (sub-)scales. A comprehensive analysis revealed that: (a) among 11 potential sources informing rubric development, test-external sources (e.g., literature review) were primarily used, whereas test-internal sources (e.g., performance samples) were much less consulted; (b) assessments primarily used analytic and task-type rubrics with an average of five performance levels and four scoring criteria. Rubric descriptors generally incorporated observable indicators of interpreting quality, employing both descriptive and evaluative rubric language; (c) rubrics were used by three main types of raters—interpreting practitioners, trainers, and students—to assess primarily spoken-language interpreting; and (d) rubric-based scores demonstrated moderate-to-high reliability and validity overall, though meta-analysis identified three significant moderators, including correlation coefficient type, assessment criterion, and rubric length. These findings are expected to provide guidance for assessment practice and research in interpreting.
Over one million international students are enrolled in US higher education institutions. To be admitted into one of these degree-granting institutions, international students must demonstrate sufficient English language proficiency (ELP)–often with a test score. Previous research on ELP tests has shed light on how scores correlate with academic success and how institutions make decisions about minimum required scores, but the extent to which US institutions vary in their use of ELP tests in international admissions is not well understood. In this study, we examined ELP test policies across all US research-intensive, doctoral-granting institutions (Carnegie R1) for general undergraduate admissions and graduate admissions. Results show that 32 different ELP tests were being used for admissions. The most accepted ELP tests were TOEFL iBT, IELTS Academic, Duolingo English Test, and Pearson Test of English Academic (in that order). A wide range of unconditional and conditional cut scores was found across institutions. Subscore requirements were not common. Institutions requiring higher ELP cut scores (specifically TOEFL iBT) tended to be more prestigious (higher ranked) and selective (admitting smaller proportions of undergraduate applicants).

