Skip to main content

Cookie settings

We use cookies to ensure the basic functionalities of the website and to enhance your online experience. You can configure and accept the use of the cookies, and modify your consent options, at any time.

Essential

Preferences

Analytics and statistics

Marketing

Measuring the AI Language Gap in International Higher Education

Avatar: Official proposal Official proposal

Team name: LinguaGap

  • Tool used: Gemini, ChatGPT, Deepl

External feedback & contributions

NameRoleType of contributionElla FaouziFellow Challenge participantWritten comment — raised the question of including European open-source multilingual models (Mistral, Aya) in the benchmarkNdeye Sophie SeckFellow Challenge participantWritten comment — raised the link between AI underperformance, cognitive fatigue, and student autonomous reflectionNipun Ranchhod NavadiaFellow Challenge participantWritten comment — challenged the methodology on bias isolation, definition of "performance," and the intervention pathwayEline VIVET MALADRYFellow Challenge participantWritten comment — pushed on the definition of equity and performance, and on the risk of treating languages as homogeneous blocks

Initial contribution: Please provide the link (URL) to your initial contribution on the platform: [Insert link]

Final contribution: LinguaGap — Measuring the AI Language Gap in International Higher Education

LinguaGap investigates a structural but largely invisible inequity: international students in higher education receive measurably lower-quality AI assistance when working in languages other than English. Our project combines technical benchmarking, student fieldwork, and economic modelling to document this gap and translate the findings into actionable policy.

Our solution has three components. First, a multilingual performance audit of leading AI academic tools (ChatGPT-4o, DeepL, Gemini) across six languages — English (baseline), French, Arabic, Portuguese, Turkish, and Georgian — using standardised academic tasks scored by bilingual raters with a rubric grounded in translation studies, not English writing norms. Second, a student fieldwork strand combining an online survey (n = 80–120 international students) with semi-structured interviews and think-aloud sessions, designed to capture how AI language gaps translate into cognitive burden, reduced academic confidence, and diminished autonomous reflection. Third, an economic modelling component estimating the compounding time-cost and outcome disadvantage of consistently lower-quality AI assistance over an academic year, producing an institutional equity score for universities.

These three strands converge in a Linguistic Equity Index — a composite measure capturing both the severity of AI performance gaps across languages and their downstream educational impact. The Index is designed to be a practical instrument for universities, education ministries, and AI developers, not merely an academic metric.

Reflection on the process

Our proposal evolved substantially during Phase 2, shaped by four rounds of peer feedback that each identified a genuine blind spot.

Ella Faouzi's comment on European open-source models prompted us to reconsider the scope of our benchmark. We had initially focused on tools students most commonly use, but her question — whether models designed with multilingual intent from the outset actually close the gap — is a meaningful research question in its own right. We have since added a comparative layer to the audit, contrasting commercially dominant tools with multilingual-by-design models such as Aya. This shifts part of our analysis from documenting the gap to explaining it.

Ndeye Sophie Seck's question on cognitive fatigue and autonomous reflection gave us clearer language for a concern that was implicit in our proposal but under-theorised. We have now made student agency an explicit dimension of the Linguistic Equity Index, and redesigned the think-aloud sessions to observe not just whether students detect AI errors, but how they respond — whether they correct, accept, or disengage. This distinction matters for understanding the long-term learning consequences of the gap.

Nipun Ranchhod Navadia's methodological challenge was the most demanding to address. His point that performance gaps might reflect differences in prompting skill rather than model bias led us to more clearly separate the two components of our study: the benchmark uses standardised team-authored prompts to hold behaviour constant, while the fieldwork treats student usage patterns as a variable of interest rather than a confound. His challenge on defining "performance" without embedding English norms also pushed us to be explicit about rubric design — we are now committed to involving bilingual subject-matter experts in rubric validation and to building in a review step where raters flag culturally inappropriate criteria. Finally, his call for a clearer intervention pathway strengthened the policy brief: we have structured it around tiered recommendations for three distinct audiences (universities, ministries, AI developers), with an explicit logic linking each finding to a concrete institutional action.

Eline VIVET MALADRY's comment on language homogeneity was the most conceptually rich. Her observation that Arabic, French, and other languages are not monolithic — and that difficulties may be disciplinary or institutional rather than purely linguistic — has directly reshaped our interview design. We now explicitly ask participants to distinguish between moments when the AI failed because of language and moments when it failed because it didn't understand the academic or disciplinary context. This distinction is analytically important: a language failure calls for better multilingual training data; a disciplinary failure calls for better domain-specific fine-tuning. The two demand different interventions.

Across all four exchanges, the core contribution of Phase 2 was to move us from a well-structured proposal to a more epistemically honest one — one that acknowledges complexity, builds in methodological safeguards, and is clear about what it can and cannot claim to measure.

Confirm

Please log in

The password is too short.

Share