Who Sets the Test for an Underserved Language?
Adding a language to a benchmark creates a place on the scoreboard. Useful support also depends on who chooses the tasks, defines acceptable language, and can challenge the answer key.
The library question
Imagine a library testing an AI tool that translates its notices into a locally used language. A reviewer prefers formal written vocabulary; a reader prefers an everyday expression. Both understand the notice. The benchmark accepts only the reviewer's version. This is a hypothetical example, but it exposes a concrete design question: is the test measuring faithful meaning, conformity to a written standard, or usefulness to the people entering the library?
Those aims can overlap. They can also demand different answers. Before counting errors, the project must decide which differences count as errors and who has authority to decide. Our argument is that language evaluation needs speaker involvement at that earlier stage. Hiring fluent annotators after fixing the task and rubric leaves some of the most consequential choices beyond their reach.
This essay develops a proposal from historical research and community statements checked on October 2, 2026. It does not rank current models or claim to speak for a language community. The practical question is how to build a test whose claim matches the people and work it actually represents.
What shared comparison buys
The FLORES-101 paper, published in 2022, describes 3,001 sentences translated across 101 languages. Its source material came from English Wikinews, Wikijunior, and Wikivoyage. Professional translation, editing, and separate quality assessment supported aligned evaluation across language pairs. The authors acknowledge that non-English source sentences are translations, with possible translationese effects. They defend the tradeoff because alignment enables comparisons for neglected translation directions.
That is a serious benefit. A common test lets a researcher ask whether a changed system improved on the same material. Abandoning shared evaluation would make comparisons harder and could leave unsupported marketing claims with less resistance. Our interpretation is that comparability and local usefulness should be separate claims, each supported by its own evidence.
For the hypothetical library, a shared translation benchmark could justify trying a system. It would not establish that the library's readers understand its opening-hours notice. A locally composed notice and a translated travel passage ask different questions. The sensible next step is to add a test of the intended use, while preserving the shared result and its narrower meaning.
Speakers as research partners
In Masakhane's 2020 participatory research case study, Nekoto and colleagues describe a process in which participants could move between roles and help shape research questions. Their Igbo evaluation exposed practical problems with typing diacritics and decisions about dialect and terminology. The authors report that inconsistent dialects in training data and output complicated evaluation when a reviewer had to choose one. This is evidence from a particular research effort, not a demonstration that participation resolves every disagreement.
Masakhane's own mission and values explicitly place Africans in the ownership and direction of African-language NLP research. Its collaboration guidance rejects engagement that treats members merely as sources of annotations or translations. These are the organization's stated commitments, not a mandate from every African language speaker.
Our inference is that expertise should be able to change the question. If a participant identifies the wrong target audience, an evaluation project needs a way to reconsider that audience. Otherwise, it can invite knowledgeable people into a process whose basic mistake they are allowed only to annotate. A fluent reviewer should be able to say that the proposed test would reward the wrong behavior.
Define the use before the answer
Bender and Friedman's 2018 proposal for data statements offers a useful starting point. It calls for descriptions of selection rationale, language variety, speakers, annotators, and the circumstances in which language was produced. It also recognizes a tension between detailed documentation and privacy, especially for small groups. The proposal helps delimit generalization; it does not itself confer authority on a project.
For our library example, we propose beginning with a plain sentence: the tool helps these readers understand these notices in this setting. Participants should then select realistic tasks and explain why they matter. Translating a schedule, finding a requested book, and explaining a membership form need not share a success criterion. Starting with the work makes it harder to substitute whichever dataset is easiest to obtain.
The intended audience should also guide the reference material. Does the service promise formal written language, familiar conversational wording, or both? Must it understand requests that mix languages? A small project can choose a narrow scope honestly. It should name what it covers and identify unanswered questions, rather than describe a convenient sample as the entire language.
Recruitment needs the same specificity. A teacher, translator, frequent library visitor, and occasional reader might bring different relevant knowledge. In this proposed design, a project records which perspectives shaped the test and invites challenges from those it missed. It avoids demanding personal disclosures simply to make its documentation appear complete.
When the answer key is disputed
Our proposed annotation process separates disagreements that require correction from disagreements that reveal several acceptable answers. A mistranslated closing time is an error. Competing expressions for the same clear instruction may require multiple references or a rubric that judges meaning and audience fit separately. The distinction must be worked through with people who understand the language and the task.
This is not a proposal to accept every output. Participants should explain their judgments, inspect context, and decide which variations preserve the intended meaning. Where disagreement remains, retain it in the research record. A majority vote can produce a convenient label without explaining whether the minority noticed a genuine error or used a different legitimate convention.
For the library trial, ask readers to act on a notice, then examine where the wording helped or confused them. Treat that as a small local study, with its sample and uncertainty visible. Keep results by task and intended variety where the sample permits. An aggregate can summarize these results, but should not conceal a failure on the very task that motivated the service.
There is a cost: richer evaluation takes time and skilled attention. The response should be to narrow the promise or fund the work. Quietly replacing a difficult assessment with a cheap proxy leaves the original promise unsupported.
Give participation a decision
The following is our recommendation, not a procedure validated by the cited studies: before collecting labels, agree which decisions participants can change. These should include task selection, acceptable variants, disputed examples, and how the project describes its coverage. Budget for their expertise and for later revisions. Give contributors a way to challenge a misleading public claim after the dataset is released.
No small panel can settle a language's future. Nor should local participation require unanimity before useful research can begin. A workable project makes bounded decisions, records unresolved disagreements, and names who may reopen them. It can produce a repeatable benchmark without presenting its temporary answer key as the final authority on correct language.
Return to the library notice. A useful evaluation would let the reader's difficulty change the test, even if the model matched the original reference perfectly. It would also retain evidence that the tool worked when it did. That combination gives the library something it can use: a reasoned account of which communications the system helps with, for whom, and where further work is needed.
Sources
- Goyal et al., The Flores-101 Evaluation Benchmark for Low-Resource and Multilingual Machine Translation. TACL, 2022.
- Nekoto et al., Participatory Research for Low-resourced Machine Translation: A Case Study in African Languages. Findings of EMNLP, November 2020.
- Masakhane, mission, values, and collaboration guidance. Undated page; consulted October 2, 2026.
- Bender and Friedman, Data Statements for Natural Language Processing: Toward Mitigating System Bias and Enabling Better Science. TACL, 2018.
Related reading
- The Language Variety Becomes the Bias Probe
- The Compressed Model Becomes the Language Tax
- The Benchmark Becomes the Curriculum
Production: commissioned by the site operator; researched and drafted by GPT-6 Astra; editorial and source review by the coordinating AI assistant.