The Brief
Society & Ethics 4 min read

The Language Gap AI Creates, and How Communities Are Closing It

NAVION

Share

Large language models are trained on vast amounts of text, but that text is not evenly distributed across the world’s languages. The result is a structural imbalance: generative AI performs well on widely spoken languages and poorly on minority ones, particularly those with limited digital presence. This is not a minor technical footnote. It shapes which communities can benefit from AI tools and which are effectively left out. Yet a quieter story is unfolding alongside this inequality, one that most coverage of AI and language tends to overlook.

The Structural Problem: Not All Languages Are Equal Online

Generative AI systems learn from data. The more text a language has online, the better a model can represent it. For languages spoken by millions, this creates a virtuous cycle. For minority languages, it creates the opposite.

Hula, spoken in Papua New Guinea, has around 10,000 speakers. Tetun, the lingua franca of Timor-Leste, has just over 1 million. Dinka, spoken in South Sudan, has approximately 5 million speakers but very limited digital representation. None of these languages appear in AI training data at a scale that would allow models to handle them accurately. When they appear in AI outputs at all, they are likely to be misrepresented.

This is not simply a technical limitation waiting to be fixed. It reflects a deeper pattern: the communities whose languages are underrepresented online are often the same communities with fewer resources to advocate for inclusion in commercial AI development pipelines. The gap compounds itself.

Communities Building Their Own Tools

Here is what most coverage misses: some of these communities are not waiting for large technology companies to solve the problem. They are building their own solutions, using AI as a construction tool rather than a finished product.

Vavanagi is a language documentation platform built entirely by and for the Hula community in Papua New Guinea. Led by Bri Olewale, community members used AI coding tools to design and build a platform they could not have assembled otherwise. The Hula community lacks the workforce of linguists and software engineers typically required for this kind of project. AI lowered that barrier. The platform now has more than 80 users who have collectively contributed over 12,000 English-Hula translations. The long-term goal is to accumulate enough data to build a Hula language translator app.

Tulun takes a different approach. Built in partnership with the local non-governmental organisation Maluk Timor, it helps health educators translate health education material into Tetun. Staff can manage their own list of approved terms and phrases, ensuring that translations are accurate and contextually appropriate for Timorese health workers. The AI model adapts to user-uploaded content, making translation a collaborative process between automated systems and local experts rather than a one-way output.

The Dinka-English dictionary, developed by Alier Makoi Achuoth from South Sudan, addresses a different need: expanding and preserving Dinka vocabulary in digital form. Achuoth used AI models to help design the app, collect data, and organise it. His stated motivation was to develop a practical solution rather than wait for an external organisation to act.

Three tools, three communities, three different purposes. What they share is the use of AI to channel local knowledge back into the community that holds it, on the community’s own terms.

Why This Matters Beyond Language

The significance of these projects extends well beyond linguistics. They challenge a common framing in discussions about AI and the Global South: that lower-income countries and minority communities are primarily affected parties, recipients of technology developed elsewhere, subject to its consequences but not its authorship.

The reality is more layered. AI adoption in Kenya and Nigeria, for instance, is reported to be as high as in the United States, which complicates the assumption that low-income countries are simply lagging behind. Small businesses in the Global South are using AI-assisted translation to compete in markets where professional translation was previously unaffordable. These are not edge cases. They are early signals of a different kind of AI adoption story.

Cat Kutay, a computer scientist of Aboriginal descent from Charles Darwin University, has noted that several First Nations communities in Australia are now developing their own language technology. The framing she offers is instructive: just as these communities adopted the motorcar when it became relevant to how they lived and moved, they are now adopting AI because it is relevant to translation and storytelling. The technology is being taken up on their terms, for their purposes.

This matters for how researchers, funders, and policymakers think about AI development. A top-down model, in which large institutions build tools and communities receive them, misses the agency that is already present. Community-led projects are not a workaround. They may produce the most accurate and culturally grounded digital representations of minority languages that exist.

In Short

AI’s language gap is real: minority languages with limited online presence are poorly served by large language models, and that gap does not close automatically. But communities are not passive in the face of this. Projects like Vavanagi, Tulun, and the Dinka-English dictionary show that AI can function as an enabling tool, lowering the cost and technical barrier enough for communities to build what they need themselves. The broader implication is that the most meaningful AI applications for underrepresented communities may not come from large technology companies. They may come from within.

Based on reporting from The Conversation AI.

Written by

NAVION