Back to Blog
Blog

Data Commons as Infrastructure for Indigenous-Led AI

PublishedOctober 5, 2026
SourceODPL · The GovLab
Data Commons as Infrastructure for Indigenous-Led AI

By Stefaan Verhulst and Andrew Zahuranec

What would it take for Indigenous communities not simply to be represented in the emerging AI ecosystem but to shape, govern, and benefit from it?

That question is what animated our event during UNGA High-Level Week on Advancing Public Interest AI Through Data Commons for Indigenous Languages and Cultures. 

The discussion brought together Indigenous leaders and technologists, funders, researchers, international organizations, civil society, and industry around a growing challenge: As AI becomes embedded in education, public services, cultural production, and everyday life, many Indigenous communities are being drawn into AI systems that they did not design, using data they may not control, and reflecting assumptions and values that may not be their own.

The stakes are particularly high for Indigenous languages and cultural knowledge. AI could become a powerful instrument for language revitalization, education, cultural transmission, and economic opportunity. But without different approaches to data, infrastructure, governance, and investment, it could just as easily reproduce longstanding patterns of extraction while amplifying misinformation. This leads to weaker foundations of reliable information for AI and other systems that subsequently harm communities.

“The challenge is to ensure that Indigenous communities can govern, use, and benefit from their own knowledge in the digital era.” (Participant)

Five themes emerged from the discussion.

Screenshot 2026 10 05 at 11.40.15 Am

1. Indigenous AI must start with Indigenous agency

A central message was that Indigenous languages cannot simply be approached as another category of “low-resource” languages to which existing AI techniques are applied. Languages exist within particular communities, histories, cultural practices, and systems of knowledge.

This means moving beyond efforts to make mainstream AI systems marginally more representative. The ambition should be to enable Indigenous-led AI solutions—built by and with communities according to their own priorities.

This is especially important given persistent inequalities in access to technology, connectivity, education, and technical expertise. Indigenous peoples remain substantially underrepresented among those designing advanced technologies while their communities and cultures continue to be subject to both romanticization and stigmatization.

One powerful idea raised during the discussion was that wisdom itself should be understood as expertise. Communities already possess knowledge, capabilities, and governance practices. The challenge is to ensure they also have the technical capacity, infrastructure, financing, and institutional leverage to translate that expertise into the digital and AI systems increasingly shaping their futures.

Indigenous voices therefore need to be present not only as stakeholders consulted about technology but where technology and data policies are actually being set.

“We are a special class within the legal infrastructure. We are not states” (Indigenous Leader)

2. Data sovereignty requires infrastructure, not just principles

A second theme concerned ownership and control.

Participants repeatedly returned to the reality that Indigenous data and cultural knowledge have often been collected, digitized, stored, or commercialized by institutions outside the communities from which they originated. Past experiences with language digitization and preservation efforts also demonstrate that people may make very different choices when they fully understand how their contributions could ultimately be used.

“Can we redirect energy so that the work is beneficial and not just an exploitative act?” (Indigenous Leader)

Consent at the moment of collection is therefore not enough.

Communities need enduring mechanisms through which they can determine who can access data, for what purposes, under what conditions, and how benefits are distributed. Possible institutional mechanisms include data trusts, community-controlled digital infrastructure, stewardship arrangements, and data commons.

The broader lesson is that Indigenous data sovereignty needs to be made operational. Governance principles matter, but they must be translated into actual institutions, technical architectures, stewardship roles, access rules, and mechanisms for accountability.

This is where data commons can be especially important—not necessarily as spaces where everything is open but as shared infrastructures through which communities can collectively govern access and reuse according to their own values.

3. The goal should be sustainability, not permanent dependence

Perhaps the strongest recurring message concerned sustainability.

Too many promising Indigenous technology and language initiatives remain dependent on short-term grants. A project receives philanthropic funding, builds a dataset or tool, demonstrates its potential—and then confronts a funding cliff.

Participants called for a transition away from this cycle toward models capable of supporting long-term Indigenous economic and technological agency.

This does not mean philanthropy has no role. Philanthropic capital can be essential for experimentation, capacity building, and infrastructure. 

Possible approaches discussed included trust funds, shared infrastructure, community-owned enterprises, benefit-sharing mechanisms, and revenue models built around Indigenous-controlled assets and services.

The objective is not simply to sustain individual projects. It is to build an ecosystem in which communities possess the resources and capabilities to continue developing solutions on their own terms.

“This can be a win-win. AI models that really take into account the central issue of common goods are the only ones that can benefit all our societies. Multilingualism is fundamental to the future of the digital world.” (Indigenous Leader)

The shift is important. It is a call to move the ecosystem from funding projects for Indigenous communities toward investing in Indigenous capacity, ownership, and economic independence.

4. Local and contextual AI may offer an alternative to one-size-fits-all models

The discussion also challenged the assumption that Indigenous inclusion requires putting every language and cultural resource into ever-larger foundation models.

Work with low-resource languages points toward another possibility: accomplishing more with carefully targeted data, community knowledge, and feedback from native speakers. Rather than creating enormous centralized datasets, smaller retrieval, augmentation, or locally controlled layers could potentially allow communities to develop useful language capabilities while retaining greater control.

This connects to a broader proposition raised during the discussion: good AI may need to be contextual AI.

“It’s about building a digital and AI ecosystem in which Indigenous languages, cultures, and knowledge have a meaningful place, where communities have greater agency over their data and where innovation serves the public interest.” (Indigenous Leader)

Instead of expecting one universal system to adequately represent thousands of languages, cultures, and epistemologies, we should explore architectures that allow AI to be adapted locally—reflecting community norms, knowledge, governance rules, and linguistic practices.

Such an approach also has implications for infrastructure. Each community does not necessarily need to recreate every component of the AI stack. Shared technical resources can coexist with local control, provided the architecture is designed around subsidiarity: decisions should be made as close as possible to the communities affected by them.

The goal is not to replace one centralized authority with another but to enable communities to collaborate as equals while retaining autonomy.

5. Collaboration is essential—but it must be centered on Indigenous leadership

None of this can be accomplished by communities acting alone. There are important roles for philanthropy, universities, international organizations, governments, civil society, and technology companies.

But collaboration needs to be structured differently.

The traditional model often begins with an external organization designing a program and subsequently inviting Indigenous communities to participate. The discussion suggested reversing that logic: start with community-defined needs and then assemble the capabilities required to address them.

Universities can contribute technical or analytic expertise; philanthropy can provide patient capital; companies can contribute infrastructure and expertise; international organizations can connect initiatives to emerging debates around data governance and benefit sharing; and civil society organizations can contribute practical experience and networks.

What matters is who sets the direction, who makes the decisions, and who ultimately owns what is created.

This also suggests a need for stronger connective infrastructure across Indigenous initiatives themselves. Directories, networks, shared frameworks, and opportunities to aggregate currently fragmented efforts could help communities learn from and support one another. The aim should not be to homogenize Indigenous communities—quite the opposite. It should be to create mechanisms through which diverse communities can share capacity without surrendering autonomy.

From principles to action

The discussion ended with a palpable sense of urgency. Communities cannot wait for the theory to be perfected before acting.

“AI is racing ahead. The urgency calls us to really work together.” (Participant)

AI systems are already shaping education, information, language, and cultural representation. Choices being made today about training data, infrastructure, intellectual property, governance, and investment will determine who possesses agency tomorrow.

The challenge, then, is not simply to preserve Indigenous languages by incorporating them into AI. It is to ask more consequential questions:

Can we build an AI ecosystem in which Indigenous communities have the capacity to determine how their languages, knowledge, and data are used—and to build and benefit from AI themselves? How can AI ecosystems involve communities directly and be more participatory?

Doing so requires moving simultaneously on several fronts: Indigenous-led technical capacity; community ownership and governance of data; sustainable economic models; contextual and locally appropriate AI architectures; and partnerships that place Indigenous communities in positions of decision-making power.

Data commons could provide an important piece of that infrastructure. But the ultimate goal is larger than any particular technical or institutional model.

It is a shift from extraction to stewardship, from consultation to governance, from grant dependence to economic agency, and from inclusion in somebody else’s AI future to the capacity to build one’s own.

***

We will strive to act on the signals received from this event through our New Commons Incubator, which seeks to provide dedicated support for Indigenous communities looking for ways to better govern and act on their data. If you’d like to support us in this effort, please reach out to us at newcommons@opendatapolicy.org.