Webinar
§ 06 · Frequently asked · Updated July 2026

Frequently Asked Questions.

Information about data commons, the New Commons Incubator, and the current focus on Indigenous languages and cultures. Still stuck? Email the team.

01

Background

What data commons are, the problems they aim to solve, and examples of existing initiatives.

4 questions

Data commons are collaboratively governed data ecosystems designed to pool and provide responsible (and governed) access to diverse, high-quality datasets from one or multiple sectors to enable the development and deployment of generative AI applications that address public-interest challenges.

This approach is distinct in that it treats data as a shared resource governed by a community rather than controlled solely as individual or institutional property. It further recognizes that open data, open by default and often fully public, does not, on its own, offer a complete solution. In an evolving data landscape, ensuring that openness supports public interest goals requires addressing how recognition and benefit flow, particularly where publicly available data can be widely reused without clear benefit returning to the communities that steward it. They can support, for instance, the responsible development of AI, climate adaptation and biodiversity conservation, public health and pandemic preparedness, sustainable agriculture, scientific research, disaster response, and the preservation, revitalization, and responsible stewardship of Indigenous knowledge, languages, cultural heritage, and data economies.

More details on data commons can be found here.

Data commons aim to address several structural failures:

  • Access asymmetries: Many large organizations seek to use open data to extract value from others while limiting access to their own proprietary data. Data commons can introduce mechanisms through which communities can influence who accesses their data, how, and for what purpose.
  • Coordination failures: There are often high transaction costs for sharing data with one actor or on a one-off basis (e.g. formulating bespoke data-sharing agreements, standards). Data commons can create more standardized processes for sharing data under consistent rules.
  • Trust deficits: In many data-sharing arrangements, the interests and expectations of the public or contributing communities are not meaningfully incorporated. This erodes the social license and limits the legitimacy of any data (re-)use. Data commons try to center these concerns through governance, oversight, and access conditions.
  • Underutilization: A problem plaguing data re-use is that the most valuable datasets are not used for the public good but instead restricted. Data commons can expand access to a broader set of actors provided they commit to upholding the standards outlined by the organizers.

In these ways, data commons can point toward more sustainable and equitable data sharing relationships for the public good. They can address some of the concerns fueling data scarcity and the ongoing “data winter”. They can recenter the conversation on AI to focus on public-interest AI and decision-making.

Our research finds that data commons arrangements tend to exhibit a set of commonly implemented characteristics that reflect the current state of the field:

  • Public Purpose: Aims to supply the data needed to solve public problems and develop public-interest AI.
  • Participatory Governance: Possesses a governance structure that offers meaningful avenues for commons members to exercise agency in how their data is used.
  • Accountability Mechanisms: Monitors the implementation of the commons’ rules on an ongoing basis. Has systems in place on how to manage disputes and how to hold those accountable who do not follow the rules.
  • Contributor Benefits: Offers clear benefits to stakeholders who have contributed to the commons. This might include monetary benefits, early access to AI tools or research developed, attribution, insights for or about them, free membership, access to infrastructure to develop new AI tools, etc.
  • Access Controls: Includes clearly defined rules governing how, when, and by whom datasets can be accessed and used. This may involve standardized frameworks such as Creative Commons licenses, as well as tiered, restricted, or consent-based access models.
  • Funding and Support: The data commons is often funded as a public good, often with support of government, philanthropy, or via a membership/cooperative model.

Data commons are an emerging practice. Relatively few existing initiatives fully meet the definition of collaboratively governed systems where communities exercise authority over data. Instead, current examples span a spectrum of models that incorporate some commons-like features, such as shared access, multi-stakeholder participation, or public-purpose orientation, without always implementing collective governance. The initiatives below illustrate this range:

  • CLARIN: CLARIN provides the infrastructure to combine language data from across European institutions for research purposes. Governance is multi-stakeholder at the institutional level, but decision-making authority is exercised by member states and organizations rather than a broader contributing community.
  • Common Voice: Common Voice is a crowdsourced open voice dataset that can be used to train AI-driven voice applications. The initiative aims to broaden access to voice data for non-English languages and other groups typically underrepresented in voice datasets. The initiative is led by the Mozilla Foundation.
  • Language Data Commons of Australia: The Language Data Commons of Australia is a partnership between the Australian Research Data Commons and the School of Languages and Cultures at The University of Queensland that seeks to make available Australian language data for both “academic and non-academic uses”.
  • Querido Diário: Open Knowledge Brasil is seeking to develop a data commons for municipal official gazettes. Its Querido Diário platform aims to generate access to these datasets and help analyze information within them. They seek to improve decision making at the local level in Brazil.
  • Malawi Voice Data Commons: NYU Peace Research and Education Program’s Malawi Voice Data Commons, developed in collaboration with Ushahidi, UNDP, and Mozilla Foundation, enables rural Malawians to report emergencies in native languages, creating multilingual, AI-ready datasets for humanitarian response and language preservation. The pilot will take place in Malawi with plans to scale across Sub-Saharan Africa.
02

About the Incubator

What the Incubator addresses, what participants gain, deadlines, eligibility, and evaluation.

6 questions

The New Commons Incubator is a structured program to help individuals and teams develop detailed and actionable proposals for the creation of data commons for the AI era.

It addresses a critical issue: the limited ability of communities to have a say in how their data is used or meaningfully benefit from that use. By creating and improving data commons, which are collaboratively governed data ecosystems that pool and provide responsible access to diverse, high-quality datasets, the gap in AI development can be filled.

Through this program, you will join a cohort of your peers for a variety of peer learning, networking opportunities, and training from internationally recognized experts. Though the Incubator does not offer direct financial backing, we will work to transform your idea for a data commons into a viable concept (and proposal) that can be presented to funders who can scale it.

Through the Incubator program, participants will be able to:

  • Connect with a global cohort of data stewards;
  • Network with and present their work to data commons funders;
  • Work with experts to develop a robust plan for setting up a data commons;
  • Attend relevant training sessions on several topics including designing governance frameworks, setting up social licenses, and building AI-ready datasets; and
  • Have their proposals reviewed by industry experts and peers.

Participants will not receive funding directly from Microsoft, UNESCO, or The GovLab—though we will make efforts to connect participants to funders at the end of the program. The Incubator will not provide computational infrastructure nor will it supply data for the proposed data commons.

We will not seek any data from participants and all data and materials will remain the sole property of applicants.

  • Applications open: 19 June 2026
  • Applications close: 14 August 2026 at 23:59 ET
  • Cohort Selection: Late August 2026
  • Kickoff: Mid-September 2026
  • Virtual programming: October 2026 – April 2027
  • Showcase: April 2027

Participants in the Incubator for the Indigenous languages and cultures must be Indigenous or have an endorsed relationship with Indigenous communities. Through the application process, we will strive to ensure that applicants have the necessary resources to establish a data commons that can meaningfully protect their community’s interests and ensure that benefits are equitably distributed.

All proposals must satisfy minimum requirements. Applicants that fulfill these minimum requirements will be reviewed by a group of expert judges (drawn from the Steering Committee or approved by them). Judges will be asked to evaluate proposals based on a rubric that examines specific criteria. The questions in the lefthand column will be included in the application form to communicate what judges are looking for specifically for each category.

Each proposal will be reviewed by three judges. Scores will be averaged and the highest scoring proposals will be shared with the collective steering committee. We will aim for fairness and will aim for all proposals to be reviewed by the same number of judges. In the event that not every judge submits scores and some applications are left with two reviews instead of three, we will remove the lowest score from any proposal that was reviewed by three judges before calculating an average.

In collaboration with the Steering Committee, the Incubator organizers will conduct due diligence of the top scoring applicants to ensure they each have a strong relationship with the Indigenous community they intend to serve. The project team will ensure that the proposals are reviewed by experts in the region of focus who can properly assess whether the relationship is sound.

Barring extenuating circumstances (e.g. an inability to confirm community support or their own credentials), teams will then be invited to join the cohort. If any participant withdraws, the next highest-ranking participant will be invited in their stead.

We ask that you submit your application using the submission system here. If you are unable to access the link you can download a submission template here. Please answer the questions to the best of your ability. You can submit this and all supplementary materials via email. Please title your message “[SUBMISSION]: New Commons Incubator”. Please try to submit all materials in one email.

03

Rules & Standards

Ethical standards, the role of each actor, core decision-making principles, and global scope.

4 questions

For Participants:

All participants must align their data commons proposals with Indigenous data sovereignty principles. Should a participant refuse to align their work with Indigenous data sovereignty principles, they will be immediately asked to leave the program. This includes but is not limited to:

  • Including unauthorized data in the commons;
  • Activating the commons without the necessary social license to operate;
  • Failing to recognize Indigenous laws, treaty obligations, and other standards; and
  • Proposing access to data in a format that would not align with Indigenous governance.

To help mitigate this risk, in the review of applications we will ask the jury to look for whether the project aligns with Indigenous data sovereignty principles. We will submit these standards to the Steering Committee. We will modify this protocol if they identify any basic principles or obligations missing.

For the project team:

The project team embraces the CARE Principles for Indigenous Data Governance. This means that, at all times, it will operate to ensure:

  • Collective Benefit: The team will seek to design the Incubator so as to ensure that Indigenous Peoples can derive benefit from the data. Guided by the steering committee, it will select proposals that evidence a concern for fair, equitable outcomes, improved governance, and citizen engagement. The Incubation training materials will promote similar outcomes.
  • Authority to Control: The Incubator recognizes Indigenous Peoples’ rights and interests in Indigenous data and seeks to empower them to control their own data. Recognizing these rights and interests, project team members will never request data from participants or mandate that they use specific systems. The data commons model seeks to offer an alternative to existing open data frameworks by addressing access asymmetries—allowing Indigenous participants themselves to decide who they want to make their data available to, under what conditions, for what purpose.
  • Responsibility: The Incubator project team recognizes that it has a responsibility to communicate how it will support Indigenous People’s self-determination and collective benefit. We seek to foster positive relationships by setting specific guardrails on ourselves and using this capacity building initiative to support the development of an Indigenous data workforce that can control its own data. We commit to working with the Indigenous Steering Committee, and individual participating teams, to ensure that all the support we offer is grounded in the worldview, lived experiences, values, and principles of Indigenous Peoples.
  • Ethics: While the Incubator will not involve data directly, one of our central focuses is to promote the ethical use of data at all stages of the data lifecycle and across the data ecosystem. We will work with our Steering Committee to address imbalances of power and use our final showcase to try to address imbalances of resources. Any materials that we use that draw from traditional knowledge will make use of Traditional Knowledge Labels. All training will acknowledge and proactively emphasize the rights of all Indigenous Peoples, cultures, and knowledges. All our work will align with the ethical United Nations Declaration on the Rights of Indigenous Peoples.

Details on how roles and responsibilities will be broken down across each of the organizations sponsoring this Incubator can be found in the section below.

The Steering Committee

The Steering Committee is a group of Indigenous experts and community representatives who have been invited to drive all Incubator activities. While the project team (The GovLab, UNESCO, and Microsoft) propose approaches, we look to the Steering Committee to co-design this work, ensuring it remains aligned with community priorities, cultural context, and ethical standards. This work may involve:

  • Strategic Guidance and model development: Contribute to the incubation of replicable models that are grounded in real-world constraints and can scale across contexts.
  • Review and validate core materials: Feedback on draft Incubator materials to ensure alignment with community priorities, sensitivity to cultural and regional dynamics, and adherence to strong ethical standards.
  • Contextual evaluation and selection: Participate in the review of submissions, helping ensure that proposals are assessed with an informed understanding of local challenges and opportunities.
  • Shaping engagement and visibility: Help define how they and others should be involved and how your contributions should be recognized publicly, in ways that are meaningful and appropriate.

Project Team

The project partners—The GovLab at NYU, UNESCO, and Microsoft—act as the main facilitators of this work, bringing together resources, knowledge, and expertise as guided by the steering committee. Drawing on their unique expertise, project team members will be involved in the following:

The GovLab: Drawing on a decade of experience designing and operating innovative and responsible data collaboratives, along with building a cohort of data stewards, The GovLab will lead all efforts to develop training, identify resources, and recruit relevant guest faculty. It will offer direct support to teams as they design programs, decide upon governance and evaluation frameworks, and seek funding. It will take particular care in designing materials that are relevant and respectful to Indigenous communities and their Traditional Knowledge. We will make use of TK Labels and ensure that all Incubator activities align with the CARE Principles (as outlined in Section 7).

UNESCO: UNESCO brings multilateral legitimacy, convening power, and a unique normative mandate as the UN’s lead agency on multilingualism, culture, open science, data governance, and AI ethics. The Incubator will be anchored in UNESCO’s Digital Policies and Digital Transformation Sector and its Data Governance in the Digital Age Initiative, which advance freedom of expression, media development, media and information literacy, universal access to information, and the preservation of documentary heritage, while promoting people-centred and rights-based approaches to digital transformation.

As the UN lead agency for the International Decade of Indigenous Languages (IDIL 2022–2032), UNESCO coordinates global efforts to safeguard, revitalize, and promote Indigenous languages and related knowledge systems. Building on long-standing relationships with Indigenous communities, UN Member States, and cultural and memory institutions, UNESCO will connect Incubator teams with relevant partners, provide strategic guidance, and support culturally grounded community outreach. It will also ensure that the Incubator’s work is aligned with and contributes to broader UN digital inclusion and language-related initiatives, including IDIL and its associated action plans.

Microsoft: Microsoft is a global technology company with experience supporting public-interest data initiatives and responsible AI efforts. Microsoft’s involvement in the Incubator will be led by the Open Innovation team, which focuses on advancing open data and data commons and promoting accessible, reusable data to address societal challenges.

Microsoft will contribute their expertise when requested, such as reviewing proposals, offering guidance on data governance, AI readiness, responsible data reuse, technical design considerations, and possible infrastructure requirements. The purpose of this support is to help participants understand options and plans for future implementation, rather than to operate, host, build, or resource the Data Commons itself.

The Project Team will not act as representatives of Indigenous communities. They will work with the Steering Committee to take active precautions to ensure that all work remains owned by the individual teams and that they, and the communities they represent, retain full agency and control over their information.

Though this project seeks to build skills and not data commons directly, we will never ask for data. We will operate with strict confidentiality through all phases of this Incubator. We will not allow access to any participant proposals or community data assets except for the purpose of direct programme support. No community data or proposal content will be accessed, retained, or used by the project team for any commercial purpose during or after the program. All submissions and related documents from applicants and participants will be destroyed after final reporting.

Institutionally, each of the organizations in the project team acknowledge and embrace the virtues of diversity, transparency, accountability, and promoting equity. We seek to center the public good, which includes empowering disadvantaged groups and fostering inclusion and equal opportunity.

For this Incubator, we prioritize two organizing principles when it comes to making major decisions about this Incubator:

  • Shared Benefits: This work must be used to help Indigenous communities better harness the data they possess for their own benefit. We seek to ensure that we do not exacerbate power asymmetries and will be cognizant of the specific moral, political, legal, and economic context to ensure that benefits are equitably distributed.
  • Indigenous-Driven: The project team recognizes the need for Indigenous people to make decisions about their own communities and their own data. We will set application requirements and work with the steering committee to ensure that project teams can speak for the communities they claim to represent. We will look to the steering committee for guidance on all resources and the training curriculum to ensure that the work is useful, relevant, and incorporates the knowledge and the guidance already developed and adopted by Indigenous communities in other contexts (e.g. CARE Principles, Traditional Knowledge Labels).

This Incubator cannot include or represent all languages, particularly not when there are thousands spoken worldwide. We will instead leave this Incubator open to any Indigenous community and include those interested in participating, who are able to meet our application criteria.

Once we have a final cohort assembled, we will work to ensure that materials are useful and relevant to those selected.

§ 05 · Still got questions

Couldn't find what you need?