Mozilla’s Common Voice now features 60 hours of speech datasets in eight new Indigenous languages. Meet the Taiwan Language volunteers at RightsCon on February 27, 2025, in Taipei
As thousands of languages face extinction, linguistic and heritage preservation has never been more critical. In Taiwan, a grassroots volunteer community is utilizing Mozilla’s Common Voice platform — the world’s largest public participation open speech dataset, to preserve Indigenous languages and help build inclusive, voice-enabled AI solutions.
Common Voice, a volunteer-led initiative with over 200 languages, including Traditional Mandarin and Taiwanese Hokkien, will now include eight Indigenous Formosan languages: Atayal, Bunun, Paiwan, Rukai, Oponoho, Teldreka, Seediq, and Sakizaya. The launch of these languages coincides with International Mother Language Day – this year celebrated on February 21, 2025.
Over 60 hours of speech data have already been collected by the local Indigenous language teachers around the island, with the help of the Mozilla Taiwan community, led by Irvin Chen, in collaboration with the Wikimedia Foundation in Taiwan. The dataset will be available to download in June.
“We carry our identity and heritage through our language. By bringing our culture into technology, we’re not just preserving words, we’re keeping our cultures alive,” says Chen.
The expansion of Taiwanese Indigenous languages is part of Mozilla’s Open Multilingual Speech initiative, a broader effort to support more ultra-low-resource and indigenous languages. This first round has included over 70 communities in Southeast Asia and beyond.
“We love seeing local communities mobilize around their languages. Common Voice is really their project. This embodies the true spirit of open-source collaboration and community engagement in shaping ethical AI,” says EM Lewis-Jong, Common Voice Product Director at Mozilla Foundation.
Common Voice's datasets can be used by anyone for free and have been utilized widely, from developing audio translation software in healthcare solutions to voice applications that teach women to better understand and exercise their land rights.
“We carry our identity and heritage through our language. By bringing our culture into technology, we’re not just preserving words, we’re keeping our cultures alive"
Irvin Chen, Taiwan Community Volunteer Lead
Meet the Volunteer community at RightsCon
The Mozilla Taiwan community will have a dedicated booth at RightsCon Taiwan on February 27, where attendees can learn more about Common Voice and help contribute to the initiative. Additionally, on Saturday, February 22, the Taiwanese language community will be participating in the g0v bi-monthly hackathon. You can also learn more about the Common Voice in Taiwan by visiting their project website: moztw.org/common-voice
Join the Movement
At Mozilla, we believe that everyone can shape AI. Be part of a global community advancing these efforts today, by joining our Common Voice community Discord and signing up forMozilla Foundation’s newsletter to get updates about other Mozilla initiatives.
Facts Only
* Common Voice now features 60 hours of speech datasets in eight new Indigenous languages: Atayal, Bunun, Paiwan, Rukai, Oponoho, Teldreka, Seediq, and Sakizaya.
* The initiative involves Taiwan Language volunteers utilizing the Common Voice platform.
* Over 60 hours of speech data have been collected by local Indigenous language teachers with help from the Mozilla Taiwan community and the Wikimedia Foundation in Taiwan.
* The dataset will be available for download in June.
* The launch coincides with International Mother Language Day on February 21, 2025.
* Irvin Chen is the Taiwan Community Volunteer Lead.
* Common Voice datasets are free to use for any purpose.
* The expansion is part of Mozilla’s Open Multilingual Speech initiative.
Executive Summary
Full Take
The mobilization around linguistic preservation through technology reveals a tension between digital access and cultural sovereignty. The narrative frames the incorporation of Indigenous languages into AI development as an act of keeping cultures alive, which operates on a humanist imperative. This framing positions technological tools not as neutral instruments but as vehicles for cultural agency; the collaboration emphasizes community-led action rather than top-down imposition, reflecting the spirit of open-source principles championed by Mozilla. The expansion showcases how global, open-source frameworks can be leveraged to address localized crises of extinction, shifting the locus of preservation from centralized institutions to grassroots communities. However, the emphasis on "keeping cultures alive" and building "voice-enabled AI solutions" necessitates deeper scrutiny regarding who controls these technologies and what definitions of "preservation" are being enacted by external frameworks. The structure highlights a pattern where localized cultural needs intersect with global technological initiatives, raising questions about whether the mechanism of inclusion truly supports the autonomy of the originating communities or merely facilitates data extraction for broader models.
What assumptions underpin the assertion that bringing culture into technology inherently serves preservation? How does the structure of "open-source collaboration" translate into genuine cognitive sovereignty when dealing with ultra-low-resource languages? Does framing this as an act of survival risk prioritizing the functionality of the tool over the intrinsic, non-quantifiable values of linguistic heritage?
Sentinel — Human
This text appears to be legitimate news reporting that effectively weaves factual data with community-focused testimonials, suggesting a human journalistic origin.
