Skip to content
Corpshore Australia

AI delivery

AI training data for Australian English and te reo Maori

By Corpshore Australia Insights Team7 min read

Models trained mainly on US or UK English routinely mishear Australian idiom and miss te reo Maori entirely. Locally grounded training data is what closes that gap.

Most large language models and speech systems are trained predominantly on American and British English, and Australian English differs from both in spelling, idiom and accent in ways that matter for real-world accuracy. A model trained mainly on US text will default to American spelling, miss Australian idiom entirely, or misinterpret it, and speech models trained mainly on US or UK accent data show measurably higher error rates on Australian accents, particularly regional ones.

Why does Australian English need its own AI training data?

For a business deploying an AI system, a voice assistant, a customer service chatbot, a document processing pipeline, into the Australian market, this is not a cosmetic issue. A support bot that consistently mishears an Australian customer's accent, or that responds in American spelling and phrasing to an Australian audience, damages trust in the interaction even when the underlying answer is correct. Locally grounded training data, text written by Australian speakers in Australian English, speech recorded from a representative range of Australian accents, is what closes that gap, and it is a distinct data problem from simply having "more English data," since volume does not fix a mismatch in dialect and accent distribution.

What does "AI training data" actually cover in practice?

It covers several distinct workstreams that sit behind any model performing well in a specific market: data collection (gathering genuinely representative text, speech or image data rather than relying on whatever is easiest to scrape), annotation and labelling (humans tagging that data so a model can learn from it accurately), speech and audio work specifically (transcription, accent and dialect tagging, text-to-speech and automatic-speech-recognition training data), and RLHF, reinforcement learning from human feedback, where human reviewers rank or correct model outputs to steer a model's behaviour toward what real users in a given market actually expect.

Corpshore AI, the AI-focused division of the Corpshore group, runs these as real service lines: data collection, annotation and labelling across image, video, text and audio, speech and audio work specifically including text-to-speech and automatic-speech-recognition training data across 35+ languages, and RLHF and preference data work for model alignment. These are not bolt-on services; they are the group's core AI data operation, delivered by dedicated data operators rather than a general BPO team asked to do data work on the side. Our AI training data service page and the broader AI services overview set out how this work is structured.

What is the actual status of te reo Maori usage in New Zealand today?

Te reo Maori usage has grown meaningfully in absolute terms over the past census cycle, though the picture is nuanced. According to Stats NZ's 2023 Census results, 213,849 people in Aotearoa New Zealand said they could hold a conversation in te reo Maori in 2023, up from 185,955 in 2018, an increase of close to 28,000 people, or 15.0 percent, making te reo Maori the most widely spoken language in the country after English. That represented 4.3 percent of the total population able to converse in the language. Regionally, Gisborne and Northland had the highest proportion of speakers, at 16.9 percent and 10.1 percent respectively.

The nuance is that growth in absolute speaker numbers has outpaced growth in the proportion of Maori specifically who speak the language, which sat at roughly 18.4 percent in 2013 and 18.6 percent in 2023, essentially static over a decade even as total speaker numbers across the whole population rose. For any organisation thinking about AI and te reo Maori, the takeaway is straightforward: there is a real, growing base of speakers, concentrated in identifiable regions, and a live public conversation in New Zealand about revitalisation, but te reo Maori remains a genuinely low-resource language by AI standards, meaning far less digital text and speech data exists for it than for English, which is exactly the gap that dedicated data collection work exists to address.

Can Corpshore build a te reo Maori dataset today?

Not an existing one, and it would be inaccurate to imply otherwise. Corpshore does not currently hold a large, ready-built te reo Maori training dataset. What Corpshore AI can genuinely offer is the operational capability to resource a te reo Maori data collection, annotation or speech-data project from scratch, using the same data collection, annotation and speech and audio workstreams the group already runs across 35+ languages, working with native speakers and appropriate community input to build a dataset that is actually representative rather than assembled from whatever scraps of digitised text happen to be available.

This is an honest but important distinction for a business evaluating AI data partners for the New Zealand market: te reo Maori data work, given the language's status and the cultural weight attached to it, should be approached deliberately, with genuine speaker and community involvement in how data is collected and labelled, not treated as a checkbox language added to an existing multilingual pipeline.

How should a business scope an AI training data project for the ANZ market?

Start with the actual use case: is the model or system aimed at general Australian-market text and speech, at a specific New Zealand audience that may include te reo Maori interactions, or both. That shapes whether the project needs primarily Australian English data (spelling, idiom, accent coverage across states and regional accents) or a genuinely dedicated te reo Maori workstream, since treating the two as interchangeable "ANZ localisation" undersells the distinctness of each.

From there, scope the specific workstream: text annotation for spelling and idiom accuracy, speech data collection across a representative spread of Australian accents for a voice product, or RLHF work to align a model's tone and register with how Australian or New Zealand users actually expect to be spoken to. Our managed AI services page covers how these workstreams get structured end to end for a client, and our AI agents and automation page is the relevant read if the end goal is deploying a model into a live customer-facing product rather than just building the underlying dataset.

Frequently asked questions

Does Corpshore already have a large te reo Maori training dataset?

No. Corpshore does not currently hold a large, ready-built te reo Maori dataset; what it offers is the operational capability, through its AI data operations, to build one from scratch with genuine speaker and community involvement.

Why can't a model trained on general English data handle Australian English well?

Because most large language models are trained predominantly on American and British English text and speech, Australian spelling, idiom and accent are underrepresented, leading to higher error rates when those models are deployed for an Australian audience.

How many people speak te reo Maori according to the latest census?

According to the 2023 New Zealand Census, 213,849 people could hold a conversation in te reo Maori, up 15.0 percent from 185,955 in 2018, making it the most widely spoken language in New Zealand after English.

What is RLHF and why does it matter for a regional market like Australia or New Zealand?

RLHF stands for reinforcement learning from human feedback, where human reviewers rank or correct a model's outputs to steer its behaviour. For a regional market, RLHF with local reviewers helps align a model's tone and cultural assumptions with what users in that specific market actually expect.

Is te reo Maori considered a low-resource language for AI purposes?

Yes. Despite growing speaker numbers, te reo Maori has far less digitised text and speech data available than English, which is the specific gap that dedicated data collection and annotation work is designed to address.

What AI data services does Corpshore actually provide today?

Corpshore AI runs data collection, annotation and labelling across image, video, text and audio, speech and audio work including text-to-speech and automatic-speech-recognition training data across 35+ languages, and RLHF and preference data work for model alignment.

Build your team with Corpshore

Tell us the work, the delivery location and the coverage you need. You will have a considered response within six hours, or book a discovery call now.

Looking for a role rather than a partner? Explore careers at Corpshore