AI delivery
AI Evaluation Specialist, Model Output Review
About the role
You will review and score AI model outputs for clients working with Corpshore AI, the AI data division of Corpshore Solutions Corporation, against a structured evaluation rubric that typically covers accuracy, safety, tone and factuality, with the exact weighting set per project. Some of that work is straightforward output review. Some of it is structured red-teaming: deliberately adversarial testing designed to surface a model's failure modes, its blind spots, its unsafe responses under pressure, before a client puts the product in front of real users. Clients specifically want this role filled from New Zealand because they want New Zealand English and Pacific-region cultural context represented in the evaluation panel, judgement that a general English-speaking reviewer elsewhere would not reliably bring.
This is a genuinely flexible contractor engagement, worked from wherever you are in New Zealand, structured around a minimum of twenty hours a week rather than fixed daily shifts.
What you will do
You will score model outputs against the rubric supplied for each active project, which will vary between straightforward factual accuracy checks and more judgement-heavy assessments of tone, helpfulness or potential harm. You will carry out structured red-teaming tasks when assigned, working from adversarial prompt sets designed to find where a model breaks down, and documenting exactly what failed and why in language a client's safety team can act on. You will apply New Zealand English usage and Pacific-region cultural context specifically where a project calls for it, catching the kind of misfire a reviewer without that background would simply not notice. You will document your reasoning on ambiguous or borderline calls rather than just submitting a score, since the written rationale is often what a client's team actually reads. You will take part in calibration sessions with the wider evaluation team so that your scoring stays consistent with everyone else's on the same rubric, and you will flag guideline gaps that produce inconsistent results rather than quietly working around them.
What you bring
Native or near-native New Zealand English, written and spoken. Sound, defensible judgement on ambiguous or borderline content, and the ability to explain a decision in writing rather than just make it. A reliable computer and internet connection. The right to work in New Zealand as a contractor. A level head around adversarial or occasionally unsettling content, since red-teaming work by its nature involves deliberately provoking a model's worst responses so they can be fixed before launch.
Nice to have
A post-secondary qualification in any field, particularly linguistics, communications, psychology or a related discipline. Prior experience in content moderation, quality assurance, editorial review, teaching or research. Familiarity with Pacific languages or Pacific community contexts beyond general New Zealand cultural fluency. Any prior exposure to AI safety, trust and safety, or red-teaming work specifically.
What we offer
This is a flexible contractor engagement, not a salaried employee position, and we would rather say that plainly than imply otherwise. What is genuinely on offer: an hourly rate that reflects the scarcity of New Zealand-specific evaluation capability, real flexibility over your working hours around the weekly minimum, and work that sits at the sharp end of frontier AI safety rather than routine review. Strong, consistent evaluators have a clear route to more hours as project volume grows, and from there into a team lead role coordinating other reviewers on the same account.
How to apply
Our process
- 1. Our talent team reviews every application against the role requirements.
- 2. Shortlisted candidates are invited to interview, which may include a role-related assessment.
- 3. We share the decision with everyone we interview, whether or not the application progresses.
Most decisions follow within a few weeks of the closing date. Urgent roles are prioritised.