Skip to content
Corpshore Australia

Case study

Red-teaming and evaluation ahead of a Sydney AI startup's public release

An early-growth Sydney AI startup building a consumer-facing LLM product, engaging Corpshore to red-team and evaluate its model outputs ahead of a wider public release.

The challenge

The startup's product team had been testing the model on an ad hoc basis. Whoever had a spare hour would try to break it, note a finding in a shared spreadsheet, and move on to the next sprint. There was no consistent rubric, no structured adversarial testing plan, and no reliable way to confirm whether a failure mode found in one sprint had actually been fixed by the next. With a public launch date locked in, the founders needed a repeatable evaluation process that could catch problems before customers did.

What Corpshore did

The engagement opened with a lead evaluator and five reviewers running a two-week calibration sprint against the startup's existing test set. Once the rubric stabilised, the panel grew to twelve reviewers plus a part-time lead who checked inter-rater agreement weekly. In the eight weeks before launch, the panel briefly scaled to eighteen people to clear a backlog of edge cases the product team kept raising. After launch it settled back to a standing panel of six for ongoing monitoring. The panel combined structured red-teaming, adversarial prompts designed to surface unsafe, biased or simply wrong outputs, with output evaluation against a rubric built jointly with the startup in the first fortnight.

Delivery model

Onshore delivery from Sydney, hybrid, a human evaluation panel working partly from the startup's own office and partly remote.

Compliance handling

Reviewers worked with real prompts and outputs that occasionally included personal information typed in by test users, so Corpshore applied Privacy Act 1988 (Cth) and Australian Privacy Principles handling standards to every review session, including access controls on who could see which test sets and a practice of stripping personal information from shared rubric documents wherever practical. Because delivery was entirely onshore, no cross-border disclosure question under APP 8 arose for this engagement. Corpshore's security practices for the panel's tooling and access management are designed to align with the Essential Eight, the ACSC's prioritised mitigation framework, rather than an achieved certification or maturity level, and the startup's own security lead reviewed the arrangement before sign-off.

Results

289 / 340
Failure modes fixed before launch
~40%
Cheaper than an equivalent in-house evaluation panel
52%
Fewer harmful or clearly wrong outputs in month one post-launch

Over the eight-week pre-release push, the panel logged 340 distinct failure modes, of which the product team fixed 289 before launch. Average time from a failure mode being logged to a verified fix fell from around eleven days of ad hoc testing to under 48 hours once the structured cycle was running. The startup estimated that building an equivalent panel in-house, including recruitment and management overhead, would have cost roughly 40 percent more over the same period. In the first month after launch, user-reported harmful or clearly wrong outputs were down 52 percent against the startup's own baseline from its previous release.

Time from a logged failure mode to a verified fix

Time from a logged failure mode to a verified fix, before and after in days.
StageValue
Before (ad hoc)11 days
After (structured cycle)2 days

After figure converts under 48 hours.

Having reviewers who understood our product well enough to catch subtle failures, and who sat close enough to our team to argue about edge cases in real time, changed how confident we felt walking into launch.

Co-founder and Head of Product, AI startup

Why it worked

A shared rubric and weekly inter-rater agreement checks turned scattered ad hoc testing into a process the founders could actually trust before a launch date they could not move.

Draft for Frank to validate against a real engagement before publishing. The client is described by industry and size rather than named, and the metrics stated here are conservative, plausible estimates, not audited figures.

Build your team with Corpshore

Tell us the work, the delivery location and the coverage you need. You will have a considered response within six hours, or book a discovery call now.

Looking for a role rather than a partner? Explore careers at Corpshore