Some campaigns see 90% of claims resolved by AI, without a person touching them. Across a typical campaign, the average sits nearer 80%. That’s the headline, already established in this series. This isn’t a victory lap, it’s an honest account of what actually got a claims process there: the design decisions that worked, the ones that didn’t at first, and where a human is still needed.

Key Takeaways

  • Automation rates reach 90% on some campaigns and average roughly 80% overall, depending on claim complexity and volume, a figure already established in Blogs 01 and 02
  • AI-reviewed claims are resolved fast: internal data shows a P95 completion time of roughly 15 seconds, about four times faster than the OCR-based process it replaced
  • Prompt engineering is treated as an ongoing discipline, not a one-time setup, with continuous iteration between the team writing prompts and the team building the system
  • OPIA re-evaluates its AI models on roughly a six-month cycle, balancing the effort of testing and validating a new model against the improvement it actually delivers
  • AI use is disclosed and governed differently by territory, some markets require customer-facing disclosure that AI is involved, others set specific rules on human oversight of customer-impacting decisions

Where We Started

This is the fourth piece in a series, not a standalone claim. our blog on AI in sales promotions told the origin story, hack days, the Zero Touch squad, and the model-agnostic decision behind OPIA’s approach to AI. This piece picks up from there: not why OPIA adopted AI, but specifically what it took to get automation rates this high, and what changed to get there.

What “90% Automation” Actually Means

This shape is already established: claims that fall outside clear thresholds, whether due to ambiguity, missing data, or fraud signals, are automatically escalated to human review, and the ratio of automated to human-reviewed claims shifts based on campaign type, risk profile, and client requirements, as covered in how OPIA deploys AI across its operational model. That’s the honest framing already live on the site, and this section should build on it rather than repeat it from scratch.

The biggest factor is campaign and validation complexity, rather than claim volume alone. Straightforward campaigns with simple eligibility rules can achieve the highest Zero Touch rates. Automation tends to reduce where customers must provide multiple documents, information needs to be compared across documents, or submission requirements are more complex.

The Three Things That Actually Got Us There

1. Prompt engineering as an operational discipline

Getting a reliable answer out of an AI model isn’t a one-time setup, it’s an ongoing practice. Clear system instructions give the model context on what it’s actually being asked to do, and getting that instruction right took real time and iteration, not a first-draft prompt that worked out of the box.

One specific realization shaped the approach early on: telling the model that only two answers, true or false, are valid can actually push it toward guessing rather than admitting uncertainty. Adding a third option, “unknown,” as a genuinely valid response improved the results, because it gave the model a way to flag a case it wasn’t confident about instead of forcing an answer either way.

This depends on close, continuous collaboration between the team writing the prompts and the team building the underlying system: flagging where a prompt isn’t producing the desired response, adjusting it, testing again. That loop, not a single well-crafted prompt, is what drives the iterative improvement behind the numbers in this piece.

“Getting to high automation isn’t about writing one perfect prompt. It’s continuous optimization, understanding why claims fall out of Zero Touch, refining the rules and prompts, testing the change and measuring whether it actually improves the customer journey without compromising validation accuracy.”

— Nina Kunaver, Automation Lead

The production validation instructions stay internal, publishing them would hand fraudsters a map of exactly what the system checks for. What can be shared is the structure behind every decision: each requirement resolves to one of three outcomes.

  • True – the requirement is met
  • False – the requirement is not met
  • Unknown – the system can’t reliably determine it, so it doesn’t force a decision, protecting accuracy over speed

2. Binary precision over probabilistic comfort

This is territory our blog on AI in sales promotions already covers in depth, the shift to deterministic yes/no/unknown decisions, and why a confidence score doesn’t actually resolve a claim. Rather than re-explain it here, this section should reference it directly and add only what’s specific to the automation-rate story, how that design choice is what makes a 90%-on-some-campaigns figure trustworthy rather than a rounding trick.

An unknown, or otherwise non-actionable, result intentionally reduces the automated rate, because that claim is routed away from Zero Touch rather than forced through. This is by design: the objective was never to maximize automation at any cost, it’s to automate claims where the system genuinely has enough information to decide, and escalate the ones it doesn’t. That’s what makes the automation percentage a meaningful number rather than an inflated one.

3. Dual-run testing

Dual-run was a real, distinct part of how this was built, not another name for dry-run or bulk-run testing. When the AI system was first introduced, it ran alongside the existing OCR process rather than replacing it outright, the AI resolved claims first, with OCR available as a fallback while the new system’s performance was being measured. Only once that comparison held up did OCR step back to a pure fallback role.

Dry-run and bulk-run are separate capabilities that exist alongside this: dry-run tests a change before it reaches live claims, and bulk-run tests a policy or prompt change against a larger set of real claims at once, useful for catching whether fixing one scenario has broken another. All three matter, but they answer different questions, and none of them should be described as another name for the others.

What Changed on the Team

The team running claim validation has effectively retrained as prompt engineers, guided by close collaboration with the development team. That’s a confirmed, real shift, not everyone doing the same job with a new tool, but a genuine change in what the role requires day to day.

Automation has changed where the team’s time and expertise go. Rather than every claim following the same manual process, a significant proportion is now resolved through Zero Touch, freeing the operational team to focus on the claims that genuinely need human judgment. 

It’s also created a more specialized automation function in its own right: instead of just maintaining validation rules, this team analyzes why claims fall out of Zero Touch, spots recurring patterns, and tests fixes at scale before rolling them out. That creates a continuous feedback loop between the automation team, operations, and development, with real claim behavior deciding what gets improved next.

What This Means for Clients

  • Faster payouts: claims resolved in roughly 15 seconds at the 95th percentile, versus a manual or OCR-based process measured in minutes or longer
  • More consistent decisions: a deterministic yes/no/unknown design applies the same standard to every claim, rather than judgment varying between reviewers or across a busy day
  • Faster escalation when review is genuinely needed: because the system is built to flag uncertainty rather than guess, claims that need a human get there quickly, not after a delayed, low-confidence automated attempt

A Note on Staying Current

AI models change faster than most infrastructure, and foundation models can be deprecated with very little notice. OPIA’s model-agnostic approach means the platform isn’t built around a single AI model or provider: different models can be tested and evaluated against real claim scenarios, and better-performing technology can be adopted as it develops without redesigning the whole validation process. OPIA reviews new models on roughly a six-month cycle, weighing the effort of testing and validating a new one against the real improvement it delivers, rather than switching every time a new version ships.

Staying Aligned with Regulation

AI-assisted decisions aren’t governed the same way everywhere. Some territories require customer communications to disclose that AI was involved in a decision. Others set specific rules on whether AI can be used to approve or reject a customer-impacting outcome at all, or on how much human oversight (human-in-the-loop) is required before a decision is final. Navigating this by market, not applying one global standard, is part of what the operational model has to account for.

FAQs

How fast is an AI-reviewed claim decision at OPIA?

Internal data shows a 95th-percentile completion time of roughly 15 seconds, about four times faster than the OCR-based process it replaced.

What happens when the AI isn’t confident about a claim?

The system is designed to return “unknown” rather than force a low-confidence guess, which routes that claim to a human reviewer instead of resolving it incorrectly.

How often does OPIA update the AI models behind claim automation?

Roughly every six months, balancing the effort of testing and validating a new model against the real improvement it delivers, rather than adopting every new release.

What percentage of claims does OPIA automate?

Automation rates reach roughly 90% on some campaigns and average around 80% overall, depending on claim complexity and volume.

What determines whether a campaign sits closer to 80% or 90% automation?

Mainly the complexity of the campaign’s validation requirements and how customers submit their proof of purchase. Campaigns with straightforward eligibility rules and clear, consistent documentation tend to reach the highest automation rates. Requirements involving multiple documents, more complex validation rules, image-based checks, or genuinely ambiguous judgment calls are more likely to need a human.

Related Posts