AI Compliance Benchmark Report

ComplyAuto Guardian · Compliance Benchmark

Guardian finds the real legal issues that today’s best AI models miss

We ran real dealership ads across 17 states through Guardian and through two of the most advanced AI models available — GPT-5.5 and Claude Opus 4.8, both at maximum reasoning — then a compliance reviewer labeled every finding.

17 states · every finding human-reviewed · generated June 26, 2026

The scoreboard

How many real issues each tool catches per ad, and how often a flag is a genuine problem rather than a false alarm.

Guardian
4.4
real issues caught / ad
93%precision
7%false-positive rate
0%hallucination rate
GPT-5.5
1.7
real issues caught / ad
52%precision
48%false-positive rate
0%hallucination rate
Claude Opus 4.8
1.3
real issues caught / ad
34%precision
66%false-positive rate
3%hallucination rate
On a typical ad, today’s best AI models failed to find about 3 real legal issues that Guardian caught.
GPT-5.5 missed 2.7 real issues per ad · Claude missed 3.1 per ad
Bottom line: Guardian caught 2.6× as many real legal issues per ad as GPT-5.5 and 3.3× as many as Claude — at 93% precision and a 0% hallucination rate. And the issues Guardian flagged were the higher-risk ones the AI models missed, while the models ran a 48%–66% false-positive rate.
Real issues caught per ad

Genuine legal issues each tool surfaces on a typical ad. Higher is better.

4.4
Guardian
1.7
GPT-5.5
1.3
Claude Opus 4.8
Precision

Of everything a tool flagged, how much was a real issue — the rest is false alarms. Higher is better.

93%
Guardian
52%
GPT-5.5
34%
Claude Opus 4.8

A closer look at the findings

It’s not just how many issues each tool finds — it’s which ones.

What each tool catches per ad, by severity

Guardian’s findings are dominated by clear-cut violations and jurisdiction-specific technical rules. The models’ real findings are mostly softer “arguable” calls.

4.4
Guardian
1.7
GPT-5.5
1.3
Claude Opus 4.8
Clear violationTechnical (state-specific)Arguable

Guardian flags roughly 2.4 clear-cut violations and 1.6 state-specific technical violations per ad — the models catch a small fraction of the first and almost none of the second.

When a tool flagged a “clear violation,” how often was it wrong?

Share of each tool’s clear-cut-violation calls that turned out to be false. Higher is worse — Guardian’s most serious flags are almost always right.

10%
Guardian
47%
GPT-5.5
63%
Claude Opus 4.8

Where today’s best AI models fall short

Consistent failure patterns across every ad we tested.

1

Blind to anything behind a click

Most of what determines compliance lives behind interactive widgets, modal pop-ups, hover tooltips, expandable fine print, and finance/lease tabs — and the models missed almost all of it. To even feed that content to a general AI model, a person would have to manually open every widget and hover, then copy and paste each piece in by hand — and the model still would not know which disclosures are required. Guardian renders the live page, automatically detects what is present and what is applicable, and checks whether each disclosure actually meets clear-and-conspicuous standards.

2

Even the issues they catch are the wrong ones

It is not just that the models find far fewer issues — the ones they do surface skew toward minor, low-relevance items. They repeatedly missed the serious, high-risk violations that Guardian flagged: the findings most likely to actually draw a regulator’s attention or a consumer lawsuit. More noise, less signal.

Guardian flags roughly 2.4 clear-cut (high-risk) violations per ad — GPT-5.5 and Claude, well under one each.

3

Weak on recent and evolving law

They almost entirely missed recent regulatory developments — the FTC’s 97 warning letters, the FTC’s updated position that mandatory fees (e.g., documentary fees) must be included in the advertised price, joint NADA & FTC webinars, and state Attorney General guidance. A general AI model reflects its training cutoff, not last quarter’s enforcement. Guardian is updated to current rules.

4

Inaccurate legal citations

When the models did cite a law, the citations were unreliable — in different ways. Claude leaned on vague, blanket statutes (bare “FTC Act §5 / UDAP” or a broad consumer-protection act) on roughly half its citations, with no specific governing rule. GPT-5.5 cited precise-looking sections, but they were frequently the wrong rule — often the same templated citation pasted across unrelated findings. A citation that sounds authoritative but is wrong is worse than none: it won’t survive scrutiny and can’t be relied on for a fix. Guardian maps each finding to the specific governing statute or regulation.

5

Mishandles state-vs-federal conflicts

When state and federal requirements conflict or stack, the models applied the wrong standard or missed the stricter rule. Guardian reconciles federal and state-specific requirements per jurisdiction.

6

Over-flags nitpicks with no real risk

They frequently raised technically-true-but-trivial “issues” that carry no real legal exposure, drowning the genuine problems in noise — exactly what the false-positive rates above reflect.

7

Creates busywork on already-compliant ads

The models don’t understand common, accepted automotive-retail practices, so they flag standard, lawful conventions as violations. Acted on, that sends dealers and their staff chasing fixes that don’t need fixing — hours spent “correcting” ads that were already compliant, with no legal benefit.

8

Speed & scale

Guardian scans an entire website — hundreds of pages — in a single, continuous pass. A leading AI model with reasoning and web search enabled takes roughly five minutes per ad, one ad at a time. Auditing a full site that way is hours of hands-on model time for a single run.

Why Guardian sees what the models can’t

Guardian isn’t a general AI model with a clever prompt — it’s purpose-built for automotive advertising compliance.

Built and maintained by dealership lawyers

A team of 10 attorneys with extensive automotive-retail legal backgrounds defines and reviews the rules — not a general-purpose model guessing at the law.

Current with the latest law

Trained on the latest federal and state advertising laws and regulations, and updated constantly as new enforcement actions, guidance, and rules emerge.

Learns from millions of real scans

Continuously refined on data from millions of dealership-website scans, so it knows the patterns that actually draw enforcement — and the accepted practices that don’t.

Grounded in real-world feedback

Informed by direct feedback from regulators, state dealer-association executives, and dealership attorneys — the people who write, interpret, and enforce the rules.

A purpose-built rules engine

An advanced, jurisdiction-aware rules engine encodes genuine industry expertise and reconciles federal and state requirements — not statistical pattern-matching.

Reads the entire live page

Advanced extraction technology renders the live website and reads content behind widgets, tabs, hovers, and fine print — exactly where the disclosures that decide compliance live.

Methodology. Real dealer ads across 17 states. The same source-ad image(s) were sent to GPT-5.5 and Claude Opus 4.8 (vision) with an identical generic compliance brief (prompt v1) at maximum reasoning effort; Guardian’s column is its production result for the same ad. A compliance reviewer labeled each finding Correct / False positive / Hallucination / Inapplicable. Precision = correct ÷ labeled findings; false-positive rate = (false positive + hallucination + inapplicable) ÷ labeled; “issues per ad” counts confirmed-correct findings only and excludes any model run that errored. Recall is not formally measured, so coverage figures are a floor. Guardian analyzes the live page (including content behind widgets and tabs); the AI models analyzed the supplied screenshot, which is also what a consumer sees.

Ready to see Guardian in action?

Don’t wait any longer. Request a free demo to speak with an expert about our latest innovations.

Scroll to Top