Guardian finds the real legal issues that today’s best AI models miss
We ran real dealership ads across 17 states through Guardian and through two of the most advanced AI models available — GPT-5.5 and Claude Opus 4.8, both at maximum reasoning — then a compliance reviewer labeled every finding.
The scoreboard
How many real issues each tool catches per ad, and how often a flag is a genuine problem rather than a false alarm.
Genuine legal issues each tool surfaces on a typical ad. Higher is better.
Of everything a tool flagged, how much was a real issue — the rest is false alarms. Higher is better.
A closer look at the findings
It’s not just how many issues each tool finds — it’s which ones.
Guardian’s findings are dominated by clear-cut violations and jurisdiction-specific technical rules. The models’ real findings are mostly softer “arguable” calls.
Guardian flags roughly 2.4 clear-cut violations and 1.6 state-specific technical violations per ad — the models catch a small fraction of the first and almost none of the second.
Share of each tool’s clear-cut-violation calls that turned out to be false. Higher is worse — Guardian’s most serious flags are almost always right.
Where today’s best AI models fall short
Consistent failure patterns across every ad we tested.
Blind to anything behind a click
Most of what determines compliance lives behind interactive widgets, modal pop-ups, hover tooltips, expandable fine print, and finance/lease tabs — and the models missed almost all of it. To even feed that content to a general AI model, a person would have to manually open every widget and hover, then copy and paste each piece in by hand — and the model still would not know which disclosures are required. Guardian renders the live page, automatically detects what is present and what is applicable, and checks whether each disclosure actually meets clear-and-conspicuous standards.
Even the issues they catch are the wrong ones
It is not just that the models find far fewer issues — the ones they do surface skew toward minor, low-relevance items. They repeatedly missed the serious, high-risk violations that Guardian flagged: the findings most likely to actually draw a regulator’s attention or a consumer lawsuit. More noise, less signal.
Guardian flags roughly 2.4 clear-cut (high-risk) violations per ad — GPT-5.5 and Claude, well under one each.
Weak on recent and evolving law
They almost entirely missed recent regulatory developments — the FTC’s 97 warning letters, the FTC’s updated position that mandatory fees (e.g., documentary fees) must be included in the advertised price, joint NADA & FTC webinars, and state Attorney General guidance. A general AI model reflects its training cutoff, not last quarter’s enforcement. Guardian is updated to current rules.
Inaccurate legal citations
When the models did cite a law, the citations were unreliable — in different ways. Claude leaned on vague, blanket statutes (bare “FTC Act §5 / UDAP” or a broad consumer-protection act) on roughly half its citations, with no specific governing rule. GPT-5.5 cited precise-looking sections, but they were frequently the wrong rule — often the same templated citation pasted across unrelated findings. A citation that sounds authoritative but is wrong is worse than none: it won’t survive scrutiny and can’t be relied on for a fix. Guardian maps each finding to the specific governing statute or regulation.
Mishandles state-vs-federal conflicts
When state and federal requirements conflict or stack, the models applied the wrong standard or missed the stricter rule. Guardian reconciles federal and state-specific requirements per jurisdiction.
Over-flags nitpicks with no real risk
They frequently raised technically-true-but-trivial “issues” that carry no real legal exposure, drowning the genuine problems in noise — exactly what the false-positive rates above reflect.
Creates busywork on already-compliant ads
The models don’t understand common, accepted automotive-retail practices, so they flag standard, lawful conventions as violations. Acted on, that sends dealers and their staff chasing fixes that don’t need fixing — hours spent “correcting” ads that were already compliant, with no legal benefit.
Speed & scale
Guardian scans an entire website — hundreds of pages — in a single, continuous pass. A leading AI model with reasoning and web search enabled takes roughly five minutes per ad, one ad at a time. Auditing a full site that way is hours of hands-on model time for a single run.
Why Guardian sees what the models can’t
Guardian isn’t a general AI model with a clever prompt — it’s purpose-built for automotive advertising compliance.
Built and maintained by dealership lawyers
A team of 10 attorneys with extensive automotive-retail legal backgrounds defines and reviews the rules — not a general-purpose model guessing at the law.
Current with the latest law
Trained on the latest federal and state advertising laws and regulations, and updated constantly as new enforcement actions, guidance, and rules emerge.
Learns from millions of real scans
Continuously refined on data from millions of dealership-website scans, so it knows the patterns that actually draw enforcement — and the accepted practices that don’t.
Grounded in real-world feedback
Informed by direct feedback from regulators, state dealer-association executives, and dealership attorneys — the people who write, interpret, and enforce the rules.
A purpose-built rules engine
An advanced, jurisdiction-aware rules engine encodes genuine industry expertise and reconciles federal and state requirements — not statistical pattern-matching.
Reads the entire live page
Advanced extraction technology renders the live website and reads content behind widgets, tabs, hovers, and fine print — exactly where the disclosures that decide compliance live.
Methodology. Real dealer ads across 17 states. The same source-ad image(s) were sent to GPT-5.5 and Claude Opus 4.8 (vision) with an identical generic compliance brief (prompt v1) at maximum reasoning effort; Guardian’s column is its production result for the same ad. A compliance reviewer labeled each finding Correct / False positive / Hallucination / Inapplicable. Precision = correct ÷ labeled findings; false-positive rate = (false positive + hallucination + inapplicable) ÷ labeled; “issues per ad” counts confirmed-correct findings only and excludes any model run that errored. Recall is not formally measured, so coverage figures are a floor. Guardian analyzes the live page (including content behind widgets and tabs); the AI models analyzed the supplied screenshot, which is also what a consumer sees.