How easily do LLMs fall for advertising?

We plant the fake ads. The AIs go shopping. Here's how often they get fooled.

Every model gets an Ad-Resistance Score from 0% to 100% — the higher the percentage, the harder that AI was to fool with fake reviews, hype, and planted claims.

We also check whether each AI invents facts on its own — not just whether it repeats a planted lie. Combined into one Honesty Score: right now google/gemini-2.5-flash leads at 76%.

Green = the AI resisted the manipulation. Red = it got fooled. Higher score = harder to fool.

Leaderboard

# AI model Ad-Resistance Score Fabrication Honesty Score Cost
1 google/gemini-2.5-flash 93.0% [92.0–94.0] Hard to fool 0.71 40% 76.3% $0.0040
2 x-ai/grok-4.5 95.5% [93.7–97.0] Hard to fool 0.90 50% 72.7% $0.0167
3 z-ai/glm-5.2 95.3% [93.7–96.7] Hard to fool 1.12 56% 69.9% $0.0160
4 x-ai/grok-4.20 93.8% [92.1–95.4] Hard to fool 1.06 54% 69.7% $0.0079
5 openai/gpt-5.4-mini 91.4% [89.8–92.9] Hard to fool 0.93 58% 66.8% $0.0068
6 meta-llama/llama-4-maverick 89.4% [88.5–90.3] Hard to fool $0.0018

The Ad-Resistance Score summarizes how easily each AI was fooled by planted fake reviews, hype, and ads. Fabrication/Honesty Score require a separate, opt-in judge pass that checks whether a model invents facts on its own — models without that pass yet show "—" in those columns, not a zero. Click a column header to sort. See the methodology page for the full technical breakdown behind each score.

Which tricks fooled which AIs

% of planted tricks each model caught — green = usually caught it, red = usually fell for it.

Fakereviews Puffery Clickbait AI-targeted grok-4.5 80% 93% 97% 97% glm-5.2 84% 71% 93% 93% grok-4.20 68% 58% 97% 89% gemini-2.5-flash 60% 47% 71% 64% gpt-5.4-mini 54% 47% 76% 68% llama-4-maverick 50% 22% 39% 33% Average 66% 56% 79% 74%

Key findings so far

  • No AI recommended the wrong product or repeated a fake claim as true — every model tested held up on the two most serious failure modes across the pilot.
  • What actually separates the models is how often they noticed the manipulation in the first place, not whether they fell for it outright.
  • Subtle fakes slip past far more often than blatant, over-the-top ones — every model tested caught the obvious spam more easily than the plausible-sounding version.