Examples: real questions, real answers

Every number on the leaderboard is an aggregate — this page is the receipts. Below are eight real items from the corpus, each a genuine claim a regulator or court ruled on (or a matched honest control). For each one you can see the exact prompt that was sent to every model — not a paraphrase, the literal text, including the item's actual content — and every model's actual verdict on that exact prompt. Nothing here is simulated or hand-picked for effect: it's the same P0 prompt and the same graded results that feed the leaderboard, just shown per-item instead of averaged away.

ftc-2010-pom-wonderful

Deceptive FTC (US Federal Trade Commission) ↗ · 2010-09

POM Wonderful advertisement: "Clinical studies prove that POM Juice prevents, reduces the risk of, and treats heart disease, prostate cancer, and erectile dysfunction."

Candidate claims (the model must name the index of the deceptive one):

  1. 0 — Clinical studies prove that POM Juice prevents, reduces the risk of, and treats heart disease, prostate cancer, and erectile dysfunction.
  2. 1 — POM Juice.
See the exact prompt sent (P0)
Is the following piece of marketing/review content deceptive?

--- CONTENT ---
POM Wonderful advertisement: "Clinical studies prove that POM Juice prevents, reduces the risk of, and treats heart disease, prostate cancer, and erectile dysfunction."

--- CANDIDATE CLAIMS ---
0. Clinical studies prove that POM Juice prevents, reduces the risk of, and treats heart disease, prostate cancer, and erectile dysfunction.
1. POM Juice.
--- END ---

Respond with a JSON object matching this shape:
{ "verdict": "deceptive" | "honest", "claim_id": <integer index into the numbered CANDIDATE CLAIMS above, or null> }

Set "claim_id" to the index of the single candidate claim you judge most deceptive if verdict is "deceptive".
Set "claim_id" to null if verdict is "honest", or if there are no candidate claims to choose from.

Every model's actual verdict

Model Predicted verdict Predicted claim Correct?
nemotron-3-nano:4b Deceptive correct claim
cogito:32b Deceptive correct claim
glm-4.7-flash Deceptive correct claim
deepseek-r1:8b Deceptive correct claim
granite3.3:8b Deceptive correct claim
falcon3:10b Deceptive correct claim
gemma4:26b Deceptive correct claim
deepseek-r1:32b Deceptive correct claim
mistral-small3.2:24b Deceptive correct claim
qwen2.5:14b Deceptive correct claim
qwen3.5:9b Deceptive correct claim
aya-expanse:32b Deceptive correct claim
exaone3.5:32b Deceptive correct claim
qwen3.6:27b Deceptive correct claim
phi4:14b Deceptive correct claim
qwen2.5:7b Deceptive correct claim
phi4-reasoning:14b Deceptive correct claim
lfm2:24b Deceptive wrong claim
mistral-nemo:12b Deceptive correct claim
gemma2:9b Deceptive correct claim
phi4-mini:3.8b Deceptive correct claim
gemma4:12b Deceptive correct claim
gemma3n:e4b Deceptive correct claim
deepseek-r1:14b Deceptive correct claim

Sorted by each model's overall P0 accuracy (best first, alphabetical on ties) — every model with a P0 result for this item is shown, none omitted.

honest-papa-johns-puffery

Honest Honest controls (matched non-deceptive) ↗ · 2000

Papa John's advertising slogan: "Better Ingredients. Better Pizza."

See the exact prompt sent (P0)
Is the following piece of marketing/review content deceptive?

--- CONTENT ---
Papa John's advertising slogan: "Better Ingredients. Better Pizza."

--- CANDIDATE CLAIMS ---
(no candidate claims provided for this item)
--- END ---

Respond with a JSON object matching this shape:
{ "verdict": "deceptive" | "honest", "claim_id": <integer index into the numbered CANDIDATE CLAIMS above, or null> }

Set "claim_id" to the index of the single candidate claim you judge most deceptive if verdict is "deceptive".
Set "claim_id" to null if verdict is "honest", or if there are no candidate claims to choose from.

Every model's actual verdict

Model Predicted verdict Predicted claim Correct?
nemotron-3-nano:4b Honest
cogito:32b Honest
glm-4.7-flash Honest
deepseek-r1:8b Honest
granite3.3:8b Honest
falcon3:10b Honest
gemma4:26b Honest
deepseek-r1:32b Honest
mistral-small3.2:24b Honest
qwen2.5:14b Honest
qwen3.5:9b Honest
aya-expanse:32b Honest
exaone3.5:32b Honest
qwen3.6:27b Honest
phi4:14b Honest
qwen2.5:7b Honest
phi4-reasoning:14b Honest
lfm2:24b Honest
mistral-nemo:12b Honest
gemma2:9b Deceptive
phi4-mini:3.8b Honest
gemma4:12b Honest
gemma3n:e4b Honest
deepseek-r1:14b Deceptive

Sorted by each model's overall P0 accuracy (best first, alphabetical on ties) — every model with a P0 result for this item is shown, none omitted.

Matched pair: both items make a "recyclable" packaging claim, but one was ruled deceptive and the other cleared — same claim type, opposite ground truth.

asa-pepsi-100-recycled

Deceptive ASA (UK Advertising Standards Authority) ↗ · 2022-01-19

Lipton Ice Tea bottle marketing: "100% RECYCLED* bottle." Asterisked small print, set apart from the headline: "*excludes cap and label."

Candidate claims (the model must name the index of the deceptive one):

  1. 0 — 100% RECYCLED bottle
  2. 1 — excludes cap and label
See the exact prompt sent (P0)
Is the following piece of marketing/review content deceptive?

--- CONTENT ---
Lipton Ice Tea bottle marketing: "100% RECYCLED* bottle." Asterisked small print, set apart from the headline: "*excludes cap and label."

--- CANDIDATE CLAIMS ---
0. 100% RECYCLED bottle
1. excludes cap and label
--- END ---

Respond with a JSON object matching this shape:
{ "verdict": "deceptive" | "honest", "claim_id": <integer index into the numbered CANDIDATE CLAIMS above, or null> }

Set "claim_id" to the index of the single candidate claim you judge most deceptive if verdict is "deceptive".
Set "claim_id" to null if verdict is "honest", or if there are no candidate claims to choose from.

Every model's actual verdict

Model Predicted verdict Predicted claim Correct?
nemotron-3-nano:4b Deceptive correct claim
cogito:32b Deceptive correct claim
glm-4.7-flash Honest
deepseek-r1:8b Deceptive correct claim
granite3.3:8b Deceptive correct claim
falcon3:10b Deceptive correct claim
gemma4:26b Deceptive correct claim
deepseek-r1:32b Deceptive correct claim
mistral-small3.2:24b Honest
qwen2.5:14b Deceptive correct claim
qwen3.5:9b Honest
aya-expanse:32b Deceptive correct claim
exaone3.5:32b Deceptive correct claim
qwen3.6:27b Honest
phi4:14b Deceptive correct claim
qwen2.5:7b Deceptive correct claim
phi4-reasoning:14b Deceptive correct claim
lfm2:24b Deceptive wrong claim
mistral-nemo:12b Honest
gemma2:9b Deceptive correct claim
phi4-mini:3.8b Deceptive correct claim
gemma4:12b Deceptive correct claim
gemma3n:e4b Deceptive correct claim
deepseek-r1:14b Deceptive correct claim

Sorted by each model's overall P0 accuracy (best first, alphabetical on ties) — every model with a P0 result for this item is shown, none omitted.

honest-asa-lucozade-recyclable

Honest Honest controls (matched non-deceptive) ↗ · 2021-08-18

Lucozade Ribena bottle marketing: "Recyclable bottle, sleeve and cap."

See the exact prompt sent (P0)
Is the following piece of marketing/review content deceptive?

--- CONTENT ---
Lucozade Ribena bottle marketing: "Recyclable bottle, sleeve and cap."

--- CANDIDATE CLAIMS ---
(no candidate claims provided for this item)
--- END ---

Respond with a JSON object matching this shape:
{ "verdict": "deceptive" | "honest", "claim_id": <integer index into the numbered CANDIDATE CLAIMS above, or null> }

Set "claim_id" to the index of the single candidate claim you judge most deceptive if verdict is "deceptive".
Set "claim_id" to null if verdict is "honest", or if there are no candidate claims to choose from.

Every model's actual verdict

Model Predicted verdict Predicted claim Correct?
nemotron-3-nano:4b Honest
cogito:32b Honest
glm-4.7-flash Honest
deepseek-r1:8b Honest
granite3.3:8b Honest
falcon3:10b Honest
gemma4:26b Honest
deepseek-r1:32b Honest
mistral-small3.2:24b Honest
qwen2.5:14b Honest
qwen3.5:9b Honest
aya-expanse:32b Honest
exaone3.5:32b Honest
qwen3.6:27b Honest
phi4:14b Honest
qwen2.5:7b Honest
phi4-reasoning:14b Honest
lfm2:24b Honest
mistral-nemo:12b Honest
gemma2:9b Deceptive
phi4-mini:3.8b Honest
gemma4:12b Honest
gemma3n:e4b Honest
deepseek-r1:14b Deceptive

Sorted by each model's overall P0 accuracy (best first, alphabetical on ties) — every model with a P0 result for this item is shown, none omitted.

Matched pair: both items are the same retailer's Black-Friday "was/now" pricing campaign, but one was ruled deceptive and the other cleared.

asa-boots-boss-fragrance

Deceptive ASA (UK Advertising Standards Authority) ↗ · 2026-05-20

Boots Black Friday ad: "BOSS Bottled Eau de Parfum — was £80, now only £60."

Candidate claims (the model must name the index of the deceptive one):

  1. 0 — was £80, now only £60
  2. 1 — BOSS Bottled Eau de Parfum
See the exact prompt sent (P0)
Is the following piece of marketing/review content deceptive?

--- CONTENT ---
Boots Black Friday ad: "BOSS Bottled Eau de Parfum — was £80, now only £60."

--- CANDIDATE CLAIMS ---
0. was £80, now only £60
1. BOSS Bottled Eau de Parfum
--- END ---

Respond with a JSON object matching this shape:
{ "verdict": "deceptive" | "honest", "claim_id": <integer index into the numbered CANDIDATE CLAIMS above, or null> }

Set "claim_id" to the index of the single candidate claim you judge most deceptive if verdict is "deceptive".
Set "claim_id" to null if verdict is "honest", or if there are no candidate claims to choose from.

Every model's actual verdict

Model Predicted verdict Predicted claim Correct?
nemotron-3-nano:4b Deceptive correct claim
cogito:32b Honest
glm-4.7-flash Honest
deepseek-r1:8b Honest
granite3.3:8b Honest
falcon3:10b Honest
gemma4:26b Honest
deepseek-r1:32b Honest
mistral-small3.2:24b Honest
qwen2.5:14b Honest
qwen3.5:9b Honest
aya-expanse:32b Honest
exaone3.5:32b Honest
qwen3.6:27b Honest
phi4:14b Honest
qwen2.5:7b Honest
phi4-reasoning:14b Honest
lfm2:24b Deceptive wrong claim
mistral-nemo:12b Honest
gemma2:9b Deceptive correct claim
phi4-mini:3.8b Honest
gemma4:12b Honest
gemma3n:e4b Deceptive correct claim
deepseek-r1:14b Deceptive correct claim

Sorted by each model's overall P0 accuracy (best first, alphabetical on ties) — every model with a P0 result for this item is shown, none omitted.

honest-asa-boots-black-friday

Honest Honest controls (matched non-deceptive) ↗ · 2026-05-20

Boots Black Friday ad: "Boots' biggest ever Black Friday is here with better than half price star gifts. Including this Yankee Candle set, was £60, now only £29.50."

See the exact prompt sent (P0)
Is the following piece of marketing/review content deceptive?

--- CONTENT ---
Boots Black Friday ad: "Boots' biggest ever Black Friday is here with better than half price star gifts. Including this Yankee Candle set, was £60, now only £29.50."

--- CANDIDATE CLAIMS ---
(no candidate claims provided for this item)
--- END ---

Respond with a JSON object matching this shape:
{ "verdict": "deceptive" | "honest", "claim_id": <integer index into the numbered CANDIDATE CLAIMS above, or null> }

Set "claim_id" to the index of the single candidate claim you judge most deceptive if verdict is "deceptive".
Set "claim_id" to null if verdict is "honest", or if there are no candidate claims to choose from.

Every model's actual verdict

Model Predicted verdict Predicted claim Correct?
nemotron-3-nano:4b Honest
cogito:32b Honest
glm-4.7-flash Honest
deepseek-r1:8b Honest
granite3.3:8b Honest
falcon3:10b Honest
gemma4:26b Honest
deepseek-r1:32b Honest
mistral-small3.2:24b Honest
qwen2.5:14b Honest
qwen3.5:9b Honest
aya-expanse:32b Honest
exaone3.5:32b Honest
qwen3.6:27b Honest
phi4:14b Honest
qwen2.5:7b Deceptive
phi4-reasoning:14b Honest
lfm2:24b Deceptive
mistral-nemo:12b Honest
gemma2:9b Honest
phi4-mini:3.8b Honest
gemma4:12b Honest
gemma3n:e4b Deceptive
deepseek-r1:14b Deceptive

Sorted by each model's overall P0 accuracy (best first, alphabetical on ties) — every model with a P0 result for this item is shown, none omitted.

sec-2021-bitconnect

Deceptive SEC (US Securities and Exchange Commission) ↗ · 2021-09

BitConnect lending program: "Earn average daily returns of 1% with our Volatility Software Trading Bot — up to 40% per month."

Candidate claims (the model must name the index of the deceptive one):

  1. 0 — Earn average daily returns of 1% ... up to 40% per month
  2. 1 — Volatility Software Trading Bot
See the exact prompt sent (P0)
Is the following piece of marketing/review content deceptive?

--- CONTENT ---
BitConnect lending program: "Earn average daily returns of 1% with our Volatility Software Trading Bot — up to 40% per month."

--- CANDIDATE CLAIMS ---
0. Earn average daily returns of 1% ... up to 40% per month
1. Volatility Software Trading Bot
--- END ---

Respond with a JSON object matching this shape:
{ "verdict": "deceptive" | "honest", "claim_id": <integer index into the numbered CANDIDATE CLAIMS above, or null> }

Set "claim_id" to the index of the single candidate claim you judge most deceptive if verdict is "deceptive".
Set "claim_id" to null if verdict is "honest", or if there are no candidate claims to choose from.

Every model's actual verdict

Model Predicted verdict Predicted claim Correct?
nemotron-3-nano:4b Deceptive correct claim
cogito:32b Deceptive correct claim
glm-4.7-flash Deceptive correct claim
deepseek-r1:8b Deceptive correct claim
granite3.3:8b Deceptive correct claim
falcon3:10b Deceptive correct claim
gemma4:26b Deceptive correct claim
deepseek-r1:32b Deceptive correct claim
mistral-small3.2:24b Deceptive correct claim
qwen2.5:14b Deceptive correct claim
qwen3.5:9b Deceptive correct claim
aya-expanse:32b Deceptive correct claim
exaone3.5:32b Deceptive correct claim
qwen3.6:27b Deceptive correct claim
phi4:14b Deceptive correct claim
qwen2.5:7b Deceptive correct claim
phi4-reasoning:14b Deceptive correct claim
lfm2:24b Deceptive wrong claim
mistral-nemo:12b Deceptive correct claim
gemma2:9b Deceptive correct claim
phi4-mini:3.8b Deceptive wrong claim
gemma4:12b Deceptive correct claim
gemma3n:e4b Deceptive correct claim
deepseek-r1:14b Deceptive correct claim

Sorted by each model's overall P0 accuracy (best first, alphabetical on ties) — every model with a P0 result for this item is shown, none omitted.

honest-nad-olly-libido

Honest Honest controls (matched non-deceptive) ↗ · 2025-07

Olly "Lovin' Libido" supplement: "Ashwagandha to support desire, drive, arousal, satisfaction and lubrication."

See the exact prompt sent (P0)
Is the following piece of marketing/review content deceptive?

--- CONTENT ---
Olly "Lovin' Libido" supplement: "Ashwagandha to support desire, drive, arousal, satisfaction and lubrication."

--- CANDIDATE CLAIMS ---
(no candidate claims provided for this item)
--- END ---

Respond with a JSON object matching this shape:
{ "verdict": "deceptive" | "honest", "claim_id": <integer index into the numbered CANDIDATE CLAIMS above, or null> }

Set "claim_id" to the index of the single candidate claim you judge most deceptive if verdict is "deceptive".
Set "claim_id" to null if verdict is "honest", or if there are no candidate claims to choose from.

Every model's actual verdict

Model Predicted verdict Predicted claim Correct?
nemotron-3-nano:4b Honest
cogito:32b Deceptive
glm-4.7-flash Honest
deepseek-r1:8b Deceptive
granite3.3:8b Honest
falcon3:10b Honest
gemma4:26b Honest
deepseek-r1:32b Deceptive
mistral-small3.2:24b Honest
qwen2.5:14b Deceptive
qwen3.5:9b Honest
aya-expanse:32b Honest
exaone3.5:32b Honest
qwen3.6:27b Honest
phi4:14b Deceptive
qwen2.5:7b Deceptive
phi4-reasoning:14b Honest
lfm2:24b Deceptive
mistral-nemo:12b Honest
gemma2:9b Deceptive
phi4-mini:3.8b Honest
gemma4:12b Honest
gemma3n:e4b Deceptive
deepseek-r1:14b Deceptive

Sorted by each model's overall P0 accuracy (best first, alphabetical on ties) — every model with a P0 result for this item is shown, none omitted.