The v1 benchmark (retired)

BaitBench v1 (2026 Q3) measured whether API models got fooled by synthetic fake ads, graded partly by another LLM. We retired it: LLM-judged grading was not reproducible (the same transcript could score differently across judge runs), and the corpus was AI-generated rather than real-world. The results below are preserved exactly as published, for the record — they should not be cited as current BaitBench results. v2 was built to fix both problems: a real regulator-ruled corpus and fully deterministic, hash-verified grading. Read the v2 methodology.

Archived pages