← Back to insights

Generative search trust evaluation dataset released

Stanford HAI released an open dataset for evaluating trust in generative search answers — 8,400 labeled examples across consumer, healthcare, and financial queries with human trust scores and citation audits.

How GEO teams can use it

The dataset includes prompt templates and rubrics for scoring whether an AI answer’s brand citation is supported, hedged, or unsupported. Teams can benchmark their corpus quality before and after GEO interventions.

Early adopters in the Poptira network use the rubric in monthly reviews to prioritize fact gaps that most affect recommendation confidence.

Learn how we run monthly citation reviews.

About our process