Benchmark · Methodology Paper · Working Draft v0.6

Consumer Research Bench (CRB)

Evaluating AI agents on consumer research tasks: signal discovery, evidence grounding, methodology fidelity, and decision-ready synthesis.

v0.1 Scope: Consumer domain·Status: Methodology published, empirical results pending

0 stars ·0 watching ·0 forks ·Version v0.6 ·12 briefs / 30–40 atomic items ·Results To Run

Abstract

Consumer research combines evidence collection, behavioral interpretation, qualitative synthesis, quantitative classification, and business recommendation. Existing AI benchmarks evaluate web research agents, long-form research agents, browsing agents, coding agents, and general-purpose assistants, but do not directly evaluate the capabilities specific to consumer research.

Consumer Research Bench evaluates whether AI systems can transform fragmented consumer-signal data into evidence-backed consumer understanding with the interpretive depth associated with primary research and the speed of digital signal analysis.

CRB v0.1 focuses on the Consumer domain and evaluates two tracks: an Objective Track for atomic research tasks and a Report Track for open-ended consumer research synthesis. The benchmark introduces RetroSignals, a frozen consumer-signal environment, and uses a data × model ablation to separate model quality from consumer-data advantage.

Key Contributions

C1 Consumer Research Construct
Decision-ready consumer-research quality per unit time, at a fixed evidence-grounding bar.
C2 Two Evaluation Tracks
Objective atomic tasks plus open-ended report evaluation.
C3 RetroSignals
A frozen consumer-signal environment for reproducible evaluation.
C4 C-RACE
Reference-based, criteria-adaptive report scoring validated against human experts.
C5 Data × Model Ablation
Separates the contribution of consumer-signal data from model capability.
C6 Validity Gates
Hard failure rules for fabricated quotes, unsupported statistics, and untraceable evidence.

Benchmark Structure

A research brief flows through evaluation to a scored, diagnostic output.

Input

  • Research brief
  • Business context
  • Methodology
  • RetroSignals environment

Evaluation

  • Objective Track
  • Report Track
  • Evidence grounding
  • Human validation
  • Data × model ablation

Output

  • Scores
  • Diagnostics
  • Validity gates
  • Leaderboard
  • Benchmark report

Evaluation Tracks

Track A

Objective Track

Scores atomic research capabilities against expert-worked ground truth.

Example tasks
Find SignalDerive MetricValidate Claim Define CohortPopulate SetCompile Dataset
Metrics
Binary accuracyRecall / F1 Probability differenceExpert-judged match
Track B

Report Track

Scores open-ended consumer research outputs using C-RACE.

Scoring dimensions
  • D1 Comprehensiveness
  • D2 Consumer Insight Depth
  • D3 Methodology Fidelity
  • D4 Evidence & Traceability
  • D5 Quant–Qual Integration
  • D6 Contextual Fidelity
  • D7 Decision Readiness
  • D8 Readability & Structure

RetroSignals

RetroSignals is a frozen, task-linked consumer-signal environment designed to make consumer research agent evaluation reproducible over time.

Social postsComments & repliesReviews Marketplace contentCommunity discussionsSearch-result pages Creator contentVideo / audio transcriptsBrand & competitor context docs
Data Breadth
Coverage across signal types and sources.
Data Depth
Granularity and richness within each source.
Data Condition
Noise, recency, and integrity of the signal.
Relevant Signal Density
Share of signal that bears on the brief.

Validity Gates

Some failures do not merely reduce score. They invalidate a run or cap its maximum score.

Invalidates the run
  • Fabricated consumer quotes
  • Fabricated statistics
  • Non-existent cited evidence
  • No source-level traceability for central claims
Caps the maximum score
  • Confusing brand-owned content with consumer signal
  • Treating news or investor commentary as consumer behavior
  • Overgeneralizing from one viral post
  • Ignoring contradictory evidence
  • Producing generic insights
  • Recommending actions without evidence
  • Failing to follow the intended methodology

Example Task

Report Track · Sample Brief

Why are young urban consumers choosing matcha over coffee?

Required output

  • 01 Executive summary
  • 02 Key motivations
  • 03 Consumption occasions
  • 04 Consumer cohorts
  • 05 Barriers & tensions
  • 06 Evidence table
  • 07 Opportunity spaces
  • 08 Brand implications
  • 09 Confidence assessment

Strong answer characteristics

A strong answer goes beyond “matcha is healthy and trendy.” It identifies specific patterns — each grounded in inspectable signal:

Clean energyMid-afternoon productivityAesthetic self-expression Coffee replacementWellness signalingAt-home ritualization PremiumizationCreator influenceTaste / price / authenticity barriers

Leaderboard

Report-only leaderboard: Consuma, Claude, Gemini and ChatGPT scored on comprehensiveness, insight depth, method fidelity and evidence.
Report-only leaderboard (CRB-RO). Four AI consumer-research tools on one brief across four weighted dimensions. Column height is the overall score; the blue cap is consumer-signal evidence — the discriminator.
Social-listening comparison: Consuma versus Winnin, Sprinklr and iGenie across four consumer-data dimensions.
Social listening. Consuma compared with social-listening tools across four consumer-data dimensions — data universe, signal-to-noise, AI insights, and measurement & enrichments — from real queries run through each tool.

Citation

Independent CRB Working Group. “Consumer Research Bench: Evaluating AI Agents for Consumer Research.” Working Draft v0.6, 2026.

@misc{crb2026,
  title={Consumer Research Bench: Evaluating AI Agents for Consumer Research},
  author={Independent CRB Working Group},
  year={2026},
  note={Working Draft v0.6}
}
Download PDF
Version History
v0.6
Current · Working DraftConstruct, two tracks, RetroSignals, C-RACE, and validity gates documented.
v0.5
Methodology DraftReport Track dimensions and C-RACE scoring drafted.
v0.4
Methodology DraftRetroSignals environment scope defined.
v0.3
Methodology DraftObjective Track task taxonomy outlined.

Participate

CRB is open to anyone — any consumer-research system, agentic, workflow-based, or human baseline.

Open, PR-based submissions. Fork the repo, run the benchmark, and open a pull request with your results and the artifacts that reproduce them — a maintainer reviews it, independently re-runs or re-scores it against the rubric and validity gates, then merges it. Entries we re-verify are tagged Verified; self-reported ones are listed as Community. Every entry discloses its models, vendor, and data sources, and the private held-out set is scored by a neutral custodian — so first-party entries can't tune to labels they can't see.

Status. The spec, rubrics, and task schema are public; the sample tasks and submission harness are being released. You can start now — read the submission guide, review the rubrics, or propose a task via PR.

Governance

CRB is intended to be governed independently of any vendor or data provider.

  • Independent task authorship
  • Blind grading
  • Pre-registered criteria
  • Private held-out test set
  • Public sample tasks
  • Human validation
  • Vendor and data-provider disclosure
  • Verified & self-reported result tiers
  • Independent re-scoring of submissions
  • Rolling benchmark refresh
Done