Skip to content
GEO

Stage 02 · Diagnose

Baseline Measurement: Establish What Is Actually Visible

Define an auditable starting point before changing content. Separate observed model answers, website analytics, self-reported operations, and inferred business impact.

Outcome

Design a baseline that tracks prompts, citations, answer share, qualified visits, and Pipeline without manufacturing certainty.

Prerequisite
Complete or review the previous stage: GEO fundamentals
Working effort
60 to 90 min learning + one working session
You leave with
A versioned query set, evidence log, metric dictionary, and baseline snapshot with explicit unknowns.

Learn → Do → Prove

01

Learn the system

Learn which GEO metrics answer visibility, influence, and revenue questions, and which evidence cannot answer them.

Build the mental model

02

Do the work

Create a repeatable query sample and record model, market, date, answer, citation, and brand presence.

Create the working artifact

03

Prove the result

Re-run the sample and show another teammate can reproduce the collection method and interpret its limits.

Check the evidence

Concept boundaries

Query set

A versioned sample of buyer questions used to observe answers under declared collection conditions.

It is a sample, not a census of every prompt buyers may use.

Answer role

The brand's function in a response, such as absent, mentioned, cited, compared, or recommended.

The categories describe the captured answer; they do not prove persuasion or revenue.

Evidence log

A record of the prompt, engine, market, date, answer, sources, method, and reviewer notes.

A log supports reproducibility only when the collection method and missing data are retained.

Core lesson

01

Separate visibility, behavior, and commercial evidence

A reliable baseline uses several evidence layers without pretending they are interchangeable.

Captured answers can show observed presence and citation. Website analytics can show attributable visits and on-site behavior. CRM records can show known opportunities and revenue. Self-reported process data can show whether a team follows an operating practice.

The layers can inform one another, but a gap remains between them. Dark traffic, changing answers, consent, attribution windows, and small samples make that gap part of the report rather than something to hide.

  • Assign every metric to one evidence layer.
  • Write what the metric cannot answer.
  • Keep raw observations available behind summaries.
02

Design a sample another person can rerun

Consistency matters more than a large dashboard that cannot be reproduced.

Define buyer stage, market, language, query wording, engine or surface, account state when relevant, and collection date. Version the query set when a question changes instead of silently overwriting it.

Use the same classification rules across runs. When reviewers disagree, preserve the disagreement and refine the rule; do not force a clean score by changing categories after seeing the answer.

  • Freeze the initial query-set version.
  • Write a role-classification rule with examples.
  • Have a second reviewer classify a small sample independently.

Decision framework

The evidence ladder

What strength of claim can this observation support?

  1. 01

    Observed

    Was the answer or event directly captured?

    Report exactly what occurred under the recorded conditions.

  2. 02

    Repeated

    Was the same method rerun across declared times or reviewers?

    Discuss stability or variance within the sample.

  3. 03

    Connected

    Can the observation be linked to a separate behavior or CRM record?

    Report the connection and attribution rule, not assumed causality.

  4. 04

    Causal

    Was a valid design used to isolate the change?

    Make causal claims only when the design actually supports them.

Worked non-client example

A team captures a brand citation in several answers and sees an increase in direct visits during the same period.

  • The citation captures are timestamped.
  • Direct traffic increased in analytics.
  • No visitor-level link joins the answers to those sessions.

Report answer presence and direct-traffic movement as parallel observations, then add a self-report or tagged journey where appropriate.

Timing alone does not connect the two datasets.

The evidence supports monitoring and a better attribution design, not a claim that the citations caused the traffic change.

Reusable work template

Baseline evidence log

Store raw captures before calculating summary metrics.

  1. 01

    Query ID and version

    Use a stable identifier and retain previous wording.

  2. 02

    Collection context

    Record engine, surface, locale, market, date, and relevant account state.

  3. 03

    Answer capture

    Preserve the response or a reviewable extract and cited sources.

  4. 04

    Brand role

    Classify using a written absent/mention/citation/comparison/recommendation rule.

  5. 05

    Linked evidence

    Reference analytics or CRM records only when the linking rule is declared.

  6. 06

    Unknowns

    List sampling, personalization, tracking, and attribution limitations.

Failure modes and corrections

Changing prompts between runs

The dashboard compares results collected with different wording but labels them as a trend.

The measurement changed with the subject.

Version prompts and compare only compatible samples.

Compressing all roles into one score

Mentions, citations, and recommendations receive an unexplained blended value.

The score hides the behavior that should drive the next decision.

Show role counts and source captures before any composite.

Backfilling certainty

Unknown answers, missing captures, or dark traffic disappear from the report.

The baseline looks cleaner while becoming less auditable.

Keep an explicit unknown state and report denominator changes.

Practice exercise

Build and rerun a baseline sample

Use a small set of real questions from one buyer stage and one market.

  1. 01Freeze the query wording and classification guide.
  2. 02Capture answers and cited sources under declared conditions.
  3. 03Ask a second reviewer to classify a subset.
  4. 04Rerun the sample and explain variance without claiming causality.

Proof artifact

A versioned query set, raw evidence log, metric dictionary, and baseline note with unknowns.

Completion rubric

  • A teammate can reproduce the method.
  • Raw captures support every summary.
  • Unknown and missing states remain visible.
  • Claims stay within the evidence ladder.

ACADEMY KNOWLEDGE LIBRARY

Browse the full library

Start with a stage, then use concept and intent signals to choose the right depth.

Show every source in this stage
considerationhubsearchable

AI Visibility Tool Selection Guide: Match Business Needs to Data Sources

Teams often select the dashboard before defining the decision. Use this map to connect the business need, data source, and tool category.

Read article
awarenessdatashareable

2026 Taiwan AI Search Citation Report: Which Traditional Chinese Sources Does ChatGPT Use?

ChatGPT does not choose Traditional Chinese sources by traffic alone. We tracked which sources appear across a set of Taiwan-focused questions and where brands should begin closing their citation gaps.

Read article
considerationdatashareable

AI Citation Volatility: Tracking the Same Prompt Set for Three Months

Ask the same question three months apart and nearly half the brand list may change. This study shows why citation stability matters more than a one-time ranking.

Read article
awarenessdatashareable

Annual Study of Perplexity's Traditional Chinese Sources: News, Wikipedia, and Brand Websites

News, Wikipedia, and official sites account for most citations in our Traditional Chinese Perplexity sample, while the source a brand controls directly is the smallest of the three.

Read article
awarenessdatashareable

B2B AI Citation Benchmarks: SaaS, Manufacturing, and Professional Services

Your industry shapes your starting point for AI citations. SaaS, professional services, and manufacturing should be measured against their own baselines, not one misleading overall average.

Read article
considerationdataboth

Content Length, Freshness, and AI Citations: Evidence from 500 Articles

Longer content is not automatically easier for AI systems to cite. This study found an inverted-U relationship with length, while freshness affected only certain kinds of queries.

Read article

Supporting field library

Continue the topic cluster

Whitepapers, guides and methodology are open to read. Email required for the 2026 GEO Trend Report.

Whitepaper6 chapters · 2026.06

Measure GEO with citations, recommendations and qualified opportunities

Define citations, unlinked brand mentions, recommendation presence in buying questions and qualified opportunities separately. Specify samples and denominators, and distinguish pipeline value, revenue and supported attribution conclusions.

Measure GEO with citations, recommendations and qualified opportunities
Report6 chapters · 2026.06

Share of Model: Compare AI Visibility With Clear Definitions

Define mentions, source links and recommendations before comparing brands. Build a useful scorecard with consistent questions, denominators and collection conditions.

Share of Model: Compare AI Visibility With Clear Definitions
Guide6 chapters · 2026.06

AI referral attribution: known sources and missing data

Read GA4 channel definitions, preserve source evidence through your CRM and compare self-reported discovery separately. Missing referral data cannot identify AI traffic on its own.

AI referral attribution: known sources and missing data

Apply the stage with a field tool

Query Evidence Lab

Create a repeatable query sample and record model, market, date, answer, citation, and brand presence.

Evidence boundary

Use the output for the decision it describes; do not treat a technical scan, self-assessment, or planning model as proof of live AI citations.

APPLY THE LEARNING

Move from the lesson to an inspectable next decision

Use the linked tool, diagnostic, or service only when its evidence base matches the decision you need to make.

Open the next action