GEO

Stage 02 · Diagnose

Baseline Measurement: Establish What Is Actually Visible

Define an auditable starting point before changing content. Separate observed model answers, website analytics, self-reported operations, and inferred business impact.

Outcome

Design a baseline that tracks prompts, citations, answer share, qualified visits, and Pipeline without manufacturing certainty.

Prerequisite
Complete or review the previous stage: GEO fundamentals
Working effort
60 to 90 min learning + one working session
You leave with
A versioned query set, evidence log, metric dictionary, and baseline snapshot with explicit unknowns.

Learn → Do → Prove

01

Learn the system

Learn which GEO metrics answer visibility, influence, and revenue questions, and which evidence cannot answer them.

Build the mental model

02

Do the work

Create a repeatable query sample and record model, market, date, answer, citation, and brand presence.

Create the working artifact

03

Prove the result

Re-run the sample and show another teammate can reproduce the collection method and interpret its limits.

Check the evidence

Concept boundaries

Query set

A versioned sample of buyer questions used to observe answers under declared collection conditions.

It is a sample, not a census of every prompt buyers may use.

Answer role

The brand's function in a response, such as absent, mentioned, cited, compared, or recommended.

The categories describe the captured answer; they do not prove persuasion or revenue.

Evidence log

A record of the prompt, engine, market, date, answer, sources, method, and reviewer notes.

A log supports reproducibility only when the collection method and missing data are retained.

Core lesson

01

Separate visibility, behavior, and commercial evidence

A reliable baseline uses several evidence layers without pretending they are interchangeable.

Captured answers can show observed presence and citation. Website analytics can show attributable visits and on-site behavior. CRM records can show known opportunities and revenue. Self-reported process data can show whether a team follows an operating practice.

The layers can inform one another, but a gap remains between them. Dark traffic, changing answers, consent, attribution windows, and small samples make that gap part of the report rather than something to hide.

  • Assign every metric to one evidence layer.
  • Write what the metric cannot answer.
  • Keep raw observations available behind summaries.
02

Design a sample another person can rerun

Consistency matters more than a large dashboard that cannot be reproduced.

Define buyer stage, market, language, query wording, engine or surface, account state when relevant, and collection date. Version the query set when a question changes instead of silently overwriting it.

Use the same classification rules across runs. When reviewers disagree, preserve the disagreement and refine the rule; do not force a clean score by changing categories after seeing the answer.

  • Freeze the initial query-set version.
  • Write a role-classification rule with examples.
  • Have a second reviewer classify a small sample independently.

Decision framework

The evidence ladder

What strength of claim can this observation support?

  1. 01

    Observed

    Was the answer or event directly captured?

    Report exactly what occurred under the recorded conditions.

  2. 02

    Repeated

    Was the same method rerun across declared times or reviewers?

    Discuss stability or variance within the sample.

  3. 03

    Connected

    Can the observation be linked to a separate behavior or CRM record?

    Report the connection and attribution rule, not assumed causality.

  4. 04

    Causal

    Was a valid design used to isolate the change?

    Make causal claims only when the design actually supports them.

Worked non-client example

A team captures a brand citation in several answers and sees an increase in direct visits during the same period.

  • The citation captures are timestamped.
  • Direct traffic increased in analytics.
  • No visitor-level link joins the answers to those sessions.

Report answer presence and direct-traffic movement as parallel observations, then add a self-report or tagged journey where appropriate.

Timing alone does not connect the two datasets.

The evidence supports monitoring and a better attribution design, not a claim that the citations caused the traffic change.

Reusable work template

Baseline evidence log

Store raw captures before calculating summary metrics.

  1. 01

    Query ID and version

    Use a stable identifier and retain previous wording.

  2. 02

    Collection context

    Record engine, surface, locale, market, date, and relevant account state.

  3. 03

    Answer capture

    Preserve the response or a reviewable extract and cited sources.

  4. 04

    Brand role

    Classify using a written absent/mention/citation/comparison/recommendation rule.

  5. 05

    Linked evidence

    Reference analytics or CRM records only when the linking rule is declared.

  6. 06

    Unknowns

    List sampling, personalization, tracking, and attribution limitations.

Failure modes and corrections

Changing prompts between runs

The dashboard compares results collected with different wording but labels them as a trend.

The measurement changed with the subject.

Version prompts and compare only compatible samples.

Compressing all roles into one score

Mentions, citations, and recommendations receive an unexplained blended value.

The score hides the behavior that should drive the next decision.

Show role counts and source captures before any composite.

Backfilling certainty

Unknown answers, missing captures, or dark traffic disappear from the report.

The baseline looks cleaner while becoming less auditable.

Keep an explicit unknown state and report denominator changes.

Practice exercise

Build and rerun a baseline sample

Use a small set of real questions from one buyer stage and one market.

  1. 01Freeze the query wording and classification guide.
  2. 02Capture answers and cited sources under declared conditions.
  3. 03Ask a second reviewer to classify a subset.
  4. 04Rerun the sample and explain variance without claiming causality.

Proof artifact

A versioned query set, raw evidence log, metric dictionary, and baseline note with unknowns.

Completion rubric

  • A teammate can reproduce the method.
  • Raw captures support every summary.
  • Unknown and missing states remain visible.
  • Claims stay within the evidence ladder.

ACADEMY KNOWLEDGE LIBRARY

Browse the full library

Start with a stage, then use concept and intent signals to choose the right depth.

Show every source in this stage
considerationhubsearchable

AI Visibility Tool Selection Guide: Match Business Needs to Data Sources

Teams often select the dashboard before defining the decision. Use this map to connect the business need, data source, and tool category.

Read article
awarenessdatashareable

2026 Taiwan AI Search Citation Report: Which Traditional Chinese Sources Does ChatGPT Use?

ChatGPT does not choose Traditional Chinese sources by traffic alone. We tracked which sources appear across a set of Taiwan-focused questions and where brands should begin closing their citation gaps.

Read article
considerationdatashareable

AI Citation Volatility: Tracking the Same Prompt Set for Three Months

Ask the same question three months apart and nearly half the brand list may change. This study shows why citation stability matters more than a one-time ranking.

Read article
awarenessdatashareable

Annual Study of Perplexity's Traditional Chinese Sources: News, Wikipedia, and Brand Websites

News, Wikipedia, and official sites account for most citations in our Traditional Chinese Perplexity sample, while the source a brand controls directly is the smallest of the three.

Read article
awarenessdatashareable

B2B AI Citation Benchmarks: SaaS, Manufacturing, and Professional Services

Your industry shapes your starting point for AI citations. SaaS, professional services, and manufacturing should be measured against their own baselines, not one misleading overall average.

Read article
considerationdataboth

Content Length, Freshness, and AI Citations: Evidence from 500 Articles

Longer content is not automatically easier for AI systems to cite. This study found an inverted-U relationship with length, while freshness affected only certain kinds of queries.

Read article

Proof task

Proof task: Baseline measurement

Re-run the sample and show another teammate can reproduce the collection method and interpret its limits.

  1. 01Capture the starting evidence — Learn which GEO metrics answer visibility, influence, and revenue questions, and which evidence cannot answer them.
  2. 02Complete the stage artifact — Create a repeatable query sample and record model, market, date, answer, citation, and brand presence.
  3. 03Review it against the outcome — Re-run the sample and show another teammate can reproduce the collection method and interpret its limits.

Deliverable

A versioned query set, evidence log, metric dictionary, and baseline snapshot with explicit unknowns.

Supporting field library

Continue the topic cluster

Use these resources for depth. Some premium whitepapers retain their existing library gate; the core stage remains open.

Whitepaper6 chapters · 2026.06

The Three-Tier GEO Metrics Whitepaper

Citation rate, answer share, Pipeline contribution: what each one measures, how to capture it, and why traffic should never be your KPI. Includes weekly-report field definitions.

The Three-Tier GEO Metrics Whitepaper
Report6 chapters · 2026.06

Share of Model: Turn AI Visibility Into a Number You Can Benchmark

Now that clicks are vanishing, how often you get mentioned and cited in AI answers (versus your competitors), has become the new KPI. This report hands you the 2026 Share of Model benchmarks and shows you how to turn them into a competitive scorecard you can put on a slide and benchmark against rivals.

Share of Model: Turn AI Visibility Into a Number You Can Benchmark
Guide6 chapters · 2026.06

57% of Traffic Is "Direct/Unknown": The Dark-Traffic Attribution Guide

AI referrals strip the referrer, and Google folds AI Mode into organic, so you can't even see AI traffic in GA4. This guide shows you how to use server-side GTM and an agent-to-pipeline framework to recover the AI traffic vanishing into "direct/unknown" and pipe it into your CRM and lead scoring.

57% of Traffic Is "Direct/Unknown": The Dark-Traffic Attribution Guide
Report6 chapters · 2026.06

90% Less Traffic, 5x More Sales: Stop Judging AI Traffic by Clicks

AI strips your clicks to the bone, yet sends you visitors who convert several times better. This report pulls together the conflicting 2026 conversion data and helps you reframe the 'AI is stealing our traffic' panic into the argument that 'AI is sending us people who are more likely to buy.'

90% Less Traffic, 5x More Sales: Stop Judging AI Traffic by Clicks

Apply the stage with a field tool

Query Evidence Lab

Create a repeatable query sample and record model, market, date, answer, citation, and brand presence.

Evidence boundary

Use the output for the decision it describes; do not treat a technical scan, self-assessment, or planning model as proof of live AI citations.

APPLY THE LEARNING

Move from the lesson to an inspectable next decision

Use the linked tool, diagnostic, or service only when its evidence base matches the decision you need to make.

Open the next action