Learn the system
Learn which GEO metrics answer visibility, influence, and revenue questions, and which evidence cannot answer them.
Build the mental model
Stage 02 · Diagnose
Define an auditable starting point before changing content. Separate observed model answers, website analytics, self-reported operations, and inferred business impact.
Outcome
Design a baseline that tracks prompts, citations, answer share, qualified visits, and Pipeline without manufacturing certainty.
Learn → Do → Prove
Learn which GEO metrics answer visibility, influence, and revenue questions, and which evidence cannot answer them.
Build the mental model
Create a repeatable query sample and record model, market, date, answer, citation, and brand presence.
Create the working artifact
Re-run the sample and show another teammate can reproduce the collection method and interpret its limits.
Check the evidence
Concept boundaries
A versioned sample of buyer questions used to observe answers under declared collection conditions.
It is a sample, not a census of every prompt buyers may use.
The brand's function in a response, such as absent, mentioned, cited, compared, or recommended.
The categories describe the captured answer; they do not prove persuasion or revenue.
A record of the prompt, engine, market, date, answer, sources, method, and reviewer notes.
A log supports reproducibility only when the collection method and missing data are retained.
Core lesson
A reliable baseline uses several evidence layers without pretending they are interchangeable.
Captured answers can show observed presence and citation. Website analytics can show attributable visits and on-site behavior. CRM records can show known opportunities and revenue. Self-reported process data can show whether a team follows an operating practice.
The layers can inform one another, but a gap remains between them. Dark traffic, changing answers, consent, attribution windows, and small samples make that gap part of the report rather than something to hide.
Consistency matters more than a large dashboard that cannot be reproduced.
Define buyer stage, market, language, query wording, engine or surface, account state when relevant, and collection date. Version the query set when a question changes instead of silently overwriting it.
Use the same classification rules across runs. When reviewers disagree, preserve the disagreement and refine the rule; do not force a clean score by changing categories after seeing the answer.
Decision framework
What strength of claim can this observation support?
Was the answer or event directly captured?
Report exactly what occurred under the recorded conditions.
Was the same method rerun across declared times or reviewers?
Discuss stability or variance within the sample.
Can the observation be linked to a separate behavior or CRM record?
Report the connection and attribution rule, not assumed causality.
Was a valid design used to isolate the change?
Make causal claims only when the design actually supports them.
Worked non-client example
A team captures a brand citation in several answers and sees an increase in direct visits during the same period.
Report answer presence and direct-traffic movement as parallel observations, then add a self-report or tagged journey where appropriate.
Timing alone does not connect the two datasets.
The evidence supports monitoring and a better attribution design, not a claim that the citations caused the traffic change.
Reusable work template
Store raw captures before calculating summary metrics.
Use a stable identifier and retain previous wording.
Record engine, surface, locale, market, date, and relevant account state.
Preserve the response or a reviewable extract and cited sources.
Classify using a written absent/mention/citation/comparison/recommendation rule.
Reference analytics or CRM records only when the linking rule is declared.
List sampling, personalization, tracking, and attribution limitations.
Failure modes and corrections
The dashboard compares results collected with different wording but labels them as a trend.
The measurement changed with the subject.
Version prompts and compare only compatible samples.
Mentions, citations, and recommendations receive an unexplained blended value.
The score hides the behavior that should drive the next decision.
Show role counts and source captures before any composite.
Unknown answers, missing captures, or dark traffic disappear from the report.
The baseline looks cleaner while becoming less auditable.
Keep an explicit unknown state and report denominator changes.
Practice exercise
Use a small set of real questions from one buyer stage and one market.
Proof artifact
A versioned query set, raw evidence log, metric dictionary, and baseline note with unknowns.
Completion rubric
ACADEMY KNOWLEDGE LIBRARY
Start with a stage, then use concept and intent signals to choose the right depth.
Teams often select the dashboard before defining the decision. Use this map to connect the business need, data source, and tool category.
Read articleChatGPT does not choose Traditional Chinese sources by traffic alone. We tracked which sources appear across a set of Taiwan-focused questions and where brands should begin closing their citation gaps.
Read articleAsk the same question three months apart and nearly half the brand list may change. This study shows why citation stability matters more than a one-time ranking.
Read articleNews, Wikipedia, and official sites account for most citations in our Traditional Chinese Perplexity sample, while the source a brand controls directly is the smallest of the three.
Read articleYour industry shapes your starting point for AI citations. SaaS, professional services, and manufacturing should be measured against their own baselines, not one misleading overall average.
Read articleLonger content is not automatically easier for AI systems to cite. This study found an inverted-U relationship with length, while freshness affected only certain kinds of queries.
Read articleProof task
Re-run the sample and show another teammate can reproduce the collection method and interpret its limits.
Deliverable
A versioned query set, evidence log, metric dictionary, and baseline snapshot with explicit unknowns.
Supporting field library
Use these resources for depth. Some premium whitepapers retain their existing library gate; the core stage remains open.
Citation rate, answer share, Pipeline contribution: what each one measures, how to capture it, and why traffic should never be your KPI. Includes weekly-report field definitions.
The Three-Tier GEO Metrics WhitepaperNow that clicks are vanishing, how often you get mentioned and cited in AI answers (versus your competitors), has become the new KPI. This report hands you the 2026 Share of Model benchmarks and shows you how to turn them into a competitive scorecard you can put on a slide and benchmark against rivals.
Share of Model: Turn AI Visibility Into a Number You Can BenchmarkAI referrals strip the referrer, and Google folds AI Mode into organic, so you can't even see AI traffic in GA4. This guide shows you how to use server-side GTM and an agent-to-pipeline framework to recover the AI traffic vanishing into "direct/unknown" and pipe it into your CRM and lead scoring.
57% of Traffic Is "Direct/Unknown": The Dark-Traffic Attribution GuideAI strips your clicks to the bone, yet sends you visitors who convert several times better. This report pulls together the conflicting 2026 conversion data and helps you reframe the 'AI is stealing our traffic' panic into the argument that 'AI is sending us people who are more likely to buy.'
90% Less Traffic, 5x More Sales: Stop Judging AI Traffic by ClicksApply the stage with a field tool
Create a repeatable query sample and record model, market, date, answer, citation, and brand presence.
Evidence boundary
Use the output for the decision it describes; do not treat a technical scan, self-assessment, or planning model as proof of live AI citations.
APPLY THE LEARNING
Use the linked tool, diagnostic, or service only when its evidence base matches the decision you need to make.