GEO

阶段 02 · 诊断

基线测量:确认当前真实可见度

在修改内容之前建立可审计起点,分开模型答案、分析数据、自报运营与业务推断。

结果

设计追踪问题、引用、答案份额、合格访问与 Pipeline 的可信基线。

先决条件
完成或复习上一阶段:GEO 基础
投入
60–90 分钟学习 + 一次工作时段
成果
版本化问题集、证据日志、指标字典与未知项清单。

学习 → 实施 → 证明

01

理解系统

学习可见度、影响和收入指标各自能回答什么。

建立模型

02

完成工作

建立可重复的问题样本并记录模型、市场、日期、答案和来源。

制作成果

03

验证结果

让另一位同事能够复现采集方法并解释其限制。

检查证据

Concept boundaries

Query set

A versioned sample of buyer questions used to observe answers under declared collection conditions.

It is a sample, not a census of every prompt buyers may use.

Answer role

The brand's function in a response, such as absent, mentioned, cited, compared, or recommended.

The categories describe the captured answer; they do not prove persuasion or revenue.

Evidence log

A record of the prompt, engine, market, date, answer, sources, method, and reviewer notes.

A log supports reproducibility only when the collection method and missing data are retained.

Core lesson

01

Separate visibility, behavior, and commercial evidence

A reliable baseline uses several evidence layers without pretending they are interchangeable.

Captured answers can show observed presence and citation. Website analytics can show attributable visits and on-site behavior. CRM records can show known opportunities and revenue. Self-reported process data can show whether a team follows an operating practice.

The layers can inform one another, but a gap remains between them. Dark traffic, changing answers, consent, attribution windows, and small samples make that gap part of the report rather than something to hide.

  • Assign every metric to one evidence layer.
  • Write what the metric cannot answer.
  • Keep raw observations available behind summaries.
02

Design a sample another person can rerun

Consistency matters more than a large dashboard that cannot be reproduced.

Define buyer stage, market, language, query wording, engine or surface, account state when relevant, and collection date. Version the query set when a question changes instead of silently overwriting it.

Use the same classification rules across runs. When reviewers disagree, preserve the disagreement and refine the rule; do not force a clean score by changing categories after seeing the answer.

  • Freeze the initial query-set version.
  • Write a role-classification rule with examples.
  • Have a second reviewer classify a small sample independently.

Decision framework

The evidence ladder

What strength of claim can this observation support?

  1. 01

    Observed

    Was the answer or event directly captured?

    Report exactly what occurred under the recorded conditions.

  2. 02

    Repeated

    Was the same method rerun across declared times or reviewers?

    Discuss stability or variance within the sample.

  3. 03

    Connected

    Can the observation be linked to a separate behavior or CRM record?

    Report the connection and attribution rule, not assumed causality.

  4. 04

    Causal

    Was a valid design used to isolate the change?

    Make causal claims only when the design actually supports them.

Worked non-client example

A team captures a brand citation in several answers and sees an increase in direct visits during the same period.

  • The citation captures are timestamped.
  • Direct traffic increased in analytics.
  • No visitor-level link joins the answers to those sessions.

Report answer presence and direct-traffic movement as parallel observations, then add a self-report or tagged journey where appropriate.

Timing alone does not connect the two datasets.

The evidence supports monitoring and a better attribution design, not a claim that the citations caused the traffic change.

Reusable work template

Baseline evidence log

Store raw captures before calculating summary metrics.

  1. 01

    Query ID and version

    Use a stable identifier and retain previous wording.

  2. 02

    Collection context

    Record engine, surface, locale, market, date, and relevant account state.

  3. 03

    Answer capture

    Preserve the response or a reviewable extract and cited sources.

  4. 04

    Brand role

    Classify using a written absent/mention/citation/comparison/recommendation rule.

  5. 05

    Linked evidence

    Reference analytics or CRM records only when the linking rule is declared.

  6. 06

    Unknowns

    List sampling, personalization, tracking, and attribution limitations.

Failure modes and corrections

Changing prompts between runs

The dashboard compares results collected with different wording but labels them as a trend.

The measurement changed with the subject.

Version prompts and compare only compatible samples.

Compressing all roles into one score

Mentions, citations, and recommendations receive an unexplained blended value.

The score hides the behavior that should drive the next decision.

Show role counts and source captures before any composite.

Backfilling certainty

Unknown answers, missing captures, or dark traffic disappear from the report.

The baseline looks cleaner while becoming less auditable.

Keep an explicit unknown state and report denominator changes.

Practice exercise

Build and rerun a baseline sample

Use a small set of real questions from one buyer stage and one market.

  1. 01Freeze the query wording and classification guide.
  2. 02Capture answers and cited sources under declared conditions.
  3. 03Ask a second reviewer to classify a subset.
  4. 04Rerun the sample and explain variance without claiming causality.

Proof artifact

A versioned query set, raw evidence log, metric dictionary, and baseline note with unknowns.

Completion rubric

  • A teammate can reproduce the method.
  • Raw captures support every summary.
  • Unknown and missing states remain visible.
  • Claims stay within the evidence ladder.

GEO 学院知识库

完整知识库

先选阶段,再用主题与阅读意图选择深度。

展开本阶段全部内容
considerationhubsearchable

AI 可见度工具选型指南(总览):从业务需求到引用来源的完整地图

多数团队的选型顺序都反了。这是一份从需求定义到工具导入的完整地图。

阅读文章
awarenessdatashareable

2026 台湾 AI 搜索引用现状报告:从 ChatGPT 引用来源看内容可见度

ChatGPT 判断内容是否值得引用时,并不会只看网站流量。我们追踪了一组来自台湾的真实问题,分析 AI 正在引用哪些来源,以及品牌应该从哪里着手提升可见度。

阅读文章
considerationdataboth

AI 可量化品牌曝光前后的变化:如何衡量 GEO 成效?

别凭感觉判断品牌有没有被 AI 看见。通过固定问题集、多引擎重复测试和四项核心指标,可以把 AI 可见度转化为可比较、可追踪的趋势曲线。

阅读文章
considerationdatashareable

AI 引用稳定性追踪:同一组问题,三个月后答案变了多少?

同一个问题,三个月后 AI 给出的品牌名单可能已有近一半发生变化。这项追踪调研揭示了 AI 引用的波动程度,也解释了为什么持续稳定性比单次排名更值得关注。

阅读文章
awarenessdatashareable

B2B 行业 AI 引用基准:SaaS、制造业与专业服务,谁更容易被提及?

品牌被 AI 引用的初始水平,很大程度上由所在行业决定。SaaS、专业服务和制造业的内容基础差异明显,用全行业平均值衡量自己,几乎一定会得出错误结论。

阅读文章
considerationdataboth

GEO 多久见效?真实项目的时间线与关键里程碑

GEO 通常分三个阶段见效:约第 4 至 8 周首次被引用,第 8 至 12 周可以量化品牌曝光,第 12 至 24 周开始带来业务机会。大多数失败项目,都是在跨过见效门槛前就停了。

阅读文章

验证任务

验证任务:基线测量

让另一位同事能够复现采集方法并解释其限制。

  1. 01保存起始证据 — 学习可见度、影响和收入指标各自能回答什么。
  2. 02完成阶段成果 — 建立可重复的问题样本并记录模型、市场、日期、答案和来源。
  3. 03按目标复核 — 让另一位同事能够复现采集方法并解释其限制。

交付成果

版本化问题集、证据日志、指标字典与未知项清单。

延伸资料库

继续主题学习

核心阶段保持开放;部分进阶白皮书继续使用原有解锁方式。

白皮书6 章 · 2026.06

GEO 三层指标白皮书

引用率、答案占有率、Pipeline 贡献 — 各自衡量什么、怎么取数、为什么流量不该是 KPI。附周报字段定义。

GEO 三层指标白皮书
报告6 章 · 2026.06

Share of Model:把 AI 声量变成能跟对手比的数字

当点击消失,「你在 AI 答案里被提及/引用的频率 vs 竞争对手」成了新的 KPI。这份报告给你 2026 年的 Share of Model 基准数据,以及如何把它做成一张能放进汇报、跟对手对标的竞争计分卡。

Share of Model:把 AI 声量变成能跟对手比的数字
指南6 章 · 2026.06

57% 流量是「直接/未知」:找回暗流量的归因指南

AI 推荐会剥掉 referrer,Google 又把 AI Mode 并进 organic,导致你在 GA4 里根本看不到 AI 流量。这份指南教你用 server-side GTM 与 agent-to-pipeline 框架,把消失在「直接/未知」里的 AI 流量找回来,并接进 CRM 与 lead scoring。

57% 流量是「直接/未知」:找回暗流量的归因指南
报告6 章 · 2026.06

流量少 90%,成交多 5 倍:别用点击数评判 AI 流量

AI 把点击量砍光,却送来转化率高出数倍的访客。这份报告梳理 2026 年各方冲突的转化数据,帮你把「AI 抢走流量」的恐慌,重新定义成「AI 送来更会买的人」的论述。

流量少 90%,成交多 5 倍:别用点击数评判 AI 流量

用工具完成本阶段

Query Evidence Lab

建立可重复的问题样本并记录模型、市场、日期、答案和来源。

证据边界

Use the output for the decision it describes; do not treat a technical scan, self-assessment, or planning model as proof of live AI citations.

应用学习

进入下一个可检查的决策

只在证据基础符合决策时使用工具、诊断或服务。

打开下一步