oGoing AI Visibility Engine

    How we score AI brand visibility

    The full methodology behind the oGoing AI Visibility Engine — what we ask AI engines, how we grade their answers, and where the limits of the score are today.

    Methodology version: 2026.11·Models updated: May 2026·Last lab audit: — (pending)

    1. Overview

    Each scan asks three flagship AI engines about your brand and grades their answers across four dimensions. Three of them (Recognition, Accuracy, Brand Authority) combine into a single 0–100 visibility score. The fourth (Recommendation) is a separate signal that asks whether AI suggests you when customers search for what you sell.

    ComponentWeightMax pointsWhat it measures
    Recognition30%30Do AI engines know your brand exists?
    Accuracy20%20Are AI responses factually correct?
    Brand Authority50%50How much authoritative detail does AI cite?
    Total100%100Final AI Visibility score

    Each engine is scored independently on these three components. The final score is the weighted blend across all responding engines. Recommendation (below) is a separate signal — it answers a different question and is not part of the 0–100 visibility score.

    Coverage cap. A brand must be recognized by all three engines (ChatGPT, Gemini, Perplexity) to be eligible for a top-tier visibility score. If a major AI engine doesn't recognize the brand, the final blended score is capped — regardless of how the other engines scored.
    Engines that recognized the brandMaximum final score
    3 of 3100 (no cap)
    2 of 370
    1 of 335
    0 of 320
    Additionally, engines that don't recognize the brand are excluded from the Accuracy and Authority denominators — they neither inflate nor dilute those sub-scores. They still cost Recognition points (the brand wasn't recognized), and they trigger the coverage cap above.

    2. Engines & models

    We use the flagship consumer-equivalent model from each major AI vendor. These are the same models a customer would encounter using ChatGPT, Gemini, or Perplexity directly — so the responses we grade are representative of real consumer experience.

    EngineModelWhy we use itLive web?
    ChatGPTopenai/gpt-5Flagship consumer model for unaided recallNo (knowledge-only)
    Geminigoogle/gemini-3.1-pro-previewFlagship consumer model for unaided recallNo (knowledge-only)
    Perplexitysonar-proLive web search with citationsYes
    A note on extraction: Authority signal extraction uses a separate, faster model (gemini-3-flash-preview) optimized for structured boolean classification. The engine cards above use flagship models for the consumer-facing responses we score. We expose the per-engine breakdown so you can see exactly what each model said about you.

    3. Recognition (30 pts)

    Recognition measures whether AI engines know your brand exists when a customer asks about you by name. It's the foundation of AI visibility — if no engine recognizes you, nothing else matters.

    Sub-signalWeightHow it's measured
    Gemini unaided recall12 pts"Tell me about {brand}" — confidence × specificity
    ChatGPT aided recognition6 ptsAided search; symmetric unaided scoring on roadmap
    Perplexity citations7.5 ptsCitation count thresholds (1 / 3 / 5 / 8+)
    Response specificity4.5 ptsAverage detail signals across engines

    4. Accuracy (20 pts)

    Accuracy asks: when AI engines do talk about you, are their hard facts right? As of v2026.08 Accuracy is computed per engine from four weighted components and reduced when this engine disagrees with the other engines on a checkable fact.

    Per-engine formula

    • 0.40 — Confidence weight (high = 1.0, medium = 0.6, low = 0.3)
    • 0.30 — Factual grounding: share of canonical facts (founding year, HQ, founder, employees) that match ≥1 other engine, plus a boost when the brand's official domain is cited
    • 0.20 — Specificity: at least 3 distinct fact types present in the response
    • 0.10 — Location coherence (only when the scan has a location constraint)

    Penalties: −0.15 per fact this engine disputes against a ≥2-engine quorum (cap −0.45), −0.20 if the engine is flagged as describing a different entity, −0.25 if the response contains explicit uncertainty phrases.

    Worked example: ChatGPT says founded 2003, Gemini and Perplexity both say 2007. ChatGPT loses 0.15 from its Accuracy quality (other components unaffected). Gemini and Perplexity are unaffected — they're the quorum.

    Factual divergence — entity vs fact

    Entity-level divergence (an engine describes a fundamentally different brand than the others) is surfaced as a transparency banner. It does reduce that engine's own Accuracy by 0.20 because it's wrong about your brand, but it does not trigger penalties at the report level — coverage is the real story there.

    Fact-level divergence (same brand, conflicting hard facts) is the core Accuracy signal. The quorum rule prevents noise: a lone outlier in a 3-way disagreement is shown but not penalised — when no two engines agree on a value, we don't claim to know who's right.

    Where the old "verification credits" went

    Two credits — official-domain citation and cross-engine match — now feed Accuracy's factual-grounding component (they're independent confirmation that this engine's facts check out). The other four credits — third-party citations, structured-facts confirmation, local-prose signals, and quantitative claims — are "evidence of knowledge" and were moved into Brand Authority where they belong. Net result: the same evidence is counted once, in the pillar it actually measures.

    5. Brand Authority (50 pts)

    Brand Authority is the largest scoring bucket because it's the strongest predictor of whether AI will speak about you with substance. Each engine's response is independently graded on 12 authority signals.

    Founding year

    Is the year the company was founded mentioned?

    Named individuals

    Are specific people (founders, leaders) named?

    Quantifiable scale

    Are employees, revenue, or customer counts cited?

    Certifications & awards

    Are credentials or recognitions listed?

    Named clients

    Are specific customers or brands they serve named?

    Geographic detail

    Are cities, regions, or service areas specified?

    Primary services

    Are the core products or services described?

    Named competitors

    Are alternatives or competitors mentioned?

    Press & media

    Are publications or media coverage cited?

    Partnerships

    Are strategic partners or integrations named?

    Longevity

    Is operational history (years in business) mentioned?

    Differentiation

    Is what makes them unique articulated?

    How signals are extracted

    Each engine's raw response is passed independently to a structured-extraction model (gemini-3-flash-preview) with a strict 12-question prompt. The output for every signal is one of: true, false, or Ambiguous. Ambiguous signals are preserved (not silently coerced to false) — see §7.

    Why a single extractor

    Cross-vendor diversity already exists at the engine layer (GPT-5, Gemini 3.1 Pro, Sonar Pro are three different vendors with three different training corpora). A second extractor model would add cost without meaningfully reducing extractor bias, because the three independent responses already provide natural cross-checking. We instead invest in cross-engine consistency rules and the labeled-set calibration described in §8.

    5b. The Evidence Matrix

    Every brand report renders the same evidence checklist, in the same order, with the same row definitions. The Matrix is a transparency layer — it does not change the score, it shows the inputs the Authority sub-score is built from. Two brands with the same score should have visibly different rows; two brands with very different scores should have visibly different row counts.

    What the Matrix shows
    • Signal — one of the 12 authority signals: 4 shared signals (Named Individuals, Founding Year, Competitive Differentiation, Revenue/Scale) plus 8 reach-specific signals (global brands get product lines, market cap, awards, etc.; local brands get certifications, reviews, service area, community ties, etc.).
    • Found — checkmark if at least one engine surfaced the signal in its response, ✗ if none did.
    • Confirmed by — which engines (Perplexity, ChatGPT, Gemini) surfaced the signal. When per-engine structured data is available, this column shows individual engine pills. When only combined-text detection is available, it shows a single "Combined responses" badge.
    • Why — a one-sentence explanation of why AI engines weight this particular signal.
    The cross-engine confirmation rule

    A signal is marked Found if any engine surfaces it. The Confirmed-by column then enumerates which. We do not require unanimous agreement to mark a signal Found, because engines have different training corpora and one engine knowing a fact is genuine signal — not noise.

    null ≠ false. When an engine returns an error or 402 (rate-limit / temporarily unavailable), that engine's cell renders as a muted dash, not an ✗. A missing response is not the same as a denial. This rule also governs the Authority sub-score itself: error responses are excluded from the per-engine average rather than scored as zero.
    Backfill policy: The Matrix does not retroactively render on pre-Stage-3 reports. Older Observatory entries show their original Authority Checklist view; new scans (and rescans) get the Matrix automatically.

    5c. Per-engine scoring

    Authority is not extracted once from a blended pile of text — it is extracted independently from each engine's own response. Every engine therefore gets its own Recognition / Accuracy / Authority triad, and the report-level number you see at the top is a weighted blend of those three triads.

    How blending works

    • Recognition — averaged across all engines that returned a response.
    • Accuracy — averaged across engines that recognised the brand; engines that disagree with the ≥2-engine quorum on a hard fact (founding year, HQ, founder, employees) have their own Accuracy reduced by 0.15 per disputed fact.
    • Authority — averaged across engines whose authority extraction completed successfully. Engines that timed out or whose extraction failed are excluded from the blend per the null ≠ false rule, never counted as zeros.
    Why the engines often disagree. GPT-5, Gemini 3.1 Pro and Sonar Pro are trained on different corpora with different refresh cycles. Entity-level divergence (one engine describing a different brand) is surfaced as a transparency banner and never penalised at the report level. Fact-level divergence (same brand, conflicting hard facts) is the core Accuracy signal — see §4.
    Backfill policy: Per-engine scoring is rendered for new scans only. Pre-Stage-4 reports show the blended score with a note inviting a rescan to populate the per-engine view.

    6. Recommendation

    Recommendation is a separate 0–100 signal that answers a different question: when a customer asks AI for a recommendation in your category — without naming you — does AI suggest you? This is the conversion-relevant metric.

    • Method: 10 reach-aware consumer queries (best, near-me, comparison, problem-led, etc.) on Perplexity Sonar.
    • Score: percentage of queries where your brand appears in the top recommendations.
    Current limitation: Recommendation today runs on Perplexity Sonar only. Multi-engine recommendation parity (ChatGPT live-search, Gemini-grounded) is on the roadmap — see §9.

    7. What counts as Ambiguous

    A signal is marked Ambiguous when the AI response hedges ("roughly three dozen employees", "founded in the early 2010s"), uses unverifiable rounded numbers, or describes the signal indirectly. Ambiguous signals are not counted as either present or absent — they're surfaced separately in the score breakdown so you can see what AI almost-knows about you. This is by design: silently coercing hedged answers to false would understate visibility; coercing them to true would inflate it.

    8. Lab calibration

    Honest disclosure: the extractor has not yet been graded against a hand-labeled answer key. We are designing a 200-brand labeled set and the audit script that will produce a monthly per-signal precision/recall table on this page.

    Calibration status

    Static disclosure — will update automatically once first audit lands.

    Labeled-set audit
    Not yet run — design in progress
    Last audited
    Next scheduled
    Q3 2026 (after labeled-set completion)
    Methodology version
    2026.04

    Target gates (per signal)

    • Precision ≥ 95% — when we say "true", we're right
    • Confidently-wrong = 0 — never "true" on a "false" ground truth
    • Ambiguous rate ≤ 15% — extractor isn't dodging the question

    Once the first audit lands, this section becomes a live per-signal table with a Last audited: YYYY-MM-DD stamp.

    9. Limitations & roadmap

    • Symmetric unaided-recall scoring across all three engines (today only Gemini gets the full unaided rubric; ChatGPT scoring is aided).
    • Multi-engine Recommendation — today Perplexity Sonar only; ChatGPT live-search and Gemini-grounded are on deck.
    • Published lab-calibration table once the 200-brand labeled set is graded — see §8.

    10. Methodology changelog

    Every material change to scoring, models, or signal definitions is logged here. When your score shifts month-over-month, this is the first place to look.

    v2026.11— Per-engine Accuracy re-synced with the pillar; prose-grounded fallbacks; Perplexity Recognition rewrite
    • Per-engine Accuracy re-sync. Per-engine Accuracy was previously written during the authority-extraction loop, before canonical facts had been attached to each result — so the cross-engine fact-match component contributed 0. The pillar then recalculated with facts attached and landed higher. The displayed engine numbers and the pillar were computed from the same formula but different inputs, which read as inconsistent. Per-engine Accuracy is now recomputed once more, after the pillar recalc, so card values and the pillar average use identical inputs.
    • Prose-grounded factual-grounding fallback (Gap 1). When an engine returns zero URL citations (Gemini always; ChatGPT often, capped at 3), factual grounding can now lift via two prose paths: official domain mentioned in the answer (e.g. hubspot.com in the body) sets the 0.5 floor; ≥3 canonical facts (year, HQ, founder, employees) that match what another engine independently said sets the 0.7 floor. Confidence weight (40%), specificity (20%), location coherence (10%), and the −0.15 per disputed fact penalty all still apply — the floor is below the 1.0 a URL-cited answer can reach.
    • Perplexity Recognition multi-factor rewrite (Gap 2). Perplexity Recognition was a pure citation-count ladder (an artifact of how many URLs Sonar surfaced). Rewrote to confidence base (high 70 / medium 50 / low 25) + specificity bonus (≥75% → +20) + citation depth (capped +10) + official-domain bonus (+5). All three engines now answer the same Recognition question (does it know the brand?), each adapted to its evidence shape.
    • ChatGPT Stage 1 SLA extended. Raised the fast-lane wall from 30s → 40s outer (inner 38s) to reduce timeout_30s failures on heavy web_search runs. Two-stage architecture unchanged; Stage 2 enrichment still backfills deeper company-profile data asynchronously.
    • No rubric weight changes (Recognition 30 / Accuracy 20 / Authority 50). No coverage-cap changes. No prompt or model swaps. Historical reports are not re-baselined; the new formulas apply to new scans only.
    v2026.08— Accuracy sharpened into a per-engine divergence score
    • Rewrote per-engine Accuracy as 0.40 confidence + 0.30 factual grounding + 0.20 specificity + 0.10 location coherence, with a −0.15 per disputed fact penalty (cap −0.45) when this engine disagrees with the ≥2-engine quorum on founding year, HQ, founder, or employees.
    • Narrowed the divergence policy: entity-level divergence (different brand) remains surfaced and never penalised at the report level. Fact-level divergence (same brand, conflicting hard facts) now reduces the outlier engine's own Accuracy — the core sharpening signal.
    • Reassigned 4 of the 6 verification credits (third-party citations, structured-facts confirmation, local-prose signals, quantitative claims) from Accuracy into Brand Authority as up to +4 pts on the 50-point scale. Official-domain citation and cross-engine match stay in Accuracy as factual-grounding signals.
    • Removed the length-nudge and loose specificity regex set from Accuracy — they were noise.
    • No rubric weight changes (Recognition 30 / Accuracy 20 / Authority 50). Historical reports are not re-baselined; the new formula applies to new scans only.
    v2026.06— Per-engine scoring transparency
    • Surfaced independent Recognition / Accuracy / Authority scores for each engine (ChatGPT, Gemini, Perplexity) on the results page and in the Per-Engine Scoring comparison panel.
    • Added §5c "Per-engine scoring" to this page documenting how engine scores blend into the report-level score.
    • Added per-engine score-contribution tooltips on every Evidence Matrix engine pill, plus a "Engines disagree" badge when at least one engine confirms a signal and another rejects it.
    • Codified the null ≠ false rule in UI: engines whose Authority extraction failed are excluded from the blend (shown as "—" or "excluded"), never penalized as zeros.
    • Added an in-card "Signals this engine confirmed" collapsible listing the specific authority labels each engine confirmed vs. missed.
    • No scoring weights or formulas changed — Stage 4 is transparency only. Underlying R/A/A inputs and the coverage curve are unchanged.
    v2026.07— Accuracy verification credits + divergence transparency
    • Removed the local/regional Accuracy hard-cap (was ceiling=70 unless trade-specific signals fired). Replaced with additive verification credits from evidence we already collect: official-domain citations, third-party sources, cross-engine agreement, structured authority extraction, and quantitative claims.
    • Added a Verification sub-row to the Per-Engine Scoring panel showing exactly which credits fired for each engine.
    • Added a Factual Divergence panel surfacing canonical scalar facts (founding year, HQ, employee band, founder) per engine. Conflicts trigger an "Engines disagree" badge. Zero scoring impact — divergence is informational only.
    • Per-engine Accuracy now receives allResults + brandName so cross-engine match and official-domain credits can be computed.
    v2026.05— Evidence Matrix transparency layer
    • Added the Evidence Matrix to every brand report — same 12-row checklist applied symmetrically across all brands.
    • Added per-engine "Confirmed by" pills with explicit null ≠ false handling for engine errors and 402s.
    • Added a one-sentence "why this matters" rationale to each authority signal, surfaced inline as tooltips.
    • Documented the cross-engine confirmation rule and the backfill policy in §5b.
    • No scoring weights changed — Stage 3 is display + symmetry only.
    v2026.04— Initial public methodology
    • Documented the 12 authority signals and the Ambiguous policy.
    • Surfaced engine-card models (GPT-5, Gemini 3.1 Pro, Sonar Pro) and the structured-extraction model (gemini-3-flash-preview).
    • Documented the lab-calibration framework and target gates (audit pending).
    • Added cross-links from the in-product score breakdown to this page.