Index methodology — how the Digital Visibility Score is measured
Every Local Digital Visibility Index is built the same way — same pillars, same weights, same data sources, same snapshot discipline. This page documents that method so the league tables stay comparable across cities and sectors, and so any reader (or any business that appears in one) can reproduce the numbers and judge what a score actually means.
The headline metric is the Digital Visibility Score — a 0–100 figure built from six pillars of objective, publicly-observable signals. That page defines the term and shows the current distribution; this one sets out the rules.
The six scoring pillars
Section titled “The six scoring pillars”| Pillar | Weight | Signals (all programmatically gettable) |
|---|---|---|
| Speed & Core Web Vitals | 20% | PageSpeed Insights / CrUX API — LCP, INP, CLS, mobile performance score |
| Technical foundation | 20% | HTTPS, mobile-friendly, indexable, XML sitemap present, robots hygiene, valid structured data |
| Local presence | 20% | Google Business Profile completeness, review count, average rating, review velocity (new reviews / 90 days), website-to-listing domain match |
| Visibility | 15% | Ranking visibility for a fixed local keyword basket; local-pack appearance |
| AI search presence | 15% | Whether the business is named in Perplexity (sonar) answers to a fixed basket of core local queries (see below) |
| Content & trust | 10% | About / team / credentials / blog links present, homepage depth, content freshness (see below) |
Each pillar is scored 0–100, then combined by the weights above into the Digital Visibility Score. Pillar scores are always published alongside the headline number so a single figure never stands alone.
The “AI search presence” pillar
Section titled “The “AI search presence” pillar”This pillar measures whether a business is actually surfaced by an AI assistant for the queries its customers use. For each index we run a fixed basket of five natural-language local queries (for example “Who are the best accountants in Gloucester?”, “I need an accountant in Gloucester — who should I contact?”) and score the percentage of those answers in which the business is named.
One engine, stated plainly: Perplexity, model sonar, via the DataForSEO llm_responses endpoint. An earlier version of this page said the pillar spanned AI Overviews, ChatGPT search, Perplexity and Gemini. It never did — the pipeline has only ever queried Perplexity — and the sentence was corrected on 30 August 2026 rather than quietly widened by adding engines. A score is only defensible if the stated method matches the code that produced it, so the narrower true claim stands until the broader one is actually implemented.
What follows from that: this pillar is evidence about one assistant, not about AI search in general. A business named by Perplexity may be absent from Gemini, and the reverse. Read a zero here as “not named by this assistant, on these five prompts, on the measurement date” — not as “invisible to AI”.
The pillar is re-run every quarter alongside the rest of the index, because AI-search visibility moves faster than classic rankings.
The “Local presence” pillar
Section titled “The “Local presence” pillar”Five components, averaged: Google Business Profile completeness, average rating, lifetime review count, review velocity (new reviews in the last 90 days), and a website-to-listing domain match.
That last component was described as “NAP consistency” until 31 August 2026, and it never was. NAP means name, address and phone; the code compares the domain on the Google listing with the domain of the site being measured, and looks at none of the other three. The claim was corrected rather than the check quietly widened, because widening it would move every Local presence score for a documentation fix. Phone numbers and postcodes are now being captured from both sides so a real comparison can be built and validated against a full quarter of data before it is allowed to change anyone’s score. A component we could not measure is excluded from the average rather than scored as zero.
Two properties of the velocity component are worth stating, because both affect how the number should be read:
-
It is scored against the cohort, not against an absolute. Full marks go to a business at its own index’s 90th percentile for new reviews, so “active” means active for that trade. Ten reviews a quarter is routine for a restaurant and exceptional for a solicitor; scoring both against the same fixed threshold measured the trade rather than the business.
-
The underlying count is censored. We retrieve the twenty most recent reviews per business, so a business with more than twenty in the window reads as twenty. The censoring point sits well above the scoring cap, so it does not affect any score; it does mean the raw count is a floor rather than a true total, and it is not published as a standalone figure.
-
In some cohorts it is excluded entirely. Where three quarters of an index receive no reviews at all — 83% of Cheltenham builders, 88% of Gloucester accountants — there is no distribution to compare against, and the component is dropped for that index. The Local pillar is then computed from its remaining parts.
This is the same rule the rest of the pipeline follows: a signal that cannot be measured is left out, never scored as zero. Normalising a cohort like that would be worse than useless. Scaling to the cohort maximum would score a builder with three reviews at 100 and forty-three builders at 0, inventing a hundred-point spread from a nought-to-three range; ranking by percentile would give the tied zeros a midrank near the 41st percentile, so a business would be rewarded for having no reviews. A builder is not digitally weak for working in a trade whose customers do not leave Google reviews.
Comparing Local presence between sectors. Because velocity is now cohort-relative, the pillar is comparable within an index by construction. Across indices it is closer to comparable than it was, but the other four components are still absolute, so treat cross-sector Local comparisons as indicative rather than exact.
This component did not work for the whole of Q3-2026. A timestamp parsing fault made it return zero for every business, and because the failure was uniform it looked like a sector where nobody reviews anything. It was found, fixed and re-measured on 30 August 2026; the changelog records what moved as a result.
The “Content & trust” pillar
Section titled “The “Content & trust” pillar”Six signals from a single fetch of the homepage, averaged: whether it links to an about, team, credentials and blog or news section; how many visible words the page carries; and how long ago the page said it last changed.
Word count is banded, not continuous — the difference between 900 words and 1,100 is noise, while the difference between 150 and 900 is a business that has written something and one that has not. Freshness is read from JSON-LD dateModified first and the Last-Modified header second, and is available for about half the businesses we measure. Where a site claims neither date, the signal is excluded rather than treated as stale: not publishing a date is not evidence of neglect.
What this pillar no longer claims. An earlier version of this page said it measured indexed page count. It never did — that field was a permanent null, and a true indexed-page count would need Search Console access to each business’s own property, which we do not and should not have. Counting the pages a site publishes is a different quantity, and calling it “indexed” would repeat the error rather than fix it. The claim was removed on 30 August 2026 and the four signals that replaced it are real ones.
This is the pillar that best predicts whether a business can be found at all — it correlates 0.53 with the Visibility pillar, more than any other. It is also only 10% of the score, which is an inconsistency we have written about rather than quietly fixed.
How a business is identified
Section titled “How a business is identified”Each scorecard states, in machine-readable form, which business it is about. That matters because the page is published by a third party: without it, a search engine has only a name and a link to go on, and names are not unique.
Two references are used, and only where the evidence supports them:
- The Google place ID from the Google Business Profile match, on 1,151 of 1,156 businesses. It is withheld in two cases. Where the profile was matched by name and category alone, the identity evidence is weakest. And where a business trades from more than one location, the stored place ID is one branch listing rather than the firm — naming a single branch as the identity of the whole business would be the same error as publishing one of its addresses, which we also refuse to do.
- The Companies House company number, on the 243 businesses where a unique name-and-town match was found on the public register.
Both are published in the open dataset as localMatch and context.companies, so the identification can be checked like any other figure rather than taken on trust. A place ID can change if Google merges or moves a listing; it is refreshed every quarter, so drift corrects itself at the next measurement.
What is deliberately not asserted. No address, no telephone number, and no rating. We hold review counts and average ratings and they are shown on the page, but they are never published as structured data about the business — third-party review markup about a business we do not own is not ours to emit. The only address available to us is a Companies House registered office, which is frequently a firm’s accountant rather than its premises, and 73 of these businesses trade from more than one location, so a single address would be wrong by construction.
The distinction is between saying where this business is authoritatively described and saying what this business is. We do the first only.
The deep scorecard
Section titled “The deep scorecard”A single 0–100 number is thin. Every business in an index gets a diagnostic scorecard, not just a row in a table. The required format:
ACME ESTATE AGENTS — Bristol Estate Agent Digital Visibility Index, Q2 2026Digital Visibility Score: 66 / 100 · Rank 12 of 18 · ▼ down 4 from Q1
PILLAR SCORE SECTOR MEDIAN READINGSpeed & CWV 51 78 ✗ Fails mobile CWV (LCP 4.8s vs 2.9s median)Technical 70 85 ⚠ No LocalBusiness schema; sitemap staleLocal presence 62 74 ⚠ 38 GBP reviews vs leader's 412; 2 in 90dVisibility 74 69 ✓ Above median on local-pack appearanceAI search presence 40 55 ✗ Absent from AI Overviews for 6/8 core termsContent & trust 68 72 ⚠ No team or credentials page; homepage 180 words
KEY FINDINGS (self-contained, attributed, quotable)• Mobile homepage LCP 4.8s — slowest quartile; likely losing mobile conversions• No structured data — invisible as an entity to Google and LLMs• Review velocity flat: 0 new reviews in 90 days vs sector median of 7• Cited by zero AI engines for "estate agents Bristol" — 5 competitors are
VS THE LEADER (Beta & Co, 87)Beta loads in 1.9s, has full schema, 412 reviews, appears in 7/8 AI answers.The gap is almost entirely speed + reviews + structured data — all fixable.
TOP 3 FIXES (ranked by impact)1. Cut mobile LCP below 2.5s (images / render-blocking) → biggest score + ranking lever2. Add LocalBusiness + Review schema → entity visibility for Google and LLMs3. Restart review generation → close the trust + local-pack gapEvery scorecard must include: the score and percentile rank within the sector; a pillar breakdown against the sector median; specific findings with measured values; quarter-on-quarter movement; the competitive gap to the leader; and prioritised, concrete fixes. Depth is deliberate — the more granular and stat-rich each scorecard is, the more useful it is to the business and the more extractable, citable facts it provides to AI search engines.
What counts as a measurement
Section titled “What counts as a measurement”A pillar signal is only used in a published index when it is:
- Publicly observable — gathered from the live site, the public Google Business Profile, public SERPs, or public AI-search answers.
- Machine-measured — produced by a documented tool or API, not a human judgement.
- Dated — every index carries the snapshot date and names the source (“Measured 18 Jun 2026 via PageSpeed Insights API”).
Prior quarters stay published so movement is visible. The most-shared asset each quarter is the movers and fallers table that diffs the current snapshot against the last.
Context data that is recorded but never scored
Section titled “Context data that is recorded but never scored”Three further sources are collected alongside the pillars. None of them carries any weight in the Digital Visibility Score, and none can move a ranking. That separation is deliberate: a published score must never change because a third party enabled an API key, or because a business is small enough to be absent from someone else’s dataset.
A first-party crawl. Each site is crawled from its homepage — obeying robots.txt, one page at a time with a delay, capped by page count, depth and total time, identifying itself as PYCLocalIndexBot with a link to this page. It records click depth to key pages, internal links per page, the share of anchors that name nothing (“click here”, “read more”), and pages listed in the sitemap that were never reached by following links. Orphan counts are only reported when a crawl finished naturally: a crawl stopped by the page cap has unvisited pages by construction, and calling those orphans would be false.
Chrome UX Report (CrUX). The Speed pillar uses PageSpeed lab data — a synthetic run on Google’s hardware. CrUX is the field equivalent: real Chrome users, 28-day p75. Where both exist, the index publishes both.
CrUX has a coverage limit that matters here. Google only reports origins with enough traffic to be statistically meaningful, and most small local firms do not reach it — of the first twelve Gloucester accountancy practices tested, none had CrUX data, while a Bristol estate agency with far more traffic did. Absence therefore correlates with being small. Scoring on it would systematically penalise exactly the businesses least able to change it, which is why it is context and not a pillar.
Companies House. Company number, incorporation date, status and SIC code, matched on an exact normalised name against an active company — anything less confident records no match rather than guessing. This gives each business an official registry identifier rather than a name we matched on, and lets a cohort be read with company age and size in view, which is the fairest answer to “you are comparing a three-person firm with a fifty-person one”. The SIC code also states, from an authoritative source, what a company is registered to do.
Professional-regulator registers (SRA, Gas Safe, Propertymark) are not used. Their published crawl policies disallow the register search endpoints, and this project will not take data it has been asked not to take. Where an official rating is genuinely open — the FSA’s food hygiene ratings, for example — it may be used and will be named.
The sampling frame, and who it cannot see
Section titled “The sampling frame, and who it cannot see”Every cohort is seeded from Google Business Profiles carrying the sector’s own Google categories, inside a radius of a chosen point. Two consequences follow, and both are limits of the method rather than judgements about a business.
A business with no Google Business Profile is structurally absent, not badly scored. It cannot appear in an index at all, however findable it is by other means. This is not hypothetical: three double glazing firms hold first-page positions in the Gloucestershire cohort — one of them position 1 for two separate queries — and have no profile within 40km of the county. They rank; they cannot be measured. An index therefore describes the businesses Google lists, which is a smaller set than the businesses that exist.
The radius has to match how the trade competes, and getting it wrong is invisible in the scores. A vet’s customers come to the premises, so a town radius describes the market. A glazier goes to the customer, so it does not. Cheltenham and Gloucester double glazing were first measured at 10km, and in both cohorts local firms outside the cohort held more of the first page than the firms inside it — 73 slots against 45, and 69 against 40. Nothing in the scores looked wrong; the cohort simply was not the market. They were replaced by a single county cohort, and the four firms the town radii had missed came in at ranks 1, 2, 9 and 12.
Where a cohort is a county rather than a town, the index says so in its name. The check for this is published with each index: heldBySeed against heldByOthers in the landscape data. When firms outside a cohort hold more of its own first page than the cohort does, the frame is wrong.
What this method will not do
Section titled “What this method will not do”These are guardrails, not preferences:
- No subjective quality or trustworthiness claims about the businesses. The indices measure digital presence, never service quality.
- Only objective, reproducible signals. If a skeptic cannot reproduce a number from this page, it does not go in an index.
- A correction process on every page. Each index and scorecard carries a “Spotted an error? Request a correction” link. Corrections are made on verification, and a changelog records them.
- UK GDPR: indices process business (not personal) data from public sources. No personal data is stored beyond what the business already publishes, and correction requests are honoured.
Publication conventions — built to be cited in AI search
Section titled “Publication conventions — built to be cited in AI search”Because the indices are unique, structured datasets, they are engineered so that AI search engines surface and cite them. These conventions are part of the methodology, not an afterthought:
- Branded entities. Each index is named as a proper noun (e.g. “The PYC Bristol Estate Agent Digital Visibility Index”), and the headline metric is consistently called the Digital Visibility Score, so the terms themselves get quoted back.
- Quotable stat sentences. Each index page carries several self-contained, dated, attributed factual sentences — for example: “As of Q2 2026, 61% of Bristol estate agents fail Google’s mobile Core Web Vitals threshold (PYC Digital Visibility Index, measured 18 Jun 2026).” These are written to be lifted verbatim with attribution.
- Question-shaped headings that mirror how people query AI (“Which Bristol estate agent has the fastest website?”).
- Structured data. Each index hub emits
DatasetandItemListschema; each scorecard references the business as aLocalBusiness. - Machine-readable exports. Every index is published as downloadable JSON and CSV, and the site’s
llms.txtpoints AI crawlers at the datasets and at this methodology.
How this connects to the rest of the site
Section titled “How this connects to the rest of the site”- Methodology → applies → Knowledge Base. Each pillar is a measurable application of the local-SEO, technical, and structured-data strategy in the KB.
- Methodology → is analysed with → Glossary. Scoring itself is plain arithmetic — ratios, weighted sums, a clamp — so that any published figure can be recomputed from the open dataset without trusting a model. The statistical methods in the glossary are the toolkit for interrogating results once published, not the scorer.
- Methodology → feeds → Local Indices. Every published index and scorecard is built to this method.
Every change to this method, and every index refresh, is dated in the changelog.
Browse the Local Digital Visibility Indices or read the Knowledge Base for the strategy each pillar measures.