We Audited 30,091 UK Business Websites.
78 Got an A.
Seventy-eight. Out of 30,091 sites we could score, 78 reached an A grade — 0.3%. The median site scored 79. That is not a long tail of stragglers; it is almost every small business in Britain sitting in the same narrow band of mediocrity — and it is the same band in Cornwall, Aberdeenshire and Belfast.
Interactive Explore all 30,091 sites by nation, county, town and trade in What towns are made of →How we measured it
Sample
36,412 independent UK businesses across 11 sectors — food and drink, retail, professional services, health, trade, accommodation and more — in 47 English counties, 1,150 towns, and all four UK nations. Sourced from OpenStreetMap; chains excluded by domain frequency. 30,091 returned a score.
Checks
134 per site — speed, structured data, local SEO, metadata, email authentication, accessibility, AI readiness.
Speed
Cold time-to-first-byte — the first uncached request, which is what a visitor and Googlebot actually pay — plus Google PageSpeed Insights for rendering.
Failures
6,317 sites could not be scored — dead domains, unreachable hosts, bot challenges. They are counted, not silently dropped.
This extends our earlier study — 3,642 UK hospitality sites audited, 2,850 scored — by 956%. Every headline claim from that study was recomputed against the larger set rather than assumed to hold. Where a figure moved, it is shown alongside the original.
1. The fastest websites in Britain are the emptiest ones
Across 30,091 sites, how fast a website answers tells you almost nothing about how good it is. The correlation between time-to-first-byte and overall score is −0.009 — statistically indistinguishable from none at all. Page render time is no better: +0.001.
The reason turns out to be worth the whole study. Sort every site by weight and the pattern is monotonic in the wrong direction:
Each bar is a quarter of the corpus, ordered lightest to heaviest.
Bars measured from 70, not 0 — the whole corpus lives in a narrow band.
The lightest quarter of British business websites — a median of 27 KB — scores 73.7. The heaviest quarter, at 553 KB, scores 81.1. Sites under 40 KB average 72.5 against 79.7 for everything else: a gap of 7.19 points at t=82.4, across 28,318 sites.
The mechanism is not subtle. Page weight correlates with checks passed at +0.375. A 27 KB homepage has no meta description to score, no structured data, no image alt text, no Open Graph tags — not because they are broken, but because there is nothing there. It is not a fast website. It is barely a website.
The empty-shell cohort
1,452 sites (5.1% of the corpus) answer in under 400 ms and weigh under 40 KB. They average 71.2 against 78.8 for everyone else, and of the 1,452, exactly one earned an A. 1276 of them are C or D.
These are the single-page holding sites, the "under construction" placeholders, the template shells nobody finished. They are the fastest websites in the country and they are worth nothing to the businesses that own them.
It also has a face. Guest houses and B&Bs score 74.3 against 77.5 for hotels — and they are the only kind of accommodation below the pack, with apartments at 77.4, campsites at 77.2, hostels at 77.0 and chalets at 76.8. It is tempting to read that as small businesses having worse websites. It is mostly page weight: the median guest house site is 46 KB against 100 KB for hotels, and 47% of them fall under the 40 KB line against 25% of hotels.
Comparing them inside each weight band separates the two explanations. In the thinnest band they are indistinguishable — both are near-empty pages, and the trade stops mattering:
| Page weight | Guest house | Hotel | Gap | t |
|---|---|---|---|---|
| Under 40 KB | 70.2 n=245 | 69.7 n=212 | 0.5 | +0.72 |
| 40 – 80 KB | 77.3 n=87 | 78.3 n=142 | -1.0 | -1.54 |
| 80 – 150 KB | 77.5 n=85 | 80.1 n=169 | -2.6 | -4.67 |
| 150 – 400 KB | 78.0 n=69 | 80.8 n=203 | -2.8 | -4.94 |
| Over 400 KB | 80.7 n=36 | 81.4 n=106 | -0.7 | -1.19 |
Both trades come from one fixed OpenStreetMap query issued identically to every county, so this is not an artefact of where we looked.
Standardising — giving guest houses the hotels' page-weight distribution — closes 1.80 of the 3.15-point gap, so 57% of it is thinness rather than the trade. The remaining 1.35 points is real and survives at matched weight. Worth noting what this is not: the two answer their servers at the same speed — 884 ms against 895 ms, t=0.6 — so this is a page-weight story, not the server story in section 2.
Which explains the strangest number in the dataset. Split the corpus by score and look at how speed relates to quality inside each band:
| Score band | Sites | corr(TTFB, score) | Reads as |
|---|---|---|---|
| Under 70 | 3,903 | +0.267 | slower sites score better |
| 70–74 | 4,137 | +0.044 | no relationship |
| 75–79 | 8,119 | -0.027 | no relationship |
| 80–84 | 10,936 | -0.177 | faster sites score better |
| 85+ | 2,996 | -0.131 | faster sites score better |
Among the worst sites, being slower predicts a better score. That is not a paradox once you know what the fast ones are: at the bottom of the table, speed is a symptom of emptiness. Above 80, the sign flips and speed becomes the thing that separates good from excellent.
The practical version. "Make your site faster" is the most common advice in this industry and, for the average British business website, it is close to useless — a site can be quick and still be a placeholder. Speed only becomes the differentiator once there is something on the page worth loading. Everything below should be read in that order: content first, then speed.
2. Hotels have a server problem. Restaurants have a page problem.
This is the finding the smaller study could not reach. With 1,667 accommodation sites now measured alongside 9,247 eating-out sites, the two behave nothing alike — and they fail for opposite reasons.
| Measure | Eating out | Staying over |
|---|---|---|
| Sites measured | 9,247 | 1,667 |
| Median time to first byte | 827 ms | 886 ms |
| Fastest quartile (p25) TTFB | 462 ms | 470 ms |
| Under Google’s 600 ms bar | 35.1% | 33.2% |
| Median page weight | 125 KB | 91 KB |
| Median first paint | 3631 ms | 3345 ms |
| Median full render (LCP) | 9.5 s | 8.4 s |
The typical accommodation site takes 886 ms to answer against 827 ms for the typical restaurant — 7% longer before a single byte arrives. Only 33.2% clear Google's 600 ms threshold, against 35.1% of restaurants: half the rate.
And yet hotel pages are lighter (91 KB vs 125 KB) and finish rendering sooner (8.4s vs 9.5s). The delay is not the page. It is everything that happens before the page is sent.
3. There is no national quality gap — we checked
Accommodation aside, the obvious question is whether websites are simply better in one nation than another. They are not. Across 30,091 sites the four nations land within one point of each other, and England versus Scotland comes to +0.04 points (t=0.4) — indistinguishable from zero on a sample this size.
| Nation | Sites | Mean score | Median TTFB | Under 600 ms | llms.txt |
|---|---|---|---|---|---|
| England | 24,094 | 77.4 | 860 ms | 33.5% | 21.9% |
| Scotland | 4,194 | 77.3 | 864 ms | 33.3% | 20.1% |
| Wales | 1,352 | 77 | 882 ms | 31.4% | 18.6% |
| N. Ireland | 371 | 76.4 | 818 ms | 32.9% | 15.9% |
Northern Ireland is 371 sites — enough to report, not enough to lead with.
A summary statistic is easy to disbelieve, so here is the whole thing. Every sector, in every nation, shaded by how far it sits from the corpus mean of 77.4:
Blank where a cell holds fewer than 30 sites. Shading is ±1.5 points around the corpus mean.
| Sector | England | Scotland | Wales | N. Ireland |
|---|---|---|---|---|
| Food & drink | 77.0 n=7,795 | 76.6 n=1,024 | 77.1 n=309 | 75.7 n=119 |
| Retail | 77.6 n=5,400 | 78.2 n=861 | 76.3 n=259 | 76.4 n=58 |
| Professional | 78.2 n=3,409 | 77.9 n=802 | 77.3 n=287 | 74.8 n=48 |
| Health | 78.6 n=1,701 | 77.8 n=350 | 77.6 n=109 | · |
| Personal care | 76.9 n=1,490 | 76.1 n=243 | 76.2 n=76 | · |
| Leisure | 77.6 n=1,805 | 76.7 n=299 | 77.9 n=111 | 79.5 n=47 |
| Accommodation | 76.4 n=1,222 | 76.4 n=334 | 77.3 n=61 | · |
Twenty-four cells, and the widest gap between any two is 4.7 points — Northern Irish professional services at 74.8 against Northern Irish leisure at 79.5, both on fewer than fifty sites. Strip those and the entire United Kingdom sits inside three points. England leads in four sectors, Scotland in one, Wales in two, Northern Ireland in one. There is no direction to it.
The pattern worth taking away: quality is remarkably uniform. Mean scores vary by 2.4 points across sectors and barely a point across nations — everyone makes the same mistakes at the same rate, wherever they are and whatever they sell. Response times vary more, but our own measurement of them has proved less reliable than the scores, so we are re-running the affected batches before drawing a geographic conclusion from speed.
Places that are genuinely different
We tested every county and town with enough sites against every metric we collect. Thirteen looked exceptional. Nine of them were not.
The trap is that a place is not a random sample of businesses. Oxfordshire's sites appeared to have the worst security scores in the country — a whole-corpus t of −7.7, the largest single deviation in the dataset. Retested inside each sector it disappears completely: food t=−0.8, retail t=−1.4, professional t=−0.4. Our Oxfordshire sample is 78% food businesses — we ran a dedicated food crawl there — food businesses score badly on security, and the county average inherited it. The county was never the cause.
That distinction matters for everything on this page. A county's mix in this corpus is a fact about which queries we ran where, not about the county: Oxfordshire's real economy is universities, publishing and motorsport. It is why the trade-mix comparisons elsewhere are confined to accommodation and to eating out, the two groups collected with one fixed query issued identically to every county.
Interactive → What towns are made of
The same question, made explorable. Drill from nation to county to town, compare any of the ten metrics side by side, and see which trades a place is actually built from — inside the two universes where the comparison is sound. It carries the sampling caveat in the interface rather than the footnotes: pick All trades and the composition charts switch themselves off.
Open the explorer →
Four survived. Each holds independently in two or more sectors:
Greater London — Time to first byte
984 ms vs 1,126 ms elsewhere
answers faster than the rest of the country · n=1,485
Food t=−3.9 · Retail t=−4.8
Edinburgh — Time to first byte
1,025 ms vs 1,127 ms elsewhere
the only other city that does · n=2,073
Food t=−3.1 · Personal care t=−5.0 · Accommodation t=−3.4
Greater London — Mobile score
9.0 / 10 vs 8.8 / 10 elsewhere
best mobile experience in the UK · n=1,485
Food t=+5.6 · Retail t=+5.7
Cumbria — llms.txt adoption
10% vs 21% elsewhere
half the national rate · n=745
Food t=−5.0 · Accommodation t=−4.2
Two cities answer faster than the rest of Britain, and it is not their business mix — London and Edinburgh hold the advantage inside food, retail, personal care and accommodation separately. That is an infrastructure story: denser hosting, better connectivity, more sites on managed platforms. London also has the best mobile scores in the country, though only in food and retail; its professional-services sites are exactly average.
And Cumbria publishes an llms.txt at half the national rate, in both food and accommodation — the two sectors that make up most of its economy. Of everything we tested, that is the only place where a genuine local deficit shows up rather than a statistical shadow.
Rejected — significant across the corpus, gone within sector
| Oxfordshire | Security | t=-7.7 | 0 of 5 sectors held |
| Cumbria | Technical SEO | t=-4.6 | 0 of 8 sectors held |
| Cumbria | Security | t=-4.5 | 0 of 8 sectors held |
| Oxfordshire | Performance | t=+5 | 0 of 6 sectors held |
| Greater London | Page size | t=+5.3 | 0 of 7 sectors held |
| Isle of Wight | First paint | t=-5.5 | 0 of 4 sectors held |
| Cirencester | AI readiness | t=+4.8 | 0 of 3 sectors held |
4 of 13 survived. If you read a claim that businesses in some county are worse at the web than businesses somewhere else, this is the check that was probably not run — and on our numbers it fails about seven times out of ten.
4. Almost nobody is good, and almost nobody is terrible
86.8% of sites scored between 70 and 89. The distribution is not a bell curve with healthy tails — it is a wall. Getting to a B is easy enough that most agencies clear it; getting to an A is rare enough that 78 sites in 30,091 managed it.
A disclosure. This is our study, run with our own audit engine, published on our own site — and three of the 81 sites scoring A or better belong to us: scrabblewordsfinder.com (91), john.xsoftlimited.com (90) and a staging host for a product we have not launched (90). Roughly one in twenty-seven of the top grades, from one small company.
That is not evidence we are good at this. It is evidence that a site built by people who know the mark scheme scores well against the mark scheme — our own flagship was designed around these 134 checks from the first commit, which is the web-development equivalent of writing the exam and then sitting it. The useful comparison is not us against the corpus; it is the corpus against itself.
The single A+ in this sample is not ours — it belongs to an independent professional-services firm, and it earned it.
The green arc is 0.3% of the circle — a hairline at this scale. That is the finding.
Explore it yourself
Nine sectors, four measures. Pick one and the ranking re-sorts. For the geography — every nation, county and town, against all ten metrics at once — open What towns are made of.
5. The render problem held, and it is not close
Of 29,627 sites where we could measure Largest Contentful Paint, 1401 — 4.7% — showed their main content within the 2.5 seconds Google treats as acceptable. The median took 9.2 seconds.
Server response got worse as the sample grew, driven almost entirely by the accommodation cohort in section 1. Strip those out and the eating-out figure (827 ms) is within 3 ms of what the earlier study reported for the same population — the original measurement held precisely.
Only the first bar clears the threshold. More sites take over fifteen seconds (24.8%) than manage under five (21.4%).
6. The maintenance marker survived a bigger sample
21.4% of sites publish an /llms.txt file
(6,440 of 30,091). They score 5.34 points
higher overall. That comparison is partly circular — three of our 134 checks test
for llms.txt itself — so we removed the entire AI Readiness category and ran it again.
The effect did not weaken with 956% more data — it strengthened, from t=22.5 to t=65.4.
There is one more way a finding like this goes wrong. If llms.txt were simply more common in sectors that score well anyway, the whole effect would be composition rather than signal — Simpson's paradox. So we ran it again inside each sector separately. It holds in every one:
| Sector | Adjusted gap | t |
|---|---|---|
| Retail | +5.78 | 41.6 |
| Trade | +5.08 | 9 |
| Motor | +4.86 | 8.4 |
| Food & drink | +4.86 | 35.4 |
| Personal care | +4.74 | 15.4 |
| Health | +4.45 | 15.4 |
| Leisure | +4.13 | 14.4 |
| Accommodation | +3.82 | 8.1 |
| Professional | +3.57 | 14.7 |
Nine sectors, nine positive gaps, none below t=8. No sector dissents, so the effect is not an artefact of which industries happen to publish the file. The honest reading is unchanged: publishing an llms.txt file does not make a website good. It marks a site someone is actively looking after, and that shows up everywhere else.
7. The sites without HTTPS are still the fast ones
The most counterintuitive result from the earlier study now rests on 14.7× the evidence, and it holds. 1613 sites still serve over plain HTTP. Their median first byte is 468 ms against 876 ms for the 28,506 sites on HTTPS.
There is no mystery in it. These are small, hand-built, static sites on simple hosting — no CMS, no plugin stack, no TLS handshake. They are fast because there is nothing there to be slow. It remains an argument for simplicity, not against encryption: every one of them shows visitors a "Not Secure" warning, and HTTPS has been a ranking signal since 2014.
The gap narrowed as the sample grew — 334 ms in the earlier study against 468 ms here — which is what you would expect if the first 110 were the easiest to find and the simplest of the group.
8. Where the points are actually lost
Categories carry different weights, so comparing raw scores just ranks them by weight. Measured as a share of the points on offer, the picture is clear.
Performance is the largest single pool of points on the board — 12 of 100 — and the industry collects two thirds of it. AI readiness scores worse in percentage terms but is worth a third as much, so it is the cheaper fix, not the more important one.
9. What actually separates an A from a B
78 sites reached an A. 13,854 stopped at B. Comparing the two groups category by category, one gap is three times larger than any other — and it is the one the section above said barely matters in the general case.
Measured as percent of the points available in each category, so a 12-point category and a 4-point one can be compared. A-grade sites earn 10.04 of 12 on Performance against 7.73 for B-grade — every other category differs by under 8%.
So both things are true at once. Speed does not predict quality across the corpus, because the fast tail is full of empty pages. But among sites that have already got the fundamentals right, speed is very nearly the only thing left to compete on.
10. How many are dead, and why we still won't say "dead"
The earlier study reported 313 businesses whose domain no longer resolved, and that number was badly low — the audit engine had been identifying dead domains by matching an error string, and a runtime change made those failures surface differently. Every dead domain after 9 August was silently filed as "unreachable" instead.
That is fixed. Every domain is now resolved independently of the audit engine, nightly, against the system resolver with two public DNS-over-HTTPS services as fallback. The current position:
3,299 of the businesses we track no longer have a working website. That is 9.2% of the sample — one in every 11 — and it is a conservative count, because it excludes anything that has failed for less than a week.
A domain earns the word here after three consecutive nights of total failure — DNS, TCP and HTTP all down — spanning at least seven days. We set that threshold by measuring rather than guessing. Across 217,181 checks, 3,407 domains have hit a three-night failure run and 29 of them later came back: 0.9%.
We originally waited thirty days, on the theory that registrar grace and redemption periods let a lapsed domain reappear weeks later. That is true in principle and, at this scale, rare enough not to justify the delay — a month-old death reported as news is not much use to anyone. One caveat we would rather state than bury: our records begin on 11 August, so a domain that failed on the 19th had two days to come back in the data, not thirty. The 0.9% is a floor. We will re-measure at sixty days and widen the window again if it climbs.
What this study cannot tell you
England dominates the sample.
24,094 of 30,091 sites are English. Scotland (4,194) and Wales (1,352) are large enough to compare; Northern Ireland (371) is reported but should not carry a conclusion on its own.
We cannot explain the accommodation gap.
Section 3 measures it carefully and rules out three explanations. It does not identify the cause, because the cause is not visible in the sites themselves. Treat it as a question we are publishing, not an answer.
Correlation, not causation.
Every relationship here is observational. Sites with llms.txt score better; nothing shown here demonstrates that adding one would raise a score.
Rendering data is third-party.
LCP and first paint come from Google PageSpeed Insights, a lab measurement on a simulated connection — directionally sound, not a substitute for field data from real visitors.
The dead-domain count is known to be low.
See section 10. We consider this a defect in the study, not a footnote.
Businesses without a website are invisible.
The sample is drawn from businesses that listed a website. Those that never had one are absent, which makes this a study of the industry’s web presence, not of the industry.
What we are measuring next
Four things, in the order we can actually do them.
Finish the sites we already hold
2,022 businesses in our lead pool resolve but have never returned a score — 996 in England, 743 in Scotland, 226 in Wales, 57 in Northern Ireland. That is a fortnight of auditing and it is already funded by work done.
Test whether the thin-site effect is our scoring or the web
Section 1 shows lighter pages scoring worse, and we cannot yet separate "these sites have less on them" from "our 134 checks reward having more". Hand-reviewing a sample of the lightest sites against a human judgement of whether they serve their owner is the only way to tell, and it is the single most important open question in this study.
Widen Wales and Northern Ireland properly
See below — the honest answer is that we cannot do this from what we have.
Sweep the corpus for WebMCP, before anyone has it
An experimental AI-readiness check across the 30,091 domains we have already audited, looking for sites that declare tools to agents rather than making them guess. We expect to find almost none, and that is the point: measuring adoption from the first month gives the baseline that made the llms.txt finding possible.
On Wales and Northern Ireland, plainly. We have 1,891 Welsh and 506 Northern Irish businesses on file. Auditing every one that still resolves would take Wales from 1,352 to 1,578 and Northern Ireland from 371 to 428 — an extra 283 sites between them. That is not widening coverage; it is finishing a list.
Genuine coverage needs a fresh harvest aimed at those two nations, and we have not done it. Until we do, treat every Welsh figure here as indicative and every Northern Irish one as a hint. We would rather say that than let four nations in a heading imply four comparable samples.
Experimental
WebMCP, and why it is on the list
Today an AI agent uses a website the way a person would and worse: it screenshots the page, guesses which pixels are the button, and clicks. WebMCP inverts that. A site registers its own capabilities as tools — a name, a plain-English description, a JSON Schema for the inputs, and a function to run — and the agent calls the tool instead of hunting for the form. There is an imperative JavaScript API and a declarative one that annotates ordinary HTML forms.
It is a proposal, not a standard. It comes out of the W3C's Web Machine Learning Community Group, written by engineers at Google and Microsoft, and it is available only behind an origin trial — Chrome from version 149, with a parallel trial in Edge. Chrome's own documentation says it is "under active discussion and subject to change". WebKit has logged concerns across API design, privacy and security, duplication and internationalisation, and has not taken a position. Anyone telling you it shipped in Chrome 146 is repeating a secondary source.
W3C Web ML Community Group
WebMCP on ChromeGoogle — API docs and origin trial
Edge origin trialMicrosoft — registration and scope
WebKit standards positionApple — the objections, in full
We are interested because it is llms.txt again. A proposal with real backing, a plausible story, and no evidence yet that anything consumes it at scale — and the only reason section 6 could say anything useful about llms.txt is that we measured adoption instead of assuming it. A sweep now costs little and gives a baseline. We fully expect the first answer to be "essentially nobody", and a null result measured properly is still a result. What a scan can see is a site declaring tools in the JavaScript it serves; it cannot see tools registered behind a login, and finding the API named is not proof it works.
For agents and developers
Query this data directly
Every figure in this study is available as a machine-readable API, so an AI agent can ask the corpus a question rather than parse a chart. It speaks Model Context Protocol over Streamable HTTP.
https://mcp.xsoftlimited.com/ corpus_stats Size, coverage and stated limits
benchmark_sector Distribution for one vertical
sector_comparison One metric across sectors
check_adoption Practice adoption and its score gap
What it will not return: a stored score for any business other than one you name yourself. This study covers thousands of real UK companies, and publishing a searchable grade for each would mean publishing a quality judgement about a named business from a dataset that — as section 7 sets out — has been wrong before, and was only found to be wrong by checking. Aggregates are safe to publish; verdicts about other people's businesses are not. Full API documentation →
Runs the same 134 checks and returns your percentile against the venues in this study.
Ask for a cut
Want a dataset we have not published?
Everything above is what we thought worth measuring. The corpus holds 134 checks against 30,091 sites across 11 sectors, 47 English counties and 4 nations, and it can answer a great deal more than this page asks of it — a single trade, one region, one check across the whole country, or the raw distribution behind any figure here. If there is a cut you want and we have the data, we will run it and send it back.
Aggregates only, and for the reason given above: we will not publish a grade for a named business other than your own. Researchers and journalists are welcome — tell us what you are trying to find out rather than the query you think you need, and we will say plainly whether this data can support the claim. Sometimes the answer is no.
Contact us about a datasetMethodology notes
Sample drawn from OpenStreetMap business listings across 47 UK counties. Chain domains are excluded by frequency — any domain shared by more than three venues — so a single hotel group cannot skew a cohort.
Time-to-first-byte is the cold measurement: the first request to an unwarmed connection, which is what a first-time visitor and Googlebot experience. Warm-connection figures are roughly half as large and flatter the whole industry.
Percentiles are linearly interpolated. Group differences use Welch's t-test, which does not assume equal variance between groups of very different size.
6,317 sites returned no score and are excluded from distributional statistics but included in every total described as "audited". Scores are never imputed.
Figures frozen 2026-08-14 from a fixed extract. The earlier study (3,642 audited / 2,850 scored, August 2026) remains published unchanged; where these results differ, both numbers are shown rather than the older one being quietly replaced.