Skip to content

Original data · 50 firms · Fetched 2026-07-14

We read the robots.txt of every top-50 UK law firm. Almost none say anything about AI crawlers.

For each firm in a published revenue ranking we fetched /robots.txt and /llms.txt ourselves, saved every raw response, and parsed each file for the 12 AI user-agent tokens that OpenAI, Anthropic, Perplexity, Google, Microsoft, Apple and Common Crawl publish. This is what the files say, and, just as important, what they do not say.

40/44

readable robots.txt files name no AI crawler at all. Silence is the default posture.

1/50

holds an AI crawler to stricter rules than it gives Googlebot: the only reading of robots.txt that counts as an AI-specific restriction.

3/50

serve a genuine llms.txt. 2 more return a 200 that is really an HTML soft-404.

n = 50 firms · 5 indeterminate: behind bot protection or unreachable, so their stance cannot be read · firm list from a published ranking (source) · every raw file retained

What the files actually show

Silence, not blocking, is the norm

Of 44 firms whose robots.txt we could read, 40 (91%) mention none of the 12 AI tokens. 4 name an AI crawler at all, and in 1 of 50 files an AI crawler is held to stricter rules than Googlebot. The dominant stance is not a decision either way.

Naming an AI crawler is not restricting it

Most firms that name AI tokens treat them exactly like mainstream crawlers. In one file the AI tokens sit in the same directive group as Googlebot and Bingbot, so the policy applies to all named crawlers alike; in others the AI tokens get Allow: / with only the same housekeeping exclusions as the default group, which is explicit permission. Per-firm notes appear under the table.

Bot protection reads more like a gate than robots.txt does

5 of 50 firms answered our request for a plain text file with a WAF challenge (Cloudflare, Vercel, Azure, Incapsula) or did not respond. That gate can turn away a compliant AI fetcher regardless of what the robots.txt would have said, and it is not a robots directive.

The search and on-demand tokens are untouched

Where firms name AI tokens at all, they name the training crawlers (GPTBot, ClaudeBot, Google-Extended, CCBot). The tokens that power live answers, OAI-SearchBot, ChatGPT-User, PerplexityBot and Perplexity-User, are almost entirely unaddressed. A rule on a training crawler does not touch retrieval.

llms.txt is barely used, and easy to fake by accident

3 firms serve a genuine llms.txt; 41 return a clean 404. 2 return HTTP 200 for /llms.txt but the body is an ordinary HTML page, a soft-404 that a careless audit would count as present. We do not.

Per-crawler summary

Counted across the 44 readable robots.txt files only. These counts describe the files' syntax shape, token by token; they are not a judgement that a firm restricts AI. "Not mentioned" is not a block; it means the token does not appear, so the site's general rules (or none) apply to it. And a token that is named is usually being permitted, not limited; the stance column in the full table below carries the actual judgement, made against what each file gives Googlebot.

AI crawler directive counts across readable robots.txt files
AI user-agent Allow Block Partial Not mentioned Named at all
GPTBot 1 1 2 40 4
OAI-SearchBot 0 0 0 44 0
ChatGPT-User 0 0 0 44 0
ClaudeBot 1 1 2 40 4
Claude-Web 0 0 0 44 0
anthropic-ai 0 0 0 44 0
PerplexityBot 1 0 1 42 2
Perplexity-User 0 0 0 44 0
Google-Extended 1 1 1 41 3
Bingbot 1 0 2 41 3
CCBot 1 1 0 42 2
Applebot-Extended 0 1 1 42 2

What a block actually prevents

A blocked user-agent token restricts one specific fetch. It does not make a brand "invisible in AI". Models retain what they were trained on before the rule existed, retrieval and training use different tokens, and one assistant often rides another's index. Precise, per-token, this is what each directive in the table above does and does not do.

  • GPTBot is OpenAI's crawler for model training and for GPTBot-driven browsing. Blocking it asks OpenAI not to collect the site for those uses. It does not remove the firm from ChatGPT's existing knowledge, and it does not stop ChatGPT's search feature, which uses OAI-SearchBot.
  • OAI-SearchBot powers results in ChatGPT search. A firm that blocks GPTBot but leaves OAI-SearchBot unmentioned, which is the common pattern here, can still be fetched and cited by ChatGPT search.
  • ChatGPT-User is the live fetch made when a user asks ChatGPT to open a specific link. Blocking it affects on-demand visits, not training.
  • ClaudeBot is Anthropic's active crawler. Claude-Web and anthropic-ai are older tokens Anthropic has published; naming ClaudeBot alone is the current control.
  • PerplexityBot is Perplexity's index crawler; Perplexity-User is its on-demand fetch when a user's question requires opening a page. They are separate decisions.
  • Google-Extended controls use of content for Google's Gemini and Vertex AI, including grounding and training. It does not affect Google Search indexing or eligibility for AI Overviews, both of which follow Googlebot. Blocking Google-Extended does not remove a firm from AI Overviews.
  • Bingbot is Bing's search crawler. Microsoft Copilot draws on the Bing index, so a Bingbot rule reaches Copilot indirectly; there is no separate Copilot opt-out token to set.
  • CCBot is Common Crawl, an open corpus that many models train from at one remove. Blocking it limits that indirect route, not any single assistant.
  • Applebot-Extended opts content out of Apple Intelligence training. Apple's search crawler, Applebot, is a separate token and is not in scope here.

And the rule that sits under all of them: robots.txt is a request, not enforcement. It expresses a preference that a crawler operator may honour or ignore, and honouring it is voluntary and varies by operator. Nothing on this page measures traffic, lost visibility or citation outcomes, and no causal claim of that kind is made. We report the directives as fetched, and only that.

Every firm, every token

All 50 firms in ranked order. Click a column header to sort. The per-token cells describe the file's syntax shape only, not a judgement that the firm restricts AI: allow (named, nothing disallowed), block (named, full disallow), part (named with path rules or an allowlist), n/m (not mentioned), n/a (file unreadable). The judgement is the AI stance column: whether any AI token is held to stricter effective rules than the same file gives Googlebot.

AI crawler directives for all 50 firms
# Firm robots AI stance GPTBotOAI-SearchBotChatGPT-UserClaudeBotClaude-Webanthropic-aiPerplexityBotPerplexity-UserGoogle-ExtendedBingbotCCBotApplebot-Extended llms
1 DLA Piper dlapiper.com WAF n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a WAF
2 Clifford Chance cliffordchance.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
3 A&O Shearman (Allen & Overy) aoshearman.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
4 Freshfields Bruckhaus Deringer freshfields.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
5 Linklaters linklaters.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
6 CMS cms.law read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
7 Herbert Smith Freehills herbertsmithfreehills.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
8 Eversheds Sutherland eversheds-sutherland.com read none part n/m n/m part n/m n/m part n/m part part n/m part 404
9 Clyde & Co clydeco.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
10 Pinsent Masons pinsentmasons.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
11 Slaughter and May slaughterandmay.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
12 Ashurst ashurst.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
13 Gowling WLG gowlingwlg.com read none part n/m n/m part n/m n/m n/m n/m n/m part n/m n/m 404
14 Bryan Cave Leighton Paisner (BCLP) bclplaw.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
15 Addleshaw Goddard addleshawgoddard.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
16 DWF dwfgroup.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
17 Simmons & Simmons simmons-simmons.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m soft
18 Bird & Bird twobirds.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
19 Taylor Wessing taylorwessing.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
20 Womble Bond Dickinson womblebonddickinson.com WAF n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a WAF
21 Osborne Clarke osborneclarke.com read none allow n/m n/m allow n/m n/m allow n/m allow allow allow n/m 404
22 Fieldfisher fieldfisher.com read AI-specific block n/m n/m block n/m n/m n/m n/m block n/m block block 404
23 Macfarlanes macfarlanes.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
24 Kennedys kennedyslaw.com WAF n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a WAF
25 DAC Beachcroft dacbeachcroft.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
26 Irwin Mitchell irwinmitchell.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
27 Withers withersworldwide.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
28 Mishcon de Reya mishcon.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m yes
29 Stephenson Harwood shlegal.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
30 HFW (Holman Fenwick Willan) hfw.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
31 Travers Smith traverssmith.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
32 Watson Farley & Williams wfw.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
33 Shoosmiths shoosmiths.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
34 Charles Russell Speechlys charlesrussellspeechlys.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
35 RPC rpclegal.com WAF n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a soft
36 TLT tltsolicitors.com 404 none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
37 Gateley gateleyplc.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m yes
38 Mills & Reeve mills-reeve.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
39 Trowers & Hamlins trowers.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
40 Knights knightsplc.com down n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a n/a down
41 Hill Dickinson hilldickinson.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
42 Burges Salmon burges-salmon.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
43 Stewarts Law stewartslaw.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
44 Freeths freeths.co.uk read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
45 Keoghs keoghs.co.uk read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
46 Weightmans weightmans.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
47 Penningtons Manches Cooper penningtonslaw.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m yes
48 Brodies brodies.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
49 Foot Anstey footanstey.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404
50 Browne Jacobson brownejacobson.com read none n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m n/m 404

Notes on every firm that names an AI token

  • Eversheds Sutherland no AI-specific restriction Named AI tokens (Applebot-Extended, ClaudeBot, GPTBot, Google-Extended, PerplexityBot) sit in the same directive group as Googlebot; the policy applies to all named crawlers alike. Nothing AI-specific.
  • Gowling WLG no AI-specific restriction Named AI tokens (ClaudeBot, GPTBot) receive rules equal to or more permissive than Googlebot's (explicit permission, not restriction).
  • Osborne Clarke no AI-specific restriction Named AI tokens (CCBot, ClaudeBot, GPTBot, Google-Extended, PerplexityBot) receive rules equal to or more permissive than Googlebot's (explicit permission, not restriction).
  • Fieldfisher AI-specific restriction Applebot-Extended, CCBot, ClaudeBot, GPTBot, Google-Extended held to stricter rules than this file gives Googlebot (full disallow while Googlebot is not equally disallowed).

The 5 firms marked n/a (DLA Piper, Womble Bond Dickinson, Kennedys, RPC, Knights) answered our request with a bot-protection challenge or did not respond, so their robots.txt stance is unknown, not inferred.

Raw responses for every firm are retained at src/data/ai-crawler-index/raw/ in the site repository, and the parsed dataset is src/data/ai-crawler-index.json.

What this means for legal practice

AI assistants are changing how potential clients and business development targets research law firms. But unlike search engine optimisation, which law firms have understood for a decade, there is no equivalent playbook yet for AI visibility. The data here shows why.

Most law firms are invisible to AI

When a prospective client asks ChatGPT or Perplexity for a law firm recommendation in a specific practice area, your firm's name may not appear because the AI model was trained on publicly available text and crawled sources. Unless your firm has third-party coverage (mentions in Legal500, The Lawyer, industry publications, or academic research), or maintains a Wikipedia entry, the models simply have no source material to cite you from.

Your robots.txt decision doesn't matter much yet

Even though only 1 of 50 firms restrict AI crawlers, this reflects practice leadership and privacy-first decision-making, not a working AI exclusion strategy. Most firms will discover that blocking training crawlers (GPTBot, ClaudeBot) is unnecessary if they were never worth crawling in the first place. The real visibility levers sit outside your robots.txt: earned media, thought leadership placement, and data partnerships.

Data partnerships are the new moat

OpenAI has licensed data directly from The New York Times, Wall Street Journal, and Financial Times. Google has similarly licensed access from Reddit and Stack Overflow. These partnerships mean AI systems cite those publishers preferentially, regardless of live crawls. For law firms, this means being featured in Legal500, Chambers & Partners directories, or specialized databases that negotiate AI-era licensing agreements is more valuable than optimising for open-web crawlers.

Mobile and international reach is expanding

ChatGPT reached approximately 900 million weekly active users by February 2026, with Perplexity and other AI assistants growing rapidly across non-English-speaking markets. For firms with international client bases or ambitions, AI visibility in multiple languages and geographies is becoming a growth vector that traditional SEO never addressed systematically.

Other sector indices

The same methodology applied to three more UK sectors: fintech companies, private healthcare providers and accountancy firms. Want to check one domain instead of fifty? Try the free AI crawler check.

Methodology

The firm list

The population is the top 50 firms from Lawyer Mag (lawyermag.co.uk), "Law Firm Rankings: 100 Leading Law Firms in UK in 2025": the top 100 UK law firms ranked primarily by revenue for 2025, of which we take the top 50. We fetched that ranking on 2026-07-14. This is a secondary compilation, not a primary revenue audit. The Lawyer UK 200 and Legal Business LB100 are the primary rankings but are paywalled. Firm-to-domain mapping was done by inspection; a wrong domain would surface below as unreachable, not as a false directive. The two primary rankings in this field, The Lawyer UK 200 and the Legal Business LB100, are the stronger sources but sit behind paywalls, so we used a freely accessible published compilation and state that plainly. The list orders firms by revenue; it is not a ranking of AI readiness, and nothing here should be read as one.

The fetch

On 2026-07-14 we requested https://<domain>/robots.txt and https://<domain>/llms.txt for each firm with curl, following redirects, and saved the exact response body for every firm as provenance. Three firms answered only on their www host, so we recorded the effective URL. Each firm's primary domain was mapped by inspection; a wrong domain would appear below as unreachable, never as a false directive.

Reading a file, and refusing to read a fake one

A response counts as a readable robots.txt only when its body is genuinely a robots file, that is, plain text carrying User-agent, Allow or Disallow lines. Several firms returned HTTP 200, 403 or 429 whose body was an HTML bot-protection challenge (Vercel, Cloudflare, Azure WAF, Incapsula). We record those as unreadable and never parse them for directives, because a challenge page is not a policy. The same discipline applies to llms.txt: a 200 whose body is an ordinary HTML page is a soft-404 and is marked as such, not counted as a genuine file.

How each token was classified

robots.txt was parsed with standard user-agent grouping: consecutive User-agent lines share the rules that follow them, until the next group. For each of the 12 AI tokens we then recorded:

  • allow: the token is named and nothing disallows it.
  • block: the token is named with Disallow: / and no Allow override, a full-site block request.
  • partial: the token is named with some paths disallowed, or an allowlist pattern (Disallow: / alongside Allow: rules that re-permit specific sections).
  • not mentioned: the token does not appear. The catch-all User-agent: * group, if any, still applies to it, but no AI-specific rule was set. This is not a block.

Those four labels describe syntax, deliberately. A token can be "partial" because it receives an allowlist, or because it shares the site's ordinary housekeeping exclusions; neither is evidence that the firm singled out AI. The judgement lives in a separate, derived field.

The stance judgement: stricter than Googlebot, or not

For each firm we derived one field, the AI-specific stance, computed from the archived raw files. The benchmark is Googlebot, the mainstream crawler every firm plainly wants. Following RFC 9309 semantics, each token obeys the group or groups naming it exactly; a token that is not named falls back to the User-agent: * group. We then ask one question: is any AI token's effective rule set clearly stricter than the same file's effective rule set for Googlebot? Bingbot sits in the per-token table because Copilot rides its index, but it is a mainstream search crawler, so it is not part of the AI-specific test.

  • Restricts AI specifically: at least one AI token is fully disallowed while Googlebot is not equally disallowed. This is the only pattern we label a restriction.
  • No AI-specific restriction: the AI tokens are unnamed (they inherit the same generic rules as every other unnamed crawler), or named with rules identical to or more permissive than Googlebot's. A restriction shared equally with Googlebot, or inherited from *, is a general crawling policy, not an AI decision.
  • Indeterminate: we could not read the file, or the differences were neither clearly stricter nor clearly equal. Ambiguity is graded conservatively, never as a restriction.

Limitations

  • This is a single snapshot on 2026-07-14. robots.txt and llms.txt change; the retained raw files fix what we saw on that date.
  • The population is a revenue ranking, not a random or complete sample of UK firms. It says nothing about firms outside the top 50.
  • Firms behind bot protection (5 of 50) have an unreadable stance here. That is a finding about access, not evidence that they block or permit any crawler.
  • robots.txt and llms.txt are voluntary requests. Compliance is at the operator's discretion, so a directive is a stated preference, not an enforced outcome, and we measure no outcome.
  • We checked 12 named AI tokens. Other agents exist (for example Amazonbot, Bytespider, meta-externalagent, which appear in some of these files); they are out of scope and not counted.

The dataset behind this page is public: src/data/ai-crawler-index.json plus one raw file per firm. Anyone can refetch a domain and check our reading against the saved response. If a firm's file has changed since 2026-07-14, that is expected, and a fresh fetch is the way to confirm it. See also how we grade and the reproduction study.