AI SearchHow we research and reviewPublished August 24, 202612 min read

2026 AI Search Readiness Benchmark: 100 B2B SaaS Websites

Original research across 100 B2B SaaS websites on llms.txt adoption, AI crawler access, JSON-LD, schema markup, canonicals, and technical readiness.

2026 AI Search Readiness Benchmark: 100 B2B SaaS Websites

Original Growthract research

Traditional technical foundations are nearly universal. AI-oriented technical signals are not.

Growthract analyzed a frozen sample of 100 B2B SaaS marketing websites across four categories. Basic search infrastructure was overwhelmingly present: titles, canonical tags and XML sitemaps were detected on almost every site where the signal could be determined.

The larger differences appeared in newer AI-oriented publishing conventions and machine-readable entity structure. 74 of 97 sites with a determinable status published an llms.txt file, while only 21 of 97 inspectable homepages exposed SoftwareApplication schema in their initial HTML.

76.3%

Published llms.txt

74 of 97 websites where llms.txt status could be determined.

99.0%

Allowed tested AI search crawlers

97 of 98 sites with known robots evidence allowed each tested search crawler at the root.

78.4%

Homepage JSON-LD

76 of 97 inspectable homepages exposed JSON-LD in the initial HTTP response.

21.6%

SoftwareApplication schema

21 of 97 inspectable homepages exposed SoftwareApplication schema.

What we studied

The benchmark covers 100 B2B SaaS companies, divided evenly into four 25-company groups:

Sales, Marketing & Customer

25 companies

Productivity & Collaboration

25 companies

Data, Analytics & Automation

25 companies

Developer, Cloud & Security

25 companies

The sample was frozen before the final scan. Companies were not removed because their sites blocked or rate-limited our research client. That matters: excluding difficult-to-fetch sites would bias the benchmark toward websites that are easier for automated clients to access.

Coverage

100

companies in frozen sample

97

homepages inspected successfully

3

unavailable to the research client

The five findings that stood out

01

llms.txt adoption was already 76.3%

We detected an llms.txt file on 74 of 97 websites with a known status.

Of those 74 files, 68 met the benchmark's basic structural checks for Markdown headings. That is 91.9% of detected llms.txt files, or 70.1% of all 97 websites where presence could be determined.

We treat llms.txt strictly as an adoption signal. This benchmark does not claim that publishing the file causes rankings, citations or inclusion in an AI answer.

llms.txt present

76.3%

74 / 97 known

Structured among detected files

91.9%

68 / 74 present

For a deeper explanation of the file itself, see our guide to llms.txt and AI search.

02

Robots.txt was rarely the limiting factor for AI search crawlers

For each of the three search-oriented crawler tokens we measured — OAI-SearchBot, Claude-SearchBot and PerplexityBot — 97 of 98 sites with determinable robots evidence allowed root access.

One site in the known denominator blocked the tested search crawler tokens at the root. Two sites were unknown because robots.txt requests were rate-limited during the final collection.

OAI-SearchBot allowed

99%

97 / 98 known

Claude-SearchBot allowed

99%

97 / 98 known

PerplexityBot allowed

99%

97 / 98 known

This is a narrow robots.txt finding. It does not prove that every page is fetchable by every AI system. CDN rules, WAFs, bot mitigation, authentication and application behavior can create additional access layers.

03

Homepage JSON-LD was common, but software-specific schema was not

JSON-LD appeared in the initial HTML of 76 of 97 inspectable homepages.

Organization schema appeared on 70 of 97 homepages. WebSite schema appeared on 39. But SoftwareApplication schema appeared on only 21.

Any homepage JSON-LD

78.4%

76 / 97

Organization

72.2%

70 / 97

WebSite

40.2%

39 / 97

SoftwareApplication

21.6%

21 / 97

Schema is not a guarantee of AI citation. The useful observation here is narrower: many SaaS homepages expose organization-level machine-readable data, while substantially fewer describe the software product itself with SoftwareApplication markup in the initial response.

See our schema markup guide for AEO for implementation context.

04

llms.txt adoption differed substantially by SaaS category

Developer, Cloud & Security companies had the highest observed adoption at 87.5% among sites with a known status.

Productivity & Collaboration was much lower at 56.5%.

Developer, Cloud & Security

87.5%

21 / 24 known

Sales, Marketing & Customer

80%

20 / 25 known

Data, Analytics & Automation

80%

20 / 25 known

Productivity & Collaboration

56.5%

13 / 23 known

The sample is intentionally balanced rather than randomly drawn from the entire SaaS market, so category differences should be read as benchmark observations rather than population estimates.

05

Traditional technical foundations were already close to universal

Titles were present on all 97 inspectable homepages. Canonical tags were detected on 96 of 97. XML sitemaps were detected on 98 of 99 sites where sitemap status could be determined.

Title present

100%

97 / 97

Meta description present

97.9%

95 / 97

H1 present

94.8%

92 / 97

Canonical present

99%

96 / 97

XML sitemap detected

99%

98 / 99 known

That changes the diagnostic question. For sophisticated SaaS websites, the first problem often is not whether a title tag or sitemap exists. The more interesting technical questions are how clearly entities, products and supporting pages are exposed to retrieval systems.

Category comparison

Each category started with 25 companies. Denominators below vary when a specific signal could not be determined.

Categoryllms.txtHomepage JSON-LDOrganization schema
Sales, Marketing & Customer80.0%20/2587.5%21/2475.0%18/24
Productivity & Collaboration56.5%13/2375.0%18/2466.7%16/24
Data, Analytics & Automation80.0%20/2580.0%20/2576.0%19/25
Developer, Cloud & Security87.5%21/2470.8%17/2470.8%17/24

Machine-readable markup dropped on supporting pages

The collector attempted to inspect a small deterministic sample of supporting page types where discoverable: a product or pricing page, a solution or use-case page, an article or resource, and a help, FAQ or documentation page.

Homepage JSON-LD adoption was 78.4%. On the supporting page types we sampled, observed adoption was closer to 60–66%.

Homepage

78.4%

76/97

Product / pricing

66.3%

63/95

Article / resource

63.9%

62/97

Help / FAQ

62.6%

57/91

Solution / use case

60.2%

56/93

This does not mean every supporting page needs the same schema type. It shows that machine-readable markup observed on homepages did not consistently extend across the broader site sample.

What this benchmark suggests for B2B SaaS teams

Do not stop at robots.txt

Search-crawler access was already nearly universal in the known sample. Technical AI readiness requires looking beyond a simple allow/block check.

Treat llms.txt as an adoption signal

The file is now common in this benchmark sample, but its presence should not be presented as a ranking factor or citation guarantee.

Inspect what the initial HTML actually exposes

A browser-rendered page can look complete while machine-readable entity or product information is missing from the initial response.

Audit supporting pages, not only the homepage

Structured-data coverage was lower across product, solution, article and help pages in our deterministic page sample.

Growthract's broader AI search visibility framework separates technical accessibility from actual brand presence in AI answers. They are related problems, but they are not the same metric.

Methodology

Sample

100 intentionally selected B2B SaaS websites, split evenly across four categories. The sample was frozen before the final collection.

Collection date

August 24, 2026.

Data source

Public HTTP responses only. No private analytics, Search Console data or customer data was used.

Model calls

None. The benchmark used deterministic technical checks and made no Gemini, ChatGPT, Claude or other model requests.

Homepage HTML

Signals such as JSON-LD, H1 and canonical markup were measured from the initial HTTP HTML response. JavaScript-injected markup may therefore not be represented.

Crawler access

robots.txt rules were evaluated for root access using the named crawler tokens. This is not a claim that every URL or every infrastructure layer is accessible.

llms.txt

Measured as file presence plus basic Markdown heading structure. It is reported as an adoption signal, not a ranking or citation factor.

Structured data

Schema types were reported when observed in JSON-LD. Absence means the type was not observed by this collector in the inspected initial HTML.

Supporting pages

Up to one discoverable page was sampled for product/pricing, solution/use-case, article/resource and help/FAQ/documentation page types. Random pages were not used to fill missing categories.

Unknown values

Sites were retained when evidence could not be collected. Unknown signals were excluded from that metric's denominator rather than counted as failures.

Limitations

  • This is a deliberately balanced benchmark sample, not a random census of the global B2B SaaS market.
  • The benchmark measures technical signals, not whether a company is actually cited or recommended by an AI assistant.
  • A robots.txt allow rule does not guarantee successful crawling through CDNs, WAFs, bot protection or application-level controls.
  • Initial HTML inspection may not capture structured data injected only after JavaScript execution.
  • llms.txt is reported as an emerging publishing convention. The benchmark does not treat it as a ranking factor.
  • Websites change frequently. These findings are a snapshot of the collection date.

Open research data

Inspect the evidence yourself

We are publishing the domain-level benchmark data, sampled page-level evidence and machine-readable summary alongside this report.

Apply the benchmark

Find the same technical signals on your own SaaS website

Growthract's AI visibility audit separates deterministic technical evidence from actual AI visibility measurement so you can see what is accessible, what is machine-readable and what still needs investigation.

Continue exploring

Explore Growthract’s SEO + AEO approach