# AI crawler access test: can answer engines read these vendors?

> 4 of 4 tool sites reachable from our build machine on 2026-09-20. 0 of 4 restrict any checked AI crawler in robots.txt (policy undetermined for 1 — robots.txt served a non-robots page). 3 of 4 ship an /llms.txt (ollama, lm-studio, open-webui). Median homepage TTFB: jan fastest at 124ms, lm-studio slowest at 301ms.

Human page: https://testedactually.com/local-llm/benchmarks/local-llm-ai-crawler-access/
Mirror: https://testedactually.com/local-llm/benchmarks/local-llm-ai-crawler-access.md
Updated: 2026-09-20

Methodology: https://testedactually.com/local-llm/benchmarks/local-llm-ai-crawler-access/ — script: `scripts/crawler-test.mjs`, n=4 tools × 3 runs, run 2026-09-20.

**AI crawler access test: can answer engines read these vendors? — n=4 tools, 3 runs per tool**

| Tool | TTFB ms (median) | HTML KB | /llms.txt | AI crawler policy | Blocked bots |
| --- | --- | --- | --- | --- | --- |
| [ollama](/local-llm/ollama/) | 262 | 41.8 | yes | unknown | none |
| [lm-studio](/local-llm/lm-studio/) | 301 | 83.7 | yes | open | none |
| [open-webui](/local-llm/open-webui/) | 201 | 40.1 | yes | open | none |
| [jan](/local-llm/jan/) | 124 | 2417.8 | no | open | none |

### Methodology

- Fetch https://<domain>/robots.txt with a plain HTTP GET and a desktop user agent.
- Parse User-agent groups line by line; a bot is blocked only if a group naming it (or the * wildcard) contains a 'Disallow: /' rule.
- Bots checked: GPTBot, ClaudeBot, CCBot, PerplexityBot, Google-Extended, Applebot-Extended, Bytespider.
- Policy: blocked = GPTBot, ClaudeBot, and CCBot all disallowed from /; partial = at least one checked bot disallowed; open = no restrictions; unknown = robots.txt unreachable.
- Fetch https://<domain>/llms.txt; llmsTxt = HTTP 200.
- Fetch the homepage 3 times; TTFB is time to response headers; report the median. HTML size is the decoded body length of the last successful fetch.

### Limitations

- Single machine, single network location — TTFB is directional, not global CDN truth.
- n=3 runs per tool; we report the median, not a distribution.
- Simplified robots.txt parsing: bot names matched case-insensitively, only Disallow rules read, path wildcards beyond 'Disallow: /' not modeled.
- Homepage only; vendor pricing pages may have different crawler policies.
- Snapshot of 2026-09-20; vendors change crawler policies without notice.
