katman / Observation log
TRANSPARENCY
Observation log
Dated, verifiable measurements made with the open-source katman-audit engine. We publish our misses as well as our wins — that is the whole point of a log. Treat every row as one dated observation, never as a verdict: crawler policies and model behavior change.
| Date | Site | Version | Result | Notes |
|---|---|---|---|---|
| 2026-08-30 | katman.pro | 0.1.0 | 1 CRITICAL | Cloudflare managed robots.txt silently blocked GPTBot, ClaudeBot, Google-Extended, CCBot and Bytespider on our own domain. Caught before public launch. → |
| 2026-08-30 | katman.pro | 0.1.0 | 1 WARN | BreadcrumbList schema missing across sampled pages — a regression from our own page-generator migration. Caught by the same CLI, fixed the same day. |
| 2026-08-30 | katman.pro | 0.1.0 | 0 crit · 0 warn | Green run after fixes. Exit code 0. Current baseline. |
| 2026-08-30 | frutti.ai | 0.1.0 | 1 WARN | robots prefix trap: /r/ does not cover the bare /r path. |
| 2026-08-30 | nytimes.com | 0.1.0 | 1 CRITICAL | GPTBot + OAI-SearchBot blocked in robots.txt. Editorial policy, presumably intentional — the tool reports, it does not judge. |
| 2026-08-30 | 22 major sites | 0.1.0 | survey | Dated survey of who blocks AI crawlers: 3/19 verifiable sites block at least one major crawler; llms.txt adoption 9/22. → |
Add your own
Run npx github:ahmetrnn/katman-audit yoursite.com, save the JSON output, and open a PR appending a row to the log in the repository. Observations only — one request per file, public pages, no judgments about intent.
Start with your own domain
The fastest observation is the one about your own site. If it comes back red, you just saved months of invisible content investment.