← All insights Case Studies

A sitemap is a hint. A link is an instruction.

Our pages sat in the sitemap for weeks and no AI crawler ever fetched one. We added a single link from a page they already read. Four fetches followed the same afternoon.

· July 28, 2026 · 4 min read · 6 views
A sitemap is a hint. A link is an instruction.

A sitemap tells a crawler a page exists. A link tells it the page is worth fetching. Those are not the same instruction, and the difference is measurable — we measured it on our own site, by accident.

Key facts
  • Our business pages sat in sitemap.xml for weeks. AI crawlers visited the site repeatedly in that time and fetched none of them.
  • Nothing on the public site linked to those pages. They were reachable only by knowing the URL.
  • We added one link, from a page crawlers already fetch, to an index of those pages.
  • Within the same day: ClaudeBot twice, GoogleOther twice — the first time any assistant's crawler had read one.

What we found in the log

We had been checking weekly what ChatGPT, Claude and Gemini said about a business, and improving the page they were supposed to be reading. Then we looked at the access log to see how often that page was actually being fetched.

The answer was never. Not slowly, not occasionally — zero, over weeks, while the same crawlers fetched the home page, the terms and the pricing anchor on a regular cycle.

The pages were in the sitemap. The sitemap was in robots.txt. Both were being read. Neither produced a single fetch of the pages they described.

Why a sitemap is not enough

A sitemap is a submission, and a crawler decides what to do with it. Search engines document it as a hint that supplements link discovery rather than replacing it, and they are explicit that inclusion does not guarantee crawling. A URL with no inbound link has nothing else recommending it: no context, no anchor text, no page vouching for it.

Links carry information a sitemap entry cannot. Where the link sits, what it says, and what else that page links to are all signals about whether the destination is worth the fetch.

The change, and what happened

We built a plain index — every business with a published page, name and trade, one link each — and put a link to it in the footer of every public page. Nothing clever. The point was that a link should exist at all, from somewhere already being crawled.

The log for the following hours, on pages that had never been fetched before:

Four fetches inside three hours, after weeks of none. We had changed one thing.

What this does not prove

It does not prove the pages will be quoted, or that anyone's ranking moves. Being fetched is necessary, not sufficient — and one site over one day is an observation, not a study. Crawl schedules vary and the timing could have been coincidence, though four fetches from two operators inside three hours of the change is a large coincidence.

What it does establish is the order of operations. A page that has never been fetched cannot be quoted by anything, however good it is.

What to check on your own site

Common questions

Should I stop maintaining a sitemap?

No. It helps with discovery of pages that are legitimately hard to reach, it carries change dates, and search consoles use it for reporting. It is a useful supplement. The mistake is treating it as the whole discovery mechanism.

How many links does a page need?

One that is itself reachable is the difference between zero and non-zero. Beyond that it becomes a question of how often the linking page is crawled and how prominent the link is, not of accumulating a count.

Does an internal link count, or does it need to be external?

Internal links are what got these pages fetched. External links carry more weight for ranking, but for discovery an internal link from a regularly crawled page does the job.

How long should I wait before concluding it is not working?

Check the access log rather than waiting on a feeling. If a fortnight passes with the page in the sitemap and no fetch of it in the log, discovery is the problem, not content.

Which crawler fetching the page actually matters?

It depends which assistant you care about. ClaudeBot feeds Anthropic, GPTBot and OAI-SearchBot feed OpenAI, and GoogleOther is Google's crawler for things other than web search. Seeing any of them is better than seeing none, and seeing the one behind the assistant your customers use is the one that counts.

Sources

The fetch timings are from our own nginx access log on 28 July 2026 and are a narubox measurement on a single site, not a controlled study. narubox is not affiliated with Google, OpenAI or Anthropic.

Does AI know your business?
Find out what it says about you today.

Check my business