LessTools
LESSCMSVisual page editor with a headless API LESSCOMMERCEStore, PIM and orders in one place LESSSEOGoogle visibility, measured daily — coming soon
One account and one invoice for all of them. Discover LessTools →
← Blog
SEO

Sitemaps, robots.txt and indexing

An XML sitemap, robots.txt and the noindex tag decide what Google finds, crawls and shows in search results. Learn how these tools differ, how to set them up correctly, and how to use Search Console to find out why a page is not indexed.

An XML sitemap is a file listing the URLs you want search engines to know about, along with when each was last updated. Together with robots.txt and meta robots tags, it forms the set of signals that tells Google what to crawl and index. These three tools are often confused, and a mistake in one of them can remove an entire section of a site from search results. Below we explain how they work, how to check them and how to fix the most common indexing problems.

Key takeaways

  • An XML sitemap helps search engines discover URLs, but it does not guarantee they will be indexed.
  • robots.txt controls crawling, not indexing, so it is not a way to hide pages from search results.
  • To keep a page out of Google, use a noindex tag and do not block the same page in robots.txt.
  • A sitemap should list only canonical URLs that return a 200 status and are meant to be indexed.
  • The Page indexing report in Google Search Console shows which URLs are indexed and why others are not.

How Google discovers, crawls and indexes a site

The process has three stages. Discovery: Google learns a URL exists from links on other sites, internal links or a sitemap. Crawling: Googlebot fetches the page and renders its content. Indexing: Google analyzes the content and decides whether to store the page in the index that search results come from. Something can go wrong at each stage, and sitemaps, robots.txt and noindex act at different stages.

What an XML sitemap is and what it should contain

A sitemap is an XML file, usually at /sitemap.xml, containing a list of URLs. Each entry has an address (loc) and optionally a last-modified date (lastmod). Google's own documentation says it ignores priority and changefreq, so they are not worth your attention. Google does use lastmod when it is reliable, meaning it only changes when the content really changes.

A simple entry looks like this:

<url>
  <loc>https://example.com/services/seo-audit/</loc>
  <lastmod>2026-08-20</lastmod>
</url>

What belongs in a sitemap

  • Canonical URLs that return a 200 status.
  • Pages you want in search results: services, posts, collection pages, language versions.

What does not belong in a sitemap

  • URLs that redirect, return 404 or carry noindex.
  • Parameter duplicates, such as sorting or filtering URLs.
  • Utility pages like the cart, form thank-you pages or internal search results.

A single sitemap can hold up to 50,000 URLs and 50 MB uncompressed. Larger sites split it into several files and tie them together with a sitemap index.

XML sitemaps on multilingual sites

If your site has several language versions, each should have its own URL, such as /en/services/ next to /uslugi/, and each should appear in the sitemap. Google treats language versions as separate pages, so leaving English URLs out slows their discovery. Relationships between language versions are described with hreflang annotations, placed either in the page code or in the XML sitemap itself. They must be reciprocal: if the Polish page points to the English one, the English page must point back.

robots.txt: what it blocks and what it does not

robots.txt is a text file in the domain root (/robots.txt) that tells crawlers which paths not to crawl. For example:

User-agent: *
Disallow: /cart/
Sitemap: https://example.com/sitemap.xml

The key and frequently misunderstood rule: robots.txt blocks crawling, not indexing. If other pages link to a blocked URL, Google can index it without knowing its content and show just the address. The Sitemap: directive in robots.txt is a simple way to point crawlers to your sitemap.

Meta robots and noindex: keeping a page out of results

If a page should not appear in Google, add <meta name="robots" content="noindex">. Google has to be able to crawl the page to see the tag, so do not also block it in robots.txt. This is one of the most common mistakes: a page blocked in robots.txt and marked noindex can stay in results because the crawler never read the tag.

ToolWhat it affectsWhen to use it
XML sitemapURL discoveryalways, for pages meant to be indexed
robots.txtcrawlingto save crawl resources on utility pages
meta robots noindexindexingto exclude a specific page from results
canonical linkwhich version gets indexedwhen the same content lives at several URLs

How to submit a sitemap in Google Search Console

  1. Verify your domain in Google Search Console.
  2. Open the Sitemaps section and enter your sitemap URL, for example sitemap.xml.
  3. Check that the status reads "Success" and how many URLs were discovered.
  4. After a few days, review the Page indexing report and compare indexed URLs with the number in your sitemap.

To check a single address, use the URL Inspection tool. It shows whether the page is indexed, when it was last crawled and which URL Google chose as canonical.

Common indexing problems and their causes

  • Discovered, currently not indexed: Google knows the URL but has not crawled it yet. Weak internal linking or thin content is often behind it.
  • Crawled, currently not indexed: Google judged the content low value or duplicate. Expanding and consolidating content helps.
  • Duplicate, Google chose a different canonical: several URLs carry similar content. Set a canonical link and reduce duplication.
  • Blocked by robots.txt: check whether the block is intentional.
  • Excluded by noindex tag: a common leftover from staging settings after launch.

Indexing checklist after launch

  1. The XML sitemap is reachable and lists only canonical URLs returning 200.
  2. robots.txt does not block important sections or CSS and JS files.
  3. No noindex tags left over from staging on production.
  4. The sitemap is submitted in Search Console without errors.
  5. Every language version has its own URLs and appears in the sitemap.
  6. Important pages are linked from the menu or content.

How LessCMS handles it

  • The XML sitemap and robots.txt are generated automatically, with no files to edit by hand.
  • New collection entries automatically get a URL and a sitemap entry, as does every language version of a page.
  • Pages are server-side rendered, so crawlers receive full HTML content straight away.
  • You set meta titles and descriptions per page, entry and language. Read more about SEO features and multilingual sites.

Frequently asked questions

Do I need an XML sitemap?

It is not mandatory, but it helps a lot, especially for new sites, large sites and sites with weak internal linking. Small, well-linked sites can be indexed without one, but a sitemap speeds up discovery of new URLs.

How do I check whether my site has a sitemap?

Open your domain with /sitemap.xml at the end, or look for a Sitemap: line in robots.txt. You can also see its status in Google Search Console.

Does robots.txt hide a page from Google?

No. robots.txt blocks crawling, but a blocked URL can still be indexed if other pages link to it. Use a noindex tag to keep a page out of results.

Why is Google not indexing my page even though it is in the sitemap?

A sitemap only announces URLs; Google decides on indexing based on content quality and uniqueness. Check the Page indexing report in Search Console for the specific reason.

How often does Google read my sitemap?

There is no fixed schedule; Google decides on its own. A reliable lastmod date and regular content updates help it pick up changes faster.

Want your sitemap and robots.txt to update themselves? See LessCMS plans and create an account, no card required.

sitemaprobots.txtindexingsearch console

Build a site like this yourself

Visual editor, content via API and SEO built in — on every plan.