robots.txt and sitemap.xml: Set Them Up Right
October 6, 2026
2 readsEvery website has two small text files that quietly influence how search engines treat it: robots.txt and sitemap.xml. They're often confused with each other, often set up wrong, and occasionally responsible for a whole site disappearing from Google. Here's what each one does, and how to set both up safely.
Two files, two different jobs
- robots.txt is a set of instructions for crawlers: "please don't visit these parts of the site."
- sitemap.xml is a list for crawlers: "here are the pages I want you to know about."
One restricts, the other invites. They work together, and they live at the top level of your site, at yourdomain.com/robots.txt and, usually, yourdomain.com/sitemap.xml.
robots.txt: the basics
A robots.txt file is made of simple groups of rules:
User-agent: *
Disallow: /admin
Disallow: /api/
Allow: /
Sitemap: https://example.com/sitemap.xml
- User-agent says which crawler the rules apply to.
*means all of them. - Disallow lists paths the crawler shouldn't fetch.
- Allow makes an exception inside a disallowed area.
- Sitemap points to your sitemap file, which helps search engines find it.
You can generate a correct file in a few clicks with the Robots.txt Generator.
What robots.txt can't do
This is where most of the mistakes come from.
It is not a security measure. The file is public, so anyone can read it, and listing a private folder in it actually advertises that the folder exists. Protect private content with a login, not a Disallow line.
It doesn't remove pages from Google. Disallow stops crawling, not indexing. A blocked page can still show up in results if other sites link to it, just without a description. To keep a page out of search results, use a noindex robots meta tag, and make sure the page is not blocked in robots.txt. Otherwise Google can't crawl it to see the noindex.
Not every crawler obeys it. Well-behaved search engines follow it. Malicious bots ignore it.
The costliest mistake
One line can take an entire site out of search:
User-agent: *
Disallow: /
It's commonly set on a development site to keep it hidden, then accidentally carried over to the live site at launch. If your traffic collapses after a relaunch, check this file first. Also avoid blocking your CSS and JavaScript, because Google needs them to see the page the way a visitor does.
sitemap.xml: the basics
A sitemap is an XML file listing your pages:
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://example.com/</loc>
<lastmod>2026-10-01</lastmod>
</url>
<url>
<loc>https://example.com/blog/my-post</loc>
<lastmod>2026-09-20</lastmod>
</url>
</urlset>
Each <loc> is a page's full address, and <lastmod> is when it last meaningfully changed. A single sitemap can hold up to 50,000 URLs or 50 MB uncompressed. Bigger sites split theirs into several files and list them in a sitemap index.
A small static site can create one with the Sitemap Generator. Many site builders and frameworks produce one automatically.
What belongs in a sitemap
Be selective. Include:
- Pages you want to appear in search
- The canonical version of each page, using the exact address you want indexed. Our Canonical URL Generator helps if you're unsure.
- Pages that return a normal, successful response
Leave out:
- Pages blocked in robots.txt
- Pages marked noindex
- Redirected or broken addresses
- Duplicates, such as the same page with different tracking parameters
- Login, cart, thank-you and admin pages
A sitemap full of redirects and errors wastes crawlers' time and weakens the signal. Check suspicious ones with the Redirect Checker and the HTTP Header Checker.
Submitting your sitemap
- Make sure the file loads in your browser at its address.
- Add the
Sitemap:line to robots.txt. - In Google Search Console, open Sitemaps, enter the path, and submit.
- Check back later. The report shows how many addresses Google found and flags problems.
Submitting doesn't guarantee that every page gets indexed. A sitemap helps Google discover pages. Whether it indexes them depends on quality and relevance.
A safe starting setup
For most small sites:
- A permissive robots.txt that blocks only private areas such as
/admin,/api/and account pages - A
Sitemap:line in that file - A sitemap that lists only your real, indexable pages, kept up to date
- Sitemap submitted in Search Console
That covers the essentials. For a fuller walk-through of what else to check, work through our technical SEO checklist, and once your pages are ready to be found, write a meta description that earns the click.
The short version
robots.txt tells crawlers where not to go. sitemap.xml tells them what to find. Don't use robots.txt for secrecy or for removing pages, double-check it doesn't say Disallow: / on your live site, keep your sitemap limited to clean, canonical pages, and submit it in Search Console.