XML Sitemap Generator

Turn a list of URLs or paths into a valid sitemap.xml. Every URL is normalised, escaped and checked against the sitemap’s host, duplicates are removed, and big lists are split into files with a sitemap index.

Generator Web & Dev Updated Oct 4, 2026
Learn how this works
How to Use
  1. Paste full URLs, one per line, or paths such as /about with your site in Base URL. You can also drop a .txt or .csv list.
  2. Optionally put a date (2026-09-30) or image file names after a URL on the same line, separated by spaces.
  3. Choose lastmod: none, today’s date for every URL, or the date from each line. Changefreq and priority are optional; Google ignores both.
  4. For a multilingual site, tick hreflang and list each language with its base URL, such as de https://example.com/de/.
  5. Check the status line and Show Work for excluded lines and duplicates, then Copy or Download each file and upload it to the root of your site.
  6. Add the Sitemap: line to robots.txt with the robots.txt Generator.
URLs
for paths; sets the sitemap’s host
one per line
After a URL, add a date (2026-09-30) or image files (hero.jpg), separated by spaces. Lines starting with # are skipped.
Google ignores changefreq and priority; they are written only if you choose a value.
↥
Or drop a .txt or .csv list of URLs (up to 20 MB)
Presets
sitemap.xml
Waiting for URLs
URLs
—
Files
—
Size
—
Dropped
—

Worked Example

Base URL https://example.com and these 7 lines:

/
/about
/blog/hello world
/products?id=7&color=red
/about
/café
/contact#form

Each path is resolved against the base URL. /blog/hello world becomes https://example.com/blog/hello%20world, and /café becomes https://example.com/caf%C3%A9: the é is the two UTF-8 bytes C3 A9. /contact#form loses its fragment, because crawlers ignore everything after #. The second /about is a duplicate of line 2 and is removed.

That leaves 6 URLs in one 494-byte sitemap.xml, far inside the limits of 50,000 URLs and 52,428,800 bytes. With 120,000 URLs the tool would write three files (50,000 + 50,000 + 20,000) and make sitemap.xml an index of them.

The common mistake: writing <loc>https://example.com/products?id=7&color=red</loc> by hand. In XML, & starts an entity, so the parser stops at that character and the whole sitemap is rejected, not just one URL. It must be ?id=7&amp;color=red. The URL a crawler reads back is unchanged.

Show Work

Paste URLs or paths above to see each one normalised, checked and sized against the limits.

Sitemap Rules

Per file
≤ 50,000 URLs and ≤ 52,428,800 bytes
50 MB uncompressed, even if you serve it gzipped
Files needed
files = ⌈URLs ÷ 50,000⌉ (+ 1 index)
120,000 URLs → 3 sitemaps and sitemap.xml as the index
Each URL
absolute, same protocol + host, < 2,048 chars
Percent-encoded (RFC 3986): space → %20, é → %C3%A9
XML escapes
& → &amp;  ' → &apos;  " → &quot;  < → &lt;  > → &gt;
Inside <loc> and every other value
lastmod
YYYY-MM-DD or YYYY-MM-DDThh:mm:ss+00:00
W3C Datetime; a time needs a time zone
Element order
loc, lastmod, changefreq, priority
Only <loc> is required; Google ignores changefreq and priority
hreflang
<xhtml:link rel="alternate" hreflang="de" href="…"/>
Every version lists every version, itself included

Where Sitemaps Came From

Google introduced Sitemaps in June 2005 as a way for site owners to list pages its crawler might otherwise miss, such as pages reached only through search forms or scripts. In November 2006 Google, Yahoo and Microsoft announced joint support for version 0.90 of the protocol and published it at sitemaps.org, which is why the namespace in every sitemap is still http://www.sitemaps.org/schemas/sitemap/0.9. In April 2007 the engines, now joined by Ask.com, added the Sitemap: line to robots.txt so crawlers could discover a sitemap without it being submitted.

The format has barely changed since; extensions added images, video, news and hreflang alternates in their own namespaces. What changed is how much of it is used. Google says it ignores <priority> and <changefreq> and trusts <lastmod> only when it is accurate. In 2022 it dropped most image-extension tags, leaving <image:loc>, and in 2023 it retired the sitemap ping URL in favour of robots.txt and Search Console.

About This Tool

This tool turns a plain list of URLs or paths into a sitemap that follows the sitemaps.org protocol. It resolves paths against your base URL, removes fragments and default ports, percent-encodes spaces and non-ASCII characters, writes international domains in punycode, escapes & and the other XML special characters, and removes duplicates that only differed in those details. Lines on another host or protocol, mailto: links and over-long URLs are excluded, each with the reason in Show Work.

Past 50,000 URLs (or a lower limit you choose) it splits the list into numbered files and writes sitemap.xml as a sitemap index. It can add hreflang alternates for language versions, image entries and lastmod dates. Everything runs in your browser; nothing you paste is uploaded, and no URL is fetched or checked online, so a listed page that returns 404 or a redirect still needs checking on the live site.

Related tools: robots.txt Generator, Meta Tag & Open Graph Generator, and XML Formatter.

Frequently Asked Questions

Does Google use changefreq and priority?

No. Google’s documentation says it ignores the <priority> and <changefreq> values. It does use <lastmod>, but only when the dates are consistently and verifiably accurate, so stamping today’s date on every URL each time you rebuild teaches it to ignore the field. Other crawlers may read changefreq and priority, which is why this tool can still write them.

How many URLs can one sitemap hold?

At most 50,000 URLs and 50 MB (52,428,800 bytes) uncompressed, whichever comes first. A typical <url> entry is 60 to 400 bytes, so the URL limit nearly always binds first; only entries averaging over 1,048 bytes (52,428,800 ÷ 50,000) hit the size limit. Above that, this tool writes sitemap-1.xml, sitemap-2.xml and so on, and makes sitemap.xml an index that lists them. An index can list up to 50,000 sitemaps.

Why is & written as &amp; in the sitemap?

A sitemap is XML, and in XML & starts an entity such as &lt;. A bare & in ?id=7&color=red makes the whole file fail to parse. The sitemap protocol requires the five characters & &apos; " < > to be written as &amp; &apos; &quot; &lt; &gt;. Crawlers read them back as the original characters, so the URL itself does not change.

Can I list http:// and https:// or www and non-www URLs together?

No. Every URL in a sitemap must use the same protocol and host as the sitemap file, so a sitemap at https://example.com/sitemap.xml cannot list http://example.com/ or https://www.example.com/. This tool takes the protocol and host from the base URL (or the first URL), excludes lines that differ and says why. A sitemap in a folder, such as /catalog/sitemap.xml, may only list URLs under /catalog/, which is why this tool places it at the root.

How do search engines find my sitemap?

Add a line such as Sitemap: https://example.com/sitemap.xml anywhere in robots.txt, which every major crawler reads, and submit it in Google Search Console and Bing Webmaster Tools to see processing errors. For a split sitemap, list or submit only the index. Google deprecated its sitemap “ping” URL in 2023, so pinging is no longer a way to notify it.

How do I use the XML Sitemap Generator?

Just pick your options. The answer shows up right away — there is no button to press. Change anything and it updates by itself.

Is it free? Does it work without internet?

Yes to both. It is free with no sign-up, and once the page has loaded it keeps working even with no internet.

Where does my data go?

Nowhere — every calculation runs on your own device. Nothing you enter is uploaded, logged, or stored.

Common Use Cases

Launching a small site

Paste 7 paths with the base URL https://example.com: the duplicate /about and the #form fragment are dropped, leaving 6 URLs in a 494-byte sitemap.xml.

Large product catalogue

120,000 product URLs become sitemap-1.xml and sitemap-2.xml with 50,000 each, sitemap-3.xml with 20,000, and sitemap.xml as the index of the three.

Multilingual site

Two English pages with en, de, fr and x-default versions become 6 <url> entries, each listing all 4 hreflang alternates including itself.

After an HTTPS move

Paste an old export: http:// lines and other hosts are excluded with the reason, so the new sitemap lists only https://example.com URLs.

Image-heavy pages

Write https://example.com/gallery hero.jpg team.webp and the page gets two <image:image> entries resolved to https://example.com/hero.jpg and /team.webp.

Cleaning a crawl export

Lines that differ only by host case, :443 or a #fragment collapse to one URL, and Show Work lists every duplicate with the line it repeats.

Last updated: