Free XML Sitemap Generator – Crawl Your Site
Discover reachable pages on a public website and turn them into standards-compatible XML you can review before publishing.
A sitemap is an indexation hint, not a list of every URL a crawler can find
The useful question is not “does this URL exist?” but “is this the canonical URL we want a search engine to index?” A clean sitemap contains successful, public, canonical pages with durable value. It excludes redirects, errors, login areas, internal search results, tracking parameters, duplicate filters, staging paths, and pages carrying a noindex directive. Including unwanted URLs creates conflicting signals and wastes the report you receive from search-engine webmaster tools.
This generator starts from one public page, follows same-host HTML links, respects robots.txt, and writes successful final URLs into XML. It cannot know private business rules or whether a reachable page deserves indexing, so the generated file is a reviewable draft rather than something to publish blindly.
How URL discovery works when the site has no existing sitemap
The crawler reads ordinary anchor links from HTML. Navigation, breadcrumbs, category pages, and contextual internal links therefore determine what it can discover. JavaScript-only routes, orphan pages, form results, authenticated areas, and URLs known only to a database will be absent. That absence is useful: if an important page is missing, visitors and conventional crawlers may also struggle to reach it.
Discovery stays on the starting hostname, follows redirects only after validating the destination, and caps the crawl at the selected number of pages. A hostname redirect such as example.com to www.example.com reaches the final start page, but subsequent discovery remains bounded to the original host policy. Review host consistency before publishing.
Remove these URL types from the generated XML
| URL type | Why it should normally be excluded | Preferred action |
|---|---|---|
| 3xx redirect | The sitemap should name the final canonical destination. | Replace it with the final 200 URL. |
| 4xx or 5xx response | The page is missing, blocked, or failing. | Repair it or leave it out until healthy. |
| Parameter duplicate | Sort, session, and tracking values can create infinite variants. | Keep the clean canonical URL only. |
noindex page | The sitemap asks for indexing while the page refuses it. | Choose one consistent instruction. |
| Canonical to another URL | The page declares that another location is authoritative. | List the canonical target. |
| Private or account page | Crawlers cannot access it and users should not find it in search. | Keep it authenticated and out of XML. |
Publish the sitemap where crawlers can verify it consistently
Save the reviewed XML at a stable HTTPS URL, commonly /sitemap.xml. Add its absolute location to robots.txt, then submit that same URL in Google Search Console and Bing Webmaster Tools. The XML must return 200 without login, redirect loops, or an HTML error template. When a sitemap exceeds 50,000 URLs or 50 MB uncompressed, split it into child files and list them in a sitemap index.
Only include a trustworthy lastmod value when it tracks a meaningful content change. Setting every page to today's date on each deployment teaches crawlers that the field is unreliable. This external generator omits fabricated modification dates because an HTML crawl cannot know when the underlying content genuinely changed.
Large or frequently changing sites should generate XML from application state
A database or routing layer knows publication status, canonical paths, locale relationships, and actual update timestamps. It can also emit every valid product or article even when internal links have not yet exposed it. Use this generator for small sites, migrations, and diagnostics; use first-party generation for a large inventory. Compare the database sitemap with the website crawler's title, heading, canonical, and status report to find orphan pages and links pointing to redirects.
XML sitemap questions to resolve before submission
Will a sitemap make a page rank?
No. It helps discovery and reporting but does not replace internal links, useful content, crawl access, or canonical consistency. Search engines decide whether a submitted URL belongs in the index.
Should images, videos, or alternate languages be added?
They can be described with supported sitemap extensions when the page genuinely owns those assets or alternates. Add extensions from first-party data rather than guessing them during a shallow crawl.
Why did the generator find fewer pages than expected?
The missing pages may be orphaned, rendered only by JavaScript, blocked by robots.txt, behind authentication, on another hostname, or beyond the selected crawl limit. Each cause calls for a different fix.
Specifications defining valid XML sitemaps and crawler discovery
The technical claims on this page are drawn from the primary specifications and vendor documentation below.
- Sitemaps XML format Sitemaps.org
- Google Search Central — Build and submit a sitemap Google