Website Crawler – Titles, Headings & Status

Diagram: A crawler reading linked pages, headings, and response status — illustrating website crawler
A crawler reading linked pages, headings, and response status.

See how a search crawler reads a small public website without installing software or connecting an analytics account.

The crawl table combines transport, document, and page-outline evidence

Each row answers a sequence of questions. Did the URL return a usable HTTP response? How long did the request take? Did the HTML expose a title and one clear primary heading? How many links could a conventional parser discover? A 200 status alone is not enough: a server can return an empty template, a login wall, or a “not found” message with 200. Conversely, a short title is not automatically a defect when it describes the page precisely.

Interpret status codes before editing titles or headings

Observed statusWhat it means in this crawlWhat to investigate
200–299The server returned a successful response.Review content, canonical, title, and headings next.
301 or 308The resource moved permanently.Update internal links to the final destination and remove chains.
302 or 307The move is declared temporary.Confirm that temporary behavior is intentional.
401 or 403Authentication or access policy blocked the crawler.Decide whether the page should be public and indexable.
404 or 410The resource is missing or intentionally gone.Repair inbound links or provide a genuine replacement.
429The server rate-limited requests.Reduce crawl volume and inspect bot or firewall rules.
500–599The origin or gateway failed.Use server logs and request IDs to find the failing dependency.

Title and H1 should agree on the page's primary purpose

The title is used in browser tabs and often influences a search-result headline. The H1 introduces the same page to the reader. They do not have to be character-for-character identical, but a crawler that finds “Cheap VPN Deals” in the title and “Network Security Guide” in the H1 has found two competing intents. Align the subject and promise, then use H2s to explain distinct questions within that subject.

Missing and duplicate titles deserve priority because they make multiple URLs indistinguishable. An H1 count above one can be legitimate in modern HTML, yet on a templated site it frequently reveals a logo, modal, imported article, or result value marked as the primary heading. Inspect the rendered HTML rather than only the content-management editor.

Use H2s to expose the information path, not reusable labels

Headings such as “Overview,” “Benefits,” “More Information,” and “Frequently Asked Questions” say little when repeated across hundreds of pages. A useful H2 tells the reader what the section resolves—for example, “DMARC alignment can fail even when SPF passes.” Scan the H2 column across a crawl and look for a template phrase dominating unrelated URLs. Replace the footprint with headings tied to the page's actual subquestions.

Heading levels should also form an understandable outline. An H3 should normally refine the H2 above it rather than appear because a component's font size looked right. Fix semantics in the template so the improvement propagates consistently.

Canonical and internal-link findings reveal duplicate or orphaned paths

A self-referencing canonical is a useful default for an indexable page. A canonical pointing elsewhere can be correct for a duplicate, but internal links should usually point directly to the preferred destination. Compare final response URLs, canonical values, and navigation links to uncover HTTP/HTTPS duplication, trailing-slash variants, uppercase paths, parameters, and old routes.

A page with few discovered links is not necessarily orphaned—the current table only counts outgoing links—but it deserves context. To prove an orphan, compare the full link graph or application inventory and count inbound references.

This crawler reads server HTML and deliberately does not execute JavaScript

The tool follows same-host anchors found in the returned document, respects robots.txt, validates every redirect, and caps pages, response bytes, request time, and concurrency. It does not log in, submit forms, run a browser, or render client-side frameworks. If a React or Vue application sends an empty shell and adds links later, the report describes that server response accurately but cannot show the hydrated interface.

Use a browser-based crawler for JavaScript rendering and a local crawler for a large site. This free scan is intended to expose a representative structural problem quickly without becoming an unrestricted scraping service. When the sample reveals failed destinations, run the broken link checker to retain each failed target's source page.

A practical order for turning crawl findings into fixes

  1. Resolve 5xx failures and broken canonical destinations.
  2. Replace internal links to redirects with their final URLs.
  3. Repair missing or duplicated titles and primary headings.
  4. Clarify vague, repeated section headings.
  5. Investigate slow outliers with server timing and application logs.
  6. Re-crawl the same starting URL and compare the changed rows.

Website crawl questions that affect the interpretation

Why is a JavaScript page nearly empty?

The server likely returned an application shell. Search engines may render it later, but server rendering or meaningful pre-rendered HTML gives crawlers and users a more dependable first response.

Does one missing meta description prevent indexing?

No. Search engines can generate snippets from page copy. A useful unique description still improves control and exposes templating gaps when many pages lack one.

Why does the report stop before the whole site?

The selected page cap, robots rules, hostname boundaries, missing links, errors, and JavaScript-only navigation all limit discovery. The displayed count is a bounded sample, not a claim of complete coverage.

Crawler and search-documentation behind the reported page signals

The technical claims on this page are drawn from the primary specifications and vendor documentation below.

  1. RFC 9110 — HTTP Semantics RFC Editor
  2. Robots Exclusion Protocol RFC Editor