An orphan page is a page on your site that no internal link points to. It loads, it may rank, and nothing on the site leads anyone to it. The awkward part is that the tool most people reach for, a link crawler, is structurally incapable of finding one by itself. Finding orphans means comparing a crawl against a second list of URLs.
What an Orphan Page Actually Is
The strict definition: a URL on your domain that zero crawlable pages on the same domain link to with an ordinary <a href>. It is a property of your internal link graph, and nothing else.
A few things get counted as orphans that are worth keeping separate:
- Pages linked only from a noindexed page. The link still exists and still passes signals. Google can follow it.
- Pages linked only from a URL blocked in robots.txt. The link is real but no crawler will read it, so the page behaves like an orphan.
- Pages reachable only through a form, a search box or a JavaScript click handler with no href. Functionally orphaned for a crawler.
Two things never count as internal links here: an entry in your XML sitemap, and a backlink from another domain. Both help discovery. Neither puts the page into your site's structure.
Why Orphan Pages Happen
Orphans are almost always a side effect of a change that moved or removed whatever was doing the linking.
- CMS migrations. Content gets ported, URL patterns change, templates get rebuilt. Some pages survive the move with their content intact and no slot in the new navigation.
- Deleted category and tag pages. The biggest single cause on content sites. If
/category/widgets/was the only page listing 60 posts, deleting it orphans all 60 at once. - Products pulled from navigation. Out of stock, discontinued or seasonal items get removed from menus and listings but stay live at their URL, often for years.
- Truncated pagination. An archive that caps at page 5, or a "load more" button that fetches results by JavaScript with no real link, cuts off everything past the first batch.
- Campaign landing pages. Built for an email or an ad, published outside the normal content structure, linked from nothing.
- Sitemap-only pages. Most sitemap plugins write every published URL into the XML file regardless of what links to it, so the sitemap keeps the page visible to Google while the site does not.
Notice the pattern: the orphan is the symptom. The cause is a missing hub, a broken template, or a migration that dropped links.
Why a Crawler Alone Can Never Find Them
This is the part most articles get wrong, and it changes the method.
A link crawler works by traversal. It fetches a seed URL, parses the HTML, extracts every href, queues the ones it has not seen, fetches those, and repeats until the queue empties. Its entire universe is the set of URLs reachable from the seeds by following links.
A page that nothing links to is never queued, never fetched, and never present in the crawl output. So it can never appear in an orphan report built from that crawl alone, because from the crawler's point of view the page does not exist. That follows from the definition, so no crawler escapes it.
The orphan set is a set difference:
Orphans = (URLs known to exist) minus (URLs reachable by crawling)
That first list has to come from somewhere other than link traversal: your XML sitemap, Google Search Console, analytics, server access logs, or a direct export from the CMS database. When a crawler does report orphans, it is doing exactly this. It ingested a sitemap or a pasted URL list, and it is showing you URLs from that source with zero recorded inlinks. Which source supplied them sets the boundary of what you can possibly find.
Why Orphan Pages Matter
Google's sitemap documentation is blunt about the limits of a sitemap as a discovery mechanism: "A sitemap helps search engines discover URLs on your site, but it doesn't guarantee that all the items in your sitemap will be crawled and indexed." The same page treats proper linking as the baseline, meaning every page you consider important can be reached through some form of navigation (Google Search Central, Sitemaps overview).
An orphan page is relying on the weaker of those two signals. Beyond discovery, the practical costs are:
- No internal link equity. Whatever authority your site has, none of it flows to a page with zero internal links.
- Low crawl priority. On large sites, pages nothing links to get revisited rarely, so content updates take a long time to register.
- No human path to the page. Visitors arrive from search, a bookmark or a paid click, with no route to it from anywhere else on your site.
Keep the scale honest. There is no orphan page penalty, and a page with zero internal links can still rank if it is good and something external points at it. The cost is that the page is weakly discovered and unsupported, and that whatever caused it probably affects more than the pages you found.
When an Orphan Page Is Completely Fine
Plenty of pages have no business in your navigation, and linking them would make the site worse.
- Thank-you pages, order confirmations, cart and checkout steps.
- Paid landing pages built for one ad group, deliberately kept out of the site structure so organic traffic goes to the main page instead.
- Unsubscribe pages, account pages and anything behind a login.
- Print views and other alternate renderings of a canonical page.
The test is one question: does this page need organic traffic? If no, being an orphan is the correct state, and the work is making that intent explicit rather than adding links.
That means two things. Add noindex if the page should stay out of the index, and remove it from your XML sitemap. A page that is deliberately unlinked and also listed in the sitemap sends contradictory signals, and it will resurface in every audit you run from now until someone fixes it.
Step by Step: Comparing a Crawl Against a Sitemap
This is the core workflow. It finds every orphan present in your sitemap, which on most sites is the majority of them.
- Run a complete crawl from the homepage. No URL cap, no depth cap, no time cap. Enable JavaScript rendering if any part of your navigation or listings is built client side. A truncated crawl manufactures false orphans, so completeness matters more here than in most audits.
- Export the crawled URL list with status code, canonical, indexability and inlink count as columns, since you will need those later.
- Expand the sitemap into a flat URL list. Start at
/sitemap.xml, and if it is a sitemap index, fetch every child sitemap it references. Watch for gzipped files and for sitemaps declared inrobots.txtbut absent from the index. Extract every<loc>value. - Normalise both lists before comparing them. This is where most comparisons produce garbage. Pick one canonical form and apply it to both sides: consistent scheme, consistent www or non-www host, consistent trailing slash, no default port, no tracking parameters, consistent percent-encoding, and a decision about
index.htmlsuffixes. Then sort and deduplicate. - Take the difference in both directions. Sitemap minus crawl gives your orphan candidates. Crawl minus sitemap is worth a look too: pages you link to internally that the sitemap omits, often the same bug seen from the other side. With two sorted text files,
comm -13 crawl.txt sitemap.txtgives you the first set. - Verify each candidate. Request the URL and check its status code, canonical target and robots meta. Then search your crawl's link data for that URL as a link destination. Zero results confirms it. One result means your normalisation missed something.
- Re-crawl with the confirmed orphans added as seeds. That gives you titles, word counts, status codes and outbound links for each one, which is what you need to decide in the next section.
Adding Search Console, Analytics and Log Files
The sitemap comparison has one blind spot: it only finds orphans that are in the sitemap. Pages that are unlinked and missing from the sitemap need another source, and each source has a different bias.
- Google Search Console. The Pages report shows indexed URLs, and the Performance report shows every URL that received an impression. The interface caps exports at 1,000 rows, so use the Search Console API for a larger site.
- Analytics. A landing page report over 12 months. Any URL that received a session exists. This misses orphans with no traffic at all, which is a real category.
- Server access logs. The most complete inventory you will get. Filter to successful HTML responses, strip query strings, deduplicate, and use a window of at least three months. Expect bot noise and requests for URLs that never existed.
- A CMS or database export. For many sites this is the cleanest answer to "what pages are published?", and it takes one query.
Union the lists, normalise them the same way, deduplicate, then subtract your crawl. That is the full candidate set.
False Positives to Rule Out Before You Act
Long orphan lists usually contain pages that are linked perfectly well. Check these before touching anything:
- Normalisation mismatch. A protocol, host or trailing slash difference between the two lists flags thousands of healthy pages at once. A suspiciously large orphan count is almost always this.
- JavaScript-rendered links. If you crawled raw HTML and the site builds its navigation, category listings or related posts client side, everything downstream of those links looks orphaned. Re-crawl with rendering enabled.
- Links inside pages the crawler never parsed. A robots.txt block, a 500, a timeout or a rate-limit response loses that page's links for the whole run.
- A crawl that stopped early. URL limits, depth limits and time limits truncate the reachable set and turn the remainder into fake orphans.
- Second-order orphans. A page linked only from another orphan. Fix the parent and the child resolves itself.
Confirm the crawl was complete and that rendering matched how the site builds its links. Then believe the list.
Deciding Per Page: Link, Redirect, or Delete
Three questions per page: does it have traffic or impressions, is the content useful and distinct, and is there an obvious place for it in the site structure?
Link it when it earns impressions or conversions, or it is useful and has a sensible parent. Add a real link from a topically related page, ideally more than one, with anchor text that describes the destination.
Redirect it when the content duplicates or is superseded by a stronger page, or when the URL is a leftover variant from a migration. Use a 301 to the closest equivalent. Avoid bulk-redirecting unrelated pages to the homepage, which Google commonly treats as a soft 404.
Delete it when the page is thin, obsolete, has no traffic, no backlinks and nothing to consolidate into. Return 410 rather than 404 to signal that the removal was deliberate.
Leave it alone when it belongs to the functional category above. Confirm it is noindexed and absent from the sitemap, and record the decision so the next audit skips it.
Before any of that: if you have several thousand orphans, you have a systemic cause rather than several thousand decisions. Sort the list by directory, template, content type and publication date. Orphans cluster. One restored category page or one fixed pagination template usually relinks hundreds of URLs at once, and what remains is short enough to judge by hand.
Where to Put the Link So It Counts
- Rebuild the hub. Category pages, tag archives, product listing pages and topic hubs are what relink content at scale. Restoring one deleted category page beats a week of manual linking.
- Fix pagination properly. Every archive page needs a real
<a href>pointing to it, and pagination needs to run through to the final page. A "load more" button with no underlying link leaves everything after the first batch unreachable. - Check that related-content modules output real links. Some recommendation widgets render as JavaScript click handlers with no href at all.
- Verify in the rendered HTML. Confirm the new link appears in what a crawler sees, not just in the CMS preview. An HTML sitemap page works as a backstop, though a link buried in a list of 2,000 passes very little.
Running This in LibreCrawl
LibreCrawl is free, MIT licensed and self-hosted, and several of its properties matter specifically for orphan work.
No URL cap. Orphan detection depends on a complete crawl. A crawl that stops at 500 URLs reports every uncrawled page as an orphan, which is a convincing way to waste an afternoon. LibreCrawl crawls unlimited URLs, so the reachable set in your export is the real one. If you are moving off a tool with a licensed URL limit, the Screaming Frog alternative comparison covers the differences.
Sitemap ingestion during the crawl. LibreCrawl can pull in sitemaps as part of a crawl, so both sides of the comparison come out of one run and share the same URL normalisation. That removes the step where most manual comparisons go wrong.
A full internal link graph with "linked from" data per URL. This is the evidence for step six. For any candidate you can read its linked-from list directly. Zero entries confirms the orphan. A single entry from a page you had forgotten about saves you a pointless fix.
JavaScript rendering via Playwright. Client-side navigation is the most common source of false orphans, and rendering the page the way a browser does removes that whole category of error.
The site structure visualisation. LibreCrawl's graph view is built on Cytoscape.js, with pages as nodes and internal links as edges. Near-orphans stand out as nodes hanging off the structure by a single thin thread, and so do clusters of pages that link only to each other. That view tells you whether you are looking at 40 separate problems or one broken template. The site structure visualisation post covers the layouts and filters.
After the fixes ship, re-crawl the same URLs. The confirmation you want is a linked-from list with entries in it.
Keeping Orphans From Coming Back
Orphans are generated by ordinary site changes, so the prevention work belongs in the process rather than in a one-off cleanup.
- Before deleting any category, tag or listing page, export its outbound links and find a new home for every one of them.
- Run the crawl-versus-sitemap comparison as a release check around every migration, once before and once after, and diff the two counts.
- Configure the sitemap generator to list what you actually want indexed. A sitemap that mirrors the database creates permanent audit noise.
- Keep the orphan check on the recurring schedule in your technical SEO audit checklist, monthly for large sites and quarterly for small ones.
The cleanup is a day of work. The prevention is a checklist item that costs ten minutes per release.
Find Your Orphan Pages with LibreCrawl
Unlimited crawling, sitemap ingestion, a full internal link graph and an interactive structure view. Free forever, MIT licensed, self-hosted.
Download LibreCrawl