An SEO crawler requests the pages of a website the way a search engine would, follows the links it finds, and records what each response contains: status code, title, meta description, headings, canonical tag, and the pages that link to it. A free SEO crawler does the same work at no cost, but the word "free" covers four different arrangements here, and the differences decide whether a tool is still useful to you in week two.

This is a guide to the category rather than a pitch. We build one of the tools below, LibreCrawl, and we say where something else is the better answer. If you already want an uncapped replacement for a capped desktop tool, our Screaming Frog alternative page covers that comparison directly.

What an SEO Crawler Actually Does

A crawl starts from a seed URL, usually your homepage. The crawler requests it, reads the HTML, extracts every link, and queues the ones on the same site. It works through that queue until no unvisited URLs remain, checking robots.txt, following redirects and recording where they land, and respecting the concurrency and delay settings you set.

The value is in what gets stored for each URL. A capable SEO crawler tool captures the status code, response time, title, meta description, headings, canonical URL, meta robots directives, hreflang annotations, word count, indexability, and the internal links in and out. It keeps the relationships too, which turns a list of URLs into a link graph: for any page you can ask what links to it, and how many clicks from the homepage it sits.

On top of that data, crawlers apply checks and present the results as issues: broken internal links, redirect chains, pages blocked from indexing that you meant to be indexed, duplicate titles, orphan pages. Those checks are rules over collected fields rather than magic. A crawler sees only your site as served, so backlinks and rankings need other tools.

What "Free" Means in This Category

This distinction is the most useful thing on the page, because all four arrangements below get marketed with the same word and behave very differently once a site grows past a few hundred URLs.

A free tier is a permanent free version of a paid product, capped below the point where a professional would depend on it. Screaming Frog's free version is the clearest example: it never expires, and it stops at 500 URLs per crawl. A free trial gives you everything on a clock, usually 7 to 14 days, which makes it good for evaluation and poor as a plan.

Open source software is free in a different sense: no paid tier to be nudged toward, because no vendor has a revenue target attached to your usage. The cost moves from money to your time, since you install it and keep it running. An open source SEO crawler usually applies no URL cap, because there is nothing for a cap to protect.

Freemium hosted tools are the arrangement to watch. The crawl is free and often fast, then the part you needed sits behind the upgrade: the full issue list rather than a summary, the export, the second project, the history that compares months. A free website crawler that blocks exports has handed you a screenshot rather than data.

Arrangement What you get Where it runs out
Free tier of a paid tool A real crawler, permanently, with a hard cap The URL cap, and features locked to the paid version
Free trial Everything, briefly The clock, typically 7 to 14 days
Open source, self-hosted All features, unlimited crawling, source you can read Your time: setup, updates, hardware
Freemium hosted tool A crawl with no installation Exports, projects and detail behind a paywall
Search engine webmaster tools Free reporting on sites you have verified Your own properties, and limited on-demand crawling

What to Check Before You Commit to One

Five things are worth testing before you build a workflow on a piece of SEO crawler software.

The URL limit, and how it is counted. Per crawl, per month, per project and per account are four different limits. Per crawl caps bite immediately; monthly quotas bite when you re-crawl to confirm a fix.

JavaScript rendering. Either the crawler executes a page's scripts before reading it, or it reads only the HTML the server sent. On a modern framework, that difference is the whole crawl.

Export restrictions. Check the row limit and the formats. CSV and Excel cover most analysis; JSON and XML matter if results feed a script or a sitemap process.

Where your data goes. A hosted crawler keeps a copy of your site structure, and often your clients' too, which may need to appear in a data processing agreement. Desktop and self-hosted crawlers keep it on hardware you control.

Whether it is maintained. Abandoned crawlers are common, because one that worked in 2019 still appears to work while quietly mishandling newer response patterns. Check the last release and whether recent issues have replies.

URL Limits: Where Free Crawlers Usually Stop

People underestimate URL counts, often by an order of magnitude. A brochure site with ten pages in the navigation can hold 60 URLs once tag archives, pagination and a forgotten author page are counted. A catalog of 2,000 products can generate tens of thousands through faceted navigation, sort parameters and session identifiers, and a crawler finds all of them because search engines do too.

That is why a 500 URL ceiling matters more than it sounds. It audits a small site properly and samples a large one, yet it is too small to answer whether anything anywhere is broken, and parameter explosions surface only after a few thousand URLs. Unlimited crawling in a self-hosted tool means no software cap, with the real ceiling set by memory, disk and patience, and it moves when you add RAM. Past a few hundred thousand URLs, technique matters as much as the tool, which our guide to large scale website crawling covers.

JavaScript Rendering Changes What You See

Without rendering, a crawler reads the HTML the server sent. With rendering, it loads the page in a real browser engine, waits for scripts to finish, and reads the resulting DOM. On a server-rendered site the two are nearly identical. On a site built with React, Vue, Angular or Next.js in client-side mode, the raw HTML can be an empty shell with no title and no links, so an unrendered crawl returns a handful of URLs and a pile of issues that exist only because the crawler never saw the real page.

Rendering costs time and memory, often five to ten times the resources per URL, since every page now involves a browser instead of one HTTP request. Crawl unrendered first, compare the URL count and title fill rate against what you expect, then switch rendering on when the numbers disagree with reality. It is also the feature most likely to sit behind a paywall: Screaming Frog locks it in the free version, and several hosted tiers render only a sample. LibreCrawl renders with Playwright driving Chromium.

Exports, Data Ownership and Maintenance

A crawl becomes useful once you can sort, filter and share it, which means exports, and export limits are where free tiers are quietest. Some cap rows in the hundreds, some strip columns, some offer a PDF summary only. Run a small crawl and export it early, so you learn the restriction cheaply.

Data ownership deserves a thought on client work, because a full crawl is a map of a business: every page, every internal link, often staging URLs nobody meant to publish. Maintenance is the quieter risk. A commercial crawler stays current because customers pay for it; an open source crawler stays current because someone is working on it, so check that activity rather than assuming it. LibreCrawl development happens in public on GitHub, where the commit history shows how active the project is.

The Main Options, Compared

These are the tools worth knowing if your crawling budget is zero.

Tool Cost Free URL limit JavaScript rendering Best for
Screaming Frog (free version) Free, permanent 500 per crawl Locked Small sites, desktop convenience, zero setup
Screaming Frog (paid) $279 per year, per seat Unlimited, memory permitting Yes Teams wanting support, integrations and log file analysis
Google Search Console Free Reporting rather than on-demand crawling Yes, in URL Inspection Seeing how Google crawled and indexed your site
Bing Webmaster Tools site scan Free Capped scan of a verified site Partial A hosted on-page check with nothing to install
LibreCrawl Free, MIT licensed No software limit, hardware bound Yes, via Playwright Unlimited crawling on infrastructure you control

Where Screaming Frog's Free Version Wins

Screaming Frog has been refined since 2010 and it shows. The free version is a full desktop application: download it, point it at a URL, and you have crawl data in a minute with no server and no dependencies. For a site under 500 URLs it is the shortest route to an answer. If you want a desktop app with no setup and your site is small, use it rather than us.

The free version caps crawls at 500 URLs and locks crawl configuration, saving and reopening crawls, scheduling, JavaScript rendering, crawl comparison, structured data validation, custom extraction, Analytics and Search Console integration, and technical support. We checked their pricing page while writing this: the license is £199 per year, shown as $279 when the page is switched to US dollars, and bulk bands cut the per-seat price from five licenses upward, reaching roughly 15 percent off at twenty or more. Prices move, so confirm before you budget.

We list what the free version includes and what it locks in our breakdown of its limits, and survey the wider field in our roundup of Screaming Frog alternatives, which recommends other tools over ours where they win.

What Google Search Console and Bing Webmaster Tools Cover

Both are free, both need verified ownership, and both are worth having whichever crawler you choose, because they report what search engines did with your pages rather than what your pages contain.

Google Search Console reports on Google's own crawling. The Pages report groups URLs by indexing status with reasons for exclusion, Crawl Stats shows request volume and response codes over time, and URL Inspection fetches one URL live and shows the rendered HTML Google sees. For diagnosing indexing problems that combination is unmatched. What it leaves out is crawling on demand and handing you a table of every title and canonical, so it complements a crawler rather than replacing one. Bing Webmaster Tools goes further with Site Scan, a capped crawl on Bing's infrastructure that reports missing titles, broken links and redirect problems for verified properties.

When a Free SEO Crawler Is Enough, and When It Is Not

A free crawler handles most technical SEO work. Finding broken links, auditing titles, checking canonical and noindex logic, mapping internal linking, verifying a migration: that is ordinary crawling, and free tools do it as accurately as paid ones, because an HTTP status code has no premium version.

Paid tools earn their money elsewhere: scheduled monitoring with alerts when something breaks overnight, log file analysis of what bots actually requested, integrations that pull Search Console and PageSpeed data into the crawl table, branded client reports, and a support contract. If your work depends on any of those, pay for the tool that covers it. Team size matters too: one person auditing their own site has little reason to pay, while an agency of twenty weighs a recurring per-seat cost against the upkeep of a self-hosted tool.

Where LibreCrawl Fits, and Where It Falls Short

LibreCrawl is a free, MIT licensed, open source SEO crawler that you host yourself. It runs as a web application, so the interface is a browser tab pointed at your own server. There is no URL cap: the limit is the hardware you give it. JavaScript rendering uses Playwright with Chromium. It builds an internal link graph with per-URL "linked from" data, draws an interactive site structure visualization with Cytoscape.js, detects issues, exports to CSV, Excel, JSON and XML, supports multiple sessions so several projects crawl at once, and takes plugins.

The trade-offs are real. You host it, which means an installation to perform, a server to keep patched and dependencies to update. Support is GitHub issues and discussions rather than a support desk with response times. It has fewer accumulated edge cases handled than a tool refined since 2010, so unusual server configurations and malformed markup are likelier to surface something we have yet to fix. Client-ready PDF reports are absent: you export the data and build the report yourself. Search Console and structured data validation integrations are absent too, so schema checking stays a separate task.

So the fit is narrow and clear: unlimited crawling, crawl data that stays on your own infrastructure, or a crawler you can modify. For a small site with no appetite for self-hosting, a desktop free tier serves you better.

Run an Unlimited Crawl for Free

MIT licensed, no URL cap, no paid tier, JavaScript rendering included. Install it on your own hardware, or try the demo first.

Get it on GitHub Try Live Demo

Running Your First Crawl and What to Look At First

Whichever tool you picked, the first crawl follows the same shape. Settle the hostname: pick the version of the domain that serves content, with or without www, on https, so the crawl spends its budget on real pages instead of a redirect chain. Leave JavaScript rendering off for this run, set concurrency at two to five threads for shared hosting, and let it finish.

Read the total URL count first. If it is far lower than expected, something blocked the crawl: robots.txt, a noindex on a hub page, a login wall, or client-side rendering hiding the links. If it is far higher, you have found a parameter or faceted navigation problem that is spending real crawl budget at Google.

Then work through five views in order. Status codes, to catch 404s and 5xx responses on pages that still receive internal links. Redirects, looking for chains longer than one hop and any loops. Indexability, confirming every page you want indexed lacks a noindex tag and points its canonical at itself. Titles and meta descriptions, filtered for missing, duplicated and truncated ones. Then the link graph: pages at depth five or more, and pages with one internal link or none, where orphan content hides. Fix, re-crawl, compare, and keep that first export as a baseline.

Frequently Asked Questions

Is there a completely free SEO crawler?

Yes, in two senses. Screaming Frog's free version costs nothing and stays free, but it caps each crawl at 500 URLs and locks configuration, saved crawls and JavaScript rendering. Open source crawlers such as LibreCrawl are free in the fuller sense: no cap, no paid tier, no trial clock, with the source published under the MIT license. The trade is that you host and maintain them yourself.

What is the best free website crawler?

It depends on the size of the site and how much setup you will tolerate. Under 500 URLs, Screaming Frog's free version is the shortest path and the data it returns is excellent. Above that, or when you need JavaScript rendering without paying, a self-hosted open source SEO crawler such as LibreCrawl covers more ground. For properties you have verified, Google Search Console and the Bing Webmaster Tools site scan add useful signals at no cost.

Do free SEO crawlers limit how many URLs you can crawl?

Most of them do. Screaming Frog's free version stops at 500 URLs per crawl. Hosted free tiers usually apply a monthly page quota per project instead, and many restrict how many rows you can export. Self-hosted open source crawlers generally apply no software limit, so the ceiling becomes the RAM, disk and patience you have available.

Is there an open source SEO crawler?

LibreCrawl is one, MIT licensed with the source on GitHub. Open source matters for three practical reasons here: you can read what the crawler does with your data, you can add checks that no vendor is going to build for you, and the license permits commercial use, so client work raises no licensing question.

Can I crawl a website for free without installing anything?

For a site you own and can verify, yes. The Bing Webmaster Tools site scan crawls your pages on Bing's infrastructure and reports on-page issues, and Google Search Console shows you how Google crawled the site. For a site outside your control, such as a competitor or a prospect, hosted free tiers tend to require verification or payment, so a desktop or self-hosted crawler is the workable route.