LibreCrawl ships with a Dockerfile and a docker-compose.yml, so putting it on your own server is a clone, a copied environment file, and one command. The decisions worth thinking about are the ones around that: which environment variables you set, where the crawl database lives, and how you reach the interface without leaving an open crawler on the public internet.

Why Run the Crawler Yourself

Four reasons come up repeatedly, and they are worth separating because only some of them will apply to you.

The data stays on your infrastructure. A crawl of a client site produces a detailed inventory of that site: every URL, every status code, every title and meta description, the full internal link graph, and often staging paths and parameter structures the client would rather keep quiet. With a hosted crawler, all of that lands in someone else's database under someone else's retention policy. Self-hosted, it lands in a SQLite file on a disk you control. For agencies under a confidentiality clause, that distinction is the whole argument, and it is a much shorter conversation with a client's security team than explaining a third-party processor.

GDPR gets simpler, though it does not disappear. Crawl output is mostly URLs and page metadata, but sites leak personal data into both: author pages, user profile URLs, email addresses in mailto links, staff directories. Running the crawler yourself keeps that inside your existing processing arrangements instead of adding another sub-processor to disclose, another data processing agreement to negotiate, and another cross-border transfer to justify. You still owe the client a retention answer, but you can actually give one, because the retention policy is your own cron job.

No per-seat licensing. LibreCrawl is MIT licensed, so the tenth person on your team costs the same as the first: nothing. Compare that with Screaming Frog at $279 per year per user, which is fine for a solo consultant and turns into a real line item for a team of eight who each need it occasionally. One self-hosted instance with accounts covers the whole team, and the cost is the server. We have a fuller feature-by-feature comparison in the Screaming Frog alternative write-up.

A server crawls better than a laptop. This is the reason people underrate. A crawl running on a VPS has a stable IP you can ask a client to allowlist, symmetric bandwidth, and no lid to close. Laptop crawls die from sleep, VPN drops, and hotel wifi, and every one of those is an interrupted job. Long crawls in particular want a machine that stays up, which is most of the argument in our guide to large-scale crawling.

When Self-Hosting Is the Wrong Choice

If you crawl one site of a few thousand pages every month or two, install a desktop tool and skip everything below. Self-hosting means you now operate a server: OS patches, a TLS certificate that expires, a disk that fills up, a container that needs rebuilding after an upgrade, and an outage that is yours to fix at the exact moment a client is waiting for an audit. That is a reasonable trade when the crawler is core to how your team works. It is a bad trade to save a licence fee you could pay in ten minutes of billable time.

Two other cases where it fits badly. If nobody on the team is comfortable on a Linux shell, the first time the container fails to start you will be stuck, and a hosted tool with support is the honest answer. And if you need scheduled crawls with alerting and historical trend dashboards out of the box, that is a monitoring product rather than a crawler, and you would be building the missing half yourself.

What Is Actually in the Container

Worth knowing before you size anything. The image is built from python:3.11-slim, installs the Python requirements, then installs Playwright and its browsers so JavaScript rendering works inside the container. That last part is why the image is large: you are shipping browser binaries and the graphics, font and audio libraries Chrome links against. The Dockerfile installs those explicitly, then runs playwright install-deps as root and playwright install as the application user.

The application runs as a non-root user called librecrawl with UID 1000, which matters when you look at file ownership in the mounted data directory. The app listens on port 5000 inside the container, served by waitress rather than the Flask development server.

The Setup

Three commands on a machine with Docker and Docker Compose installed:

# Clone the repository
git clone https://github.com/PhialsBasement/LibreCrawl.git
cd LibreCrawl

# Copy environment file
cp .env.example .env

# Start LibreCrawl
docker compose up -d

Edit .env before that last command rather than after, because the values in it decide both how the app authenticates and which network interface the port is published on. The defaults in .env.example are deliberately the safe ones: authentication on, bound to localhost.

If you would rather have the repository decide, there are startup scripts that detect Docker and fall back to a Python install if it is missing:

chmod +x start-librecrawl.sh
./start-librecrawl.sh

Those are convenient on a workstation. On a server, run compose directly so you can see what it does.

The Environment Variables That Matter

Everything below comes from .env.example and is consumed by docker-compose.yml. The compose file turns several of them into command line flags for main.py, so setting the variable and passing the flag are the same thing.

  • HOST_BINDING - The interface Docker publishes the port on. 127.0.0.1 means localhost only, 0.0.0.0 means every interface. Defaults to 127.0.0.1 in the compose file if unset. This is the single most important line in the file.
  • EXTERNAL_PORT - The host port mapped to the container's 5000. Defaults to 5000. Change it if something else on the box already owns that port.
  • LOCAL_MODE - true disables authentication entirely and grants every visitor admin access with no rate limits. Defaults to false. Only set it true when the port is reachable from your own machine alone.
  • SECRET_KEY - Signs session cookies. Left unset, the app generates a random key at startup, which is secure but logs out every user on every restart. Set it for anything long-lived.
  • REGISTRATION_DISABLED - true stops new accounts being created. On a team instance, register your accounts, then set this to true so a discovered URL cannot become an open sign-up form.
  • DEMO_MODE - true applies a 1.5GB per-user memory limit. Intended for public demos, and a cap you do not want on a real audit.
  • DANGEROUSLY_SKIP_AUTH - true lets anyone log in as any username with no password, with the username used only to keep sessions separate. The name is accurate. Leave it false.
  • SMTP settings - SMTP_HOST, SMTP_PORT, SMTP_USER, SMTP_PASSWORD, SMTP_FROM and SMTP_FROM_NAME drive verification emails, and MAIN_APP_URL sets the base URL those emails link to. They are commented out in the compose file, so uncomment them there as well as filling them in .env. Skip all of this if registration is disabled.

Generate a real secret key with the one-liner from the repository:

python -c "import secrets; print(secrets.token_hex(32))"

Paste the output into SECRET_KEY in .env and keep that file out of version control.

Local Mode Versus Authentication

These are two different deployments that happen to share a compose file.

Local mode is for one person on one machine. Every visitor is auto-logged-in as admin, tier limits and rate limits are off, and there is no login screen to get past. The .env for it:

# .env file
LOCAL_MODE=true
HOST_BINDING=127.0.0.1
REGISTRATION_DISABLED=false

Local mode only makes sense paired with HOST_BINDING=127.0.0.1. The combination of LOCAL_MODE=true and HOST_BINDING=0.0.0.0 publishes an admin console with no password to every interface on the box, and on a cloud VM that means the internet.

Authenticated mode is for anything shared. Accounts, login, and the tier system are all active, guest users are capped at three crawls per 24 hours by IP, and every session is isolated with its own crawler instance and data:

# .env file
LOCAL_MODE=false
HOST_BINDING=0.0.0.0
REGISTRATION_DISABLED=false
# Generate with: python -c "import secrets; print(secrets.token_hex(32))"
SECRET_KEY=replace-with-a-long-random-string

Even here, consider keeping HOST_BINDING=127.0.0.1 and letting a reverse proxy on the same host do the listening. More on that shortly.

Running directly under Python instead of Docker, the equivalent flags are --local or -l for local mode, --disable-register or -dr, --disable-guest or -dg, --demo or -dm, and --dangerously-skip-auth or -dsa:

# Local mode (all users get admin tier, no rate limits)
python main.py --local

Where the Data Lives

The compose file mounts one volume:

volumes:
  # Persist the user database and settings
  - ./data:/app/data

Inside that directory sits users.db, a SQLite database holding both the user accounts and the persisted crawl data. Because it is a bind mount to ./data in the cloned repository, it survives docker compose down, image rebuilds and host reboots. Remove that volume line and every account and every crawl disappears with the container.

Three practical consequences:

  • Back up the data directory, and nothing else. It is the entire state of the instance. The rest of the tree is a git checkout you can recreate.
  • Watch its size. Crawl data accumulates there across jobs, so a busy team instance grows steadily. Put it on a disk with room, and delete crawls you have already exported.
  • Mind the file ownership. The container runs as UID 1000, so the files it writes into ./data are owned by UID 1000 on the host. If your host user has a different UID, you will need to adjust permissions before poking at the database directly.

One thing that does not live there: crawler settings. Those are stored in browser localStorage per browser, along with any custom CSS theme, so clearing site data resets them and each team member configures their own. Sessions expire after an hour of inactivity. If you want configuration and exports driven from somewhere other than a browser, the API docs cover the endpoints, and a script against the API is the right place to put anything you need reproducible.

Memory, Shared Memory and Playwright

The compose file sets shm_size: '2gb'. Chrome uses shared memory heavily, and Docker's 64MB default causes browser crashes under any real rendering load. Keep that line. If you build your own compose file from scratch, add it.

Rendering also changes the shape of resource use. HTML-only crawling is mostly network bound; rendering launches browser contexts and makes the job CPU and memory hungry, and those browser processes usually outweigh the crawl data itself. Rather than guessing at RAM, run a crawl and read the live per-URL memory figures in the UI, which is the sizing method described in the large-scale crawling guide. Measuring on your own sites beats any number we could print here.

The container also has restart: unless-stopped, so it comes back after a host reboot or a crash without you logging in.

Reaching It Safely

This is the part that matters most, and the part most often skipped.

An SEO crawler with a web UI should never be published straight to the internet. A crawler that anyone can reach is a crawler anyone can point at any site, from your IP address, at whatever concurrency they like. The consequences land on you: your server gets blocked by the target, your IP shows up in someone's abuse report, your hosting provider gets the complaint. An open crawler is an open proxy for generating traffic, and it will be found, because scanners sweep port 5000 continuously.

The setup to use instead:

  • Keep HOST_BINDING=127.0.0.1. The container port is then only reachable from the host itself, which is the default in .env.example for good reason.
  • Put a reverse proxy in front. Nginx, Caddy or Traefik on the same host, terminating TLS and proxying to 127.0.0.1:5000. The proxy is the only thing listening publicly.
  • Terminate TLS properly. Login credentials and crawl output both cross that connection. Use a real certificate, and redirect plain HTTP to HTTPS.
  • Add a second layer of access control at the proxy. HTTP basic auth, an allowlist of office and VPN addresses, or an SSO forward-auth setup. The application's own login is one gate; the proxy is the gate that stops unauthenticated traffic ever reaching the app.
  • Set REGISTRATION_DISABLED=true once your accounts exist. An open registration form on an internet-reachable crawler is the same problem wearing a login page.
  • Consider skipping public exposure entirely. A WireGuard or Tailscale network, or an SSH tunnel to localhost, gives the team access with nothing published at all. For a small team this is less work than certificates and stays safe by construction.

The two settings that should never appear together on a public interface are LOCAL_MODE=true and DANGEROUSLY_SKIP_AUTH=true. Both disable authentication by design. Both are fine behind localhost and dangerous anywhere else.

The Maintenance You Are Signing Up For

Upgrades are a pull and a rebuild, since the compose file builds from the local Dockerfile rather than a published image:

git pull
docker compose up -d --build

The ./data bind mount is outside the image, so accounts and crawls carry across. Check the logs afterwards to confirm the container came up rather than restart-looping:

docker compose logs -f librecrawl

Beyond that, the standing jobs are the ordinary ones: host OS patches, certificate renewal if you are terminating TLS yourself, disk space monitoring on the volume holding data, and an actual backup of that directory tested by restoring it once. A rebuild after git pull refetches Playwright browsers, so expect it to take a while and to consume disk on old image layers. None of this is difficult. It is simply real work that recurs, and it is the honest cost side of the ledger against the licence fees you are avoiding.

Key takeaways:

  • Self-host when data residency, team size, or long unattended crawls justify running a server, and use a desktop tool when they do not
  • Clone, cp .env.example .env, edit the file, then docker compose up -d
  • HOST_BINDING and LOCAL_MODE together decide whether you have a private tool or a public admin console
  • Set a real SECRET_KEY so sessions survive restarts, and set REGISTRATION_DISABLED=true once accounts exist
  • The ./data bind mount holds users.db with every account and crawl, so back it up and watch its size
  • Bind to localhost and reach the UI through a reverse proxy with TLS, or over a private network, rather than publishing port 5000

Run LibreCrawl on Your Own Server

MIT licensed, no per-seat cost, no crawl limits, and a Dockerfile and compose file in the repository. Your crawl data stays on your infrastructure.

Download LibreCrawl