timeout
The page did not answer within the crawl timeout. Found by SEO crawler →
5xx
The server returns an error instead of the page. Found by Technical SEO audit →

Requirements

SpiderHead runs on your own machine. The core needs Python and a few packages; everything else is optional.

  • Python 3.10+
  • macOS · Windows · Linux
  • No hosted account

Components by operating system

Components by operating system
ComponentmacOSWindowsLinuxNeedsStatusNotes
Core packageSupportedSupportedSupportedPython 3.10+, pip, gitRequiredInstalled into the project's virtual environment
CLISupportedSupportedSupportedCore packageRequiredUbuntu Server over SSH documented for headless use
MCP serverSupportedSupportedSupportedmcp extra, MCP SDK 1.29+OptionalLocal stdio; any MCP client
Desktop app.pkgsetuptarballInstaller for each OSOptionalSame core; scans shared with CLI and MCP
JS renderingSupportedSupportedSupportedChromium via Playwright, ~150 MBOptionalOnly for rendering checks
Screaming Frog CLISupportedSupportedSupportedActive SF licenceOptionalOnly to drive live SF crawls; exports need nothing
Docker imageSupportedSupportedSupportedDockerOptionalSlim 440 MB, full 1.08 GB

Disk space

Disk space by component
WhatSizeWhen you need it
Core with all extras≈ 860 MBVirtual environment, measured on macOS with .[all]
Chromium≈ 150 MBOnly for JavaScript rendering
Docker image440 MB · 1.08 GBSlim · full
Scan storagegrows with crawlsKeeps at least 1 GiB free; warns past 20 GiB of history

Plan for about 2 GB

Core with all extras, Chromium and room for the first scans. Large sites need more: scan size depends on page count and whether bodies are kept.

Memory and budgets

Large crawls run under budgets you set: wall time, peak memory, database size. A budget hit stops the run as blocked, never as a pass.

Data sources

Search Console, GA4, Metrika, Webmaster, CrUX need your own credentials and are read-only. Paid providers are off by default.

No provider credentials are needed for the core crawl and audit.