Cache Warming: GridPane Sitemap-Driven Crawler

13 min read

Introduction

Every time someone visits an uncached WordPress page, your server has to build it from scratch, running PHP and querying the database before anything loads. Page caching skips that work by saving a ready-made copy of each page and serving it instantly, which means a faster site for your visitors and less strain on your server.

Pre-warming takes this a step further by crawling your pages ahead of time, so the cache is already populated before real visitors arrive. Without it, the first visitor to each page triggers a full render, which is slow and can hurt both user experience and Core Web Vitals. With it, every visitor gets a fast, cached response from the start, and search engine crawlers see consistently fast load times too.

Cache warming can also be useful when migrating higher-traffic websites from one server to another, or in from a different host. By pre-warming the cache on the new server before switching your DNS, the site is fast and ready the moment traffic cuts over, rather than the new server taking the full brunt of cold requests on day one.

This article will walk you through how to get started with our sitemap crawler CLI tool and configure automatic cache warming after automatic server reboots.

How it Works

The sitemap crawler reads one or more XML sitemaps, builds a deduplicated list of URLs, and crawls each one so that the cache is already populated when visitors land. It works with any page cache, including our server-level FastCGI and Redis page caching.

The tool can be used to crawl one or multiple sitemaps at a configurable, paced rate, and has two primary functions:

  1. Generate: Parses the sitemap(s), including sitemap-index files and nested indexes, and writes a deduplicated list to urls.txt. You can inspect, filter, or reuse this list before crawling.
  2. Crawl: Requests every URL to warm the cache. Progress is recorded in done.txt, so an interrupted run resumes where it left off when you re-run the same command.

The urls.txt file is editable, so you can have the tool crawl and compile your sitemap, and then edit anything within that file per your needs.

Getting Started

Before You Begin

This is a CLI-based tool, and so everything detailed below is run directly on the command line. If this is your first time connecting to a server over SSH, please see the following articles to get started:

Sitemaps

The crawler requires your website sitemap/s to gather a list of pages to warm. Typically, this is either /sitemap_index.xml or /sitemap.xml.

Compile a list of the sitemaps you want to crawl on the specific server you want to run this on – note that this is a server-specific tool, not an account-wide tool, so be sure to only choose sites hosted on that specific server.

Side Note

This is also a good time and opportunity to check over your sitemaps and make sure they look correct (you may be surprised to find templates, contact forms, and other items you weren't expecting to see).

Full Command List and Format

Below is the full command list for the crawler.  In the following section, you’ll find example commands that show you how the pieces fit together.

sitemap-warm-crawl [options] <sitemap-url> [<sitemap-url> ...]
--ip <addr>          Force every host to resolve to this IP (curl --resolve, :80+:443).
                     Omit to use live DNS. THIS is the hosts-redirect target.
--pause <sec>        Pause time after each request, per worker (default 1; decimals ok, e.g. 0.5).
--concurrency <N>    Parallel workers (default 1 = strict serial). >1 for many-core boxes.
--include <ERE>      Keep only URLs matching this extended regex.
--exclude <ERE>      Drop URLs matching this extended regex (applied after --include).
--out <dir>          Work dir (default ./warm-<host>[-<sub>]).
--list-only          Generate the URL list and stop (no crawl).
--crawl-only <file>  Skip sitemap parsing; crawl an existing list file.
--max-time <sec>     Per-request timeout (default 180 — a cold render must finish + store).
-h, --help           Full help.

Pause and Concurrency

The two options that you’re most likely going to want to edit are:

  1. --pause This sets the time the crawler pauses between pages
  2. --concurrency This sets how many pages can be crawled simultaneously.

Here are our guidelines, based on how many v/CPUs your server has:

Server Size Start with: If load stays comfortable, re-run with:
2 CPU --pause 5 --concurrency 1 --pause 2 --concurrency 1
4 CPU --pause 1 --concurrency 1 --pause 0.5 --concurrency 2
6+ CPU --pause 1 --concurrency 1 --pause 0.5 --concurrency 2–4

Warming the Cache

Optional: Create Your URLs to Crawl Lists

The tool allows you to work from your existing sitemap or create your own list of pages to crawl. Depending on the website’s size and whether it publishes new content regularly, it may be easier to let the crawler work through the sitemap. For smaller brochure sites, you can edit and tweak the crawl list if you wish to do so.

Generate Your Crawl Lists

Swap in your own domain and sitemap URL:

screen -S warm
sitemap-warm-crawl --list-only https://site.url/sitemap_index.xml

For example:

screen -S warm
sitemap-warm-crawl --list-only https://example.com/sitemap_index.xml

Viewing and Editing Your Crawl List Files

Each website will be given its own folder and files. You can edit your URL list for a given site from the /root directory (the directory you start in when you connect to your server) with the following command:

nano warm-site.url/urls.txt

For example:

nano warm-mywebsite.com/urls.txt

This will open the file in the nano editor. Each URL has its own line. Once you’re done with your edits, you can save the file with CTRL+O followed by Enter, and exit nano with CTRL+X.

Manual Cache Warming

Below are a list of example commands that you can copy and modify. We recommend that you run commands in a screen per the details below.

Run Commands in a Screen

The following command will open up a screen for your cache warming. Run this before running the crawl commands:

screen -S warm

Here’s a quick rundown on how and why to use a screen, and how to watch the tool’s progress:

  • screen -S warm keeps the job running if your connection drops.
  • Detach with Ctrl-A then D, and reattach with screen -r warm.
  • Stop it any time with Ctrl-C. Re-running the same command picks up where it left off.
  • Watch progress from another shell: tail -f ./warm-yourdomain.com/crawl.log

Warm the Cache for One Website

The following command can be used to warm the cache for a specific website – swap out the highlighted domain and sitemap URL for your own website:

sitemap-warm-crawl --ip 127.0.0.1 --concurrency 1 --pause 5 \
https://yourdomain.com/sitemap_index.xml

This will crawl 1 page at a time and pause for 5 seconds between pages. 

Warm the Cache for Multiple Websites

The following example can be used to run multiple websites, one after the other, 1 page at a time, with 5 seconds between each page:

sitemap-warm-crawl --ip 127.0.0.1 --concurrency 1 --pause 5 \
https://exampleone.com/sitemap_index.xml
sitemap-warm-crawl --ip 127.0.0.1 --concurrency 1 --pause 5 \
https://second-example.com/sitemap_index.xml
sitemap-warm-crawl --ip 127.0.0.1 --concurrency 1 --pause 5 \
https://thirdwebsite.org/sitemap_index.xml

Warm the cache but skip long-tail pages

The following example warms the money pages on a busy website on a larger server, skips longer-tail, lower-traffic pages, and runs 4 pages at a time, with 0.5 seconds between each set of pages:

sitemap-warm-crawl --ip 127.0.0.1 --concurrency 4 --pause 0.5 \
--exclude '/[^/]+/blog/' \
https://example.com/sitemap_index.xml

Warm the cache with extended logs

By default, the first stage of the run (setting up the list of pages) shows its messages and warnings on screen only, and they aren’t saved anywhere. If you’d like one file that captures the whole story, add 2>&1 | tee yourfile.log to the end of your command. This shows everything on screen while also saving a copy.

This example captures the whole run —  the generation phase, warnings, and the crawl:

sitemap-warm-crawl --ip 127.0.0.1 --pause 1 \
https://example.com/sitemap_index.xml \
2>&1 | tee ~/warm-example.com.run.log

You can view your custom log with:

cat warm-example.com.run.log

And you can remove it once you no longer need it with:

rm warm-example.com.run.log

Boot-Time Auto-Warm (Automatically Re-Warm Your Caches After an Overnight Restart)

Your server will occasionally restart itself overnight to install security updates. When that happens, your website caches are wiped, so the first visitors in the morning could find the site slower than usual.To prevent this, you can use our warming tool to re-warm it automatically after a restart takes place, before morning traffic arrives.

The following command will configure cache warming only after a restart in the late-night window you choose, so a restart at any other time won’t trigger them.

Activate Automatic Cache Warming

You can set everything up with one command,with one --site entry per website:

gp-cache-warm-setup --from 2 --to 5 \
--site "https://siteone.com/sitemap_index.xml 1 5" \
--site "https://site-two.com/sitemap_index.xml 1 5" \
--site "https://examplethree.org/sitemap_index.xml 1 5"

In this example:

  • --from 2 --to 5 means the refresh only runs if the restart happens between 2 AM and 5 AM. This covers the overnight automatic security update reboot. A daytime reboot won’t trigger it.
  • Each --site value is "<sitemap-url> <workers> <pause-in-seconds>".
  • --site tells it which website to refresh.
  • The two numbers after the address set the pace: In the above example, 1 means it visits 1 page at a time, and 5 means it pauses for 5 seconds between visits.

How to Verify Your Pages are Cached

A crawled page should be served from cache on a follow-up request. Here’s a command you can modify to check the cache header on one of your websites:

curl -sk --resolve example.com:443:127.0.0.1 -o /dev/null -D - \
"https://example.com/example-page/" | grep -i x-grid

For example:

curl -sk --resolve example.com:443:127.0.0.1 -o /dev/null -D - \ 
"https://gridpane.com/kb/" | grep -i x-grid

This will give you the cache TTL and whether it’s a cache HIT, MISS, or STALE. 

Header Value: X-Grid-Cache: HIT Or X-Grid-SRCache-Fetch: HIT

The page has been served from the cache (warmed, within TTL).

Header Value: X-Grid-Cache: STALE / UPDATING

This is specific to FastCGI caching. The cache is still being served but it’s past its TTL. It will be refreshed in the background. 

Header Value: X-Grid-Cache: MISS Or X-Grid-SRCache-Fetch: MISS

The page is NOT cached – see Troubleshooting.

Logging and troubleshooting

Files and Logging

The tool creates 3 files per website. These are straightforward and automatically generated. Files are written to the work directory (default ./warm-<slug>/, for example: ./warm-example.com.
File Contents
urls.txt The generated URL list
crawl.log One line per request: HTTP code, cache header, time, size, URL
done.txt URLs warmed successfully (HTTP 200); the resume ledger

When you first connect to your server, run the following to display your directories in a list:

ls -l

In this list you will see any directories that have been created by running the crawling tool. You can view the above files with the cat command, for example:

cat warm-example.com/crawl.log

Logging

The crawl log’s cache= column records the header seen at warm time (MISS on the cold first hit is expected — that hit is what populates the cache).

Each time the tool runs, it saves its results in the folder you chose with --out. Two files are created there:

  • crawl.log: a diary of what happened. It records when the run started and finished (times are in UTC, with totals at the end), plus one line for every web page it visited. Each page line shows whether the page loaded successfully, how long it took, and how large it was. You can watch this on screen as the run progresses, and it’s saved to a file at the same time.
  • done.txt: a checklist of every page that has already been processed. If a run is interrupted, the tool uses this list to pick up where it left off instead of starting over.

Troubleshooting

Pages keep showing cache=MISS every time you check

The cache is being skipped for these pages, so they’re rebuilt from scratch on every visit. Common causes:
  • The page address has extra bits on the end (like ?something=123)
  • A login cookie is attached to the visit
  • A “no-cache” rule is set for the page
  • Caching isn’t turned on for that part of the site
What to do: Look at your site’s caching settings (sometimes called “skip rules” or “exclusions”) and make sure the page isn’t being excluded. Note that the cache-warming tool visits the exact address listed in your sitemap and doesn’t log in, so it only warms pages that logged-out visitors can see.

Very large pages (several MB) time out or never get cached

When a page is too big for the server’s working memory, it gets temporarily saved to disk instead, and that can cause the caching to fail. You’ll usually see a long pause, and then the page still isn’t cached.

What to do:
Increase the server’s memory allowance for page responses by raising the fastcgi_buffer_size and fastcgi_buffers settings, so the whole page fits in memory. To make the change permanent, put it in the file conf.d/_custom/fastcgi.conf.

The cache status shows cache=- (blank)

The server didn’t report any caching information for this page. Either that page is handled in a way that doesn’t use caching, or page caching is switched off entirely.

What to do:
Check which type of caching is currently active on your site and confirm it’s turned on.

The log keeps showing HTTP=000

The tool couldn’t connect to the server, or it gave up waiting for a reply.

What to do:

  1. Check that the IP address you gave with --ip is correct.
  2. Make sure the web server (nginx) is running.
  3. If your pages are just slow to load, increase the wait time using --max-time.

The log shows HTTP=301 or HTTP=302

The address in your sitemap forwards visitors to a different address (a redirect). The cache-warming tool only visits the exact address listed and won’t follow the redirect, so the page doesn’t get warmed.

What to do:
Usually nothing. If the address it redirects to is also in your sitemap, that one will be warmed automatically. If it isn’t, consider updating your sitemap to list the final address instead.