Media is usually the biggest thing on a web server’s disk, and the least interesting thing for that server to be doing. Serving it from object storage through a CDN takes the weight off the disk and the bandwidth off the server. The risk is in the move. Here is how I approach it, so that nothing is ever deleted before a good copy has been proved to exist.
Design the bucket before you copy anything
- Private by default. Block all public access on the bucket. Visitors never talk to it directly.
- The CDN is the only reader, and only of the public folder. Private uploads never go there.
- Sites never hold cloud keys. A small service on the host does the uploading, so a hacked site has no credentials to steal.
- Versioning and an independent backup, so a bad sync or a bad day can be undone.
aws s3api put-public-access-block --bucket example-media \
--public-access-block-configuration \
BlockPublicAcls=true,IgnorePublicAcls=true,BlockPublicPolicy=true,RestrictPublicBuckets=true
aws s3api put-bucket-versioning --bucket example-media \
--versioning-configuration Status=EnabledThen let only your CDN distribution read the public prefix:
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Principal": { "Service": "cloudfront.amazonaws.com" },
"Action": "s3:GetObject",
"Resource": "arn:aws:s3:::example-media/sites/example/public/*",
"Condition": {
"StringEquals": { "AWS:SourceArn": "arn:aws:cloudfront::111122223333:distribution/EXAMPLE" }
}
}]
}Copy, then prove it
Copy first. Delete nothing yet:
rclone copy /srv/example/uploads s3:example-media/sites/example/public \
--transfers 4 --checkers 8 --size-onlyBefore any local file goes, check every one against the bucket. I only consider files older than 48 hours, so nothing that is still being written is touched:
cd /srv/example/uploads
find . -type f -mmin +2880 -printf '%P\n' > older.txt
rclone check . s3:example-media/sites/example/public --one-way \
--files-from older.txt \
--match ok.txt --differ differ.txt --missing-on-dst missing.txt
wc -l ok.txt differ.txt missing.txtOnly files in ok.txt are candidates for removal. Anything in differ.txt or missing.txt stays exactly where it is until you know why.
Large files uploaded in parts do not get a simple checksum in S3, so rclone can only compare their sizes. That is fine for files your site can rebuild, but be stricter with originals you could never recreate.
Rewrite the links carefully
WordPress stores URLs inside serialised data, so a plain find and replace in SQL can corrupt it. WP-CLI understands the format. Take an export first, then dry-run:
wp db export - | gzip > before-media-$(date +%F).sql.gz
wp search-replace 'https://example.com/wp-content/uploads/' \
'https://media.example.com/' --all-tables-with-prefix --precise --dry-run
# Happy with the counts? Run it again without --dry-run, then flush the cache.
wp cache flushKeep old links working for anyone who saved them, and for search engines that indexed them:
location ~ "^/wp-content/uploads/(\d{4}/.+)$" {
return 301 https://media.example.com/$1;
}Only then, free the space
Remove the verified files in small batches, keep the lists of what went, and keep the bucket’s versioning and backups switched on. Check the site, the CDN and the error log between batches.
Done this way, a media move is boring. That is exactly what you want from it.