The caddy-data volume is a podman named volume, so it sits under /var/lib/containers on the boot disk rather than the separately managed data disk. An instance replacement re-registers the ACME account and re-issues, and enough of those in a week hits Let's Encrypt's duplicate-certificate limit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
11 KiB
Runbook
Post-deploy verification
Work top to bottom the first time. Several of these controls fail silently, so the drills matter more than the status output.
Infrastructure
cd infra && pulumi preview # clean, no diff
dig +short A gitea.jasonmross.dev # the static IP
gcloud compute instances describe gitea-vm --zone <zone> \
--format='value(disks[].deviceName, shieldedInstanceConfig)'
Host
make ssh
mount | grep /var/lib/gitea # PD mounted, xfs
systemctl list-dependencies gitea.service | grep mount # the ordering dep exists
ls -Zd /var/lib/gitea # container_file_t
systemctl status gitea caddy nftables fail2ban
podman ps # both healthy
The nftables reload check — this is what catches an accidental global flush:
sudo nft list ruleset | grep -E '^table (inet gitea_filter|inet netavark|ip netavark)'
sudo systemctl reload nftables
sudo podman ps # container networking must still work
curl -fsS https://gitea.jasonmross.dev/api/healthz
Subnet agreement — a mismatch here is what silently breaks fail2ban:
sudo podman network inspect gitea --format '{{range .Subnets}}{{.Subnet}}{{end}}'
sudo grep REVERSE_PROXY_TRUSTED_PROXIES /etc/gitea/app.ini
TLS / DNS-01
sudo journalctl -u caddy | grep -i 'acme\|challenge'
Look for the dns-01 challenge. If Caddy fell back to http-01, the
googleclouddns plugin or its credentials are not working — check
/etc/gitea/caddy-network and see Caddy cannot get a certificate below.
curl -vI https://gitea.jasonmross.dev # valid Let's Encrypt cert
DNS-01 means renewal does not need inbound port 80 at all. That is testable:
temporarily remove the gitea-allow-web port 80 rule and force a renewal.
The ACME account and certificates live in the caddy-data podman volume, under
/var/lib/containers on the boot disk, not the data disk. Replacing the
instance therefore re-registers and re-issues on first start. That is fine
occasionally, but Let's Encrypt allows 5 duplicate certificates per week, so
several replacements in a few days can lock issuance out until the window
rolls over.
fail2ban — drill it, do not trust the status output
# 1. Fail a web login, then read the log line. It must show YOUR ip,
# not a 10.89.x address. If it shows Caddy, REVERSE_PROXY_TRUSTED_PROXIES is wrong.
sudo journalctl CONTAINER_NAME=gitea | grep -i 'failed authentication'
# 2. The jail is live.
sudo fail2ban-client status gitea
# 3. Ban a throwaway address you control, then verify from that host that
# both 443 and 2222 are genuinely unreachable.
sudo fail2ban-client set gitea banip <ip>
sudo nft list set inet f2b-prerouting f2b-gitea-v4
sudo fail2ban-client set gitea unbanip <ip>
A ban that appears in fail2ban-client status but still lets traffic through
means the prerouting action is not in effect — the default INPUT-hook actions
never see DNAT'd container traffic.
WAF — git must still work, and blocks must still happen
Two drills. The first is the one that catches a broken product; run it after any change to the Caddyfile matcher.
# 1. git still works through the proxy (the bypass is intact)
git clone https://gitea.jasonmross.dev/<you>/<repo>.git /tmp/wafdrill
cd /tmp/wafdrill && dd if=/dev/urandom of=blob.bin bs=1M count=20
git add -A && git commit -qm 'waf drill' && git push
A 403 on push means the @gittransport matcher no longer covers the git
routes. See waf.md.
# 2. the WAF is actually inspecting the web branch
curl -s -o /dev/null -w '%{http_code}\n' \
'https://gitea.jasonmross.dev/?file=../../../../etc/passwd'
Expect 403 when gitea:wafMode is On, and 200 in DetectionOnly — in
detection mode, confirm it was recorded instead:
make ssh
sudo journalctl CONTAINER_NAME=caddy --since '5 min ago' | grep 949110
If neither blocks nor records, the WAF module is not in the request path — check
order coraza_waf first survived the last Caddyfile edit.
fail2ban's WAF jail follows the mode
sudo fail2ban-client status # caddy-coraza listed only when wafMode=On
grep -A2 '^\[caddy-coraza\]' /etc/fail2ban/jail.d/gitea.local
enabled = false in DetectionOnly is correct, not a bug: the WAF is not
refusing anything, so there is no verdict to escalate into a ban.
Auto-update and rollback
sudo podman auto-update --dry-run # lists both units; UPDATED = false
If that errors on authentication, /etc/containers/ar-auth.json is stale or
missing — sudo systemctl start gitea-ar-auth.service and check the timer.
Rollback drill. Push a deliberately broken :prod (a bad CMD is enough),
run sudo systemctl start podman-auto-update.service, and confirm:
sudo journalctl -u podman-auto-update | grep -i rollback
sudo podman inspect gitea --format '{{.ImageName}}'
If it does not roll back, Notify=healthy is not taking effect on this
podman version. Fall back to HealthCmd plus HealthOnFailure=stop in
vm/quadlets/gitea.container and record it here.
The weekly rebuild — fire it once by hand
Do this immediately after the first pulumi up. It is the only check here
that cannot wait for its schedule, because a failure is completely silent: the
job fires at 04:00 on a Sunday, gets a 400, and layer 3 of the update story
quietly stops feeding layer 2. Nothing alerts.
gcloud scheduler jobs run gitea-weekly-rebuild --location <region>
sleep 15
gcloud builds list --region <region> --limit 5 # a build must have started
# (2nd-gen triggers are regional;
# the default is the global region)
gcloud scheduler jobs describe gitea-weekly-rebuild --location <region> \
--format='value(status)'
The job POSTs an empty body to the regional
.../locations/{region}/triggers/{id}:run endpoint. That is deliberate:
RunBuildTriggerRequest.source is a 1st-generation RepoSource that cannot name
a 2nd-gen repository, and omitting it tells Cloud Build to use the trigger's own
configured repository and branch — the REST equivalent of
gcloud builds triggers run TRIGGER --region=… with no --branch.
If it returns 400 or 404, check scheduleRebuild in
infra/pkg/build/build.go: a 404 usually means the URL used the global
.../projects/{p}/triggers/{id}:run path instead of the regional one.
Until it passes, run make build manually to pick up base-image security fixes.
Memory on a 2 GB instance
e2-small is 2 shared vCPU / 2 GB RAM. Measured on this exact image set:
| idle | peak during a 60 MB / 400-object push | |
|---|---|---|
| gitea | 105 MB | 375 MB |
| caddy + Coraza + full CRS | 50 MB | 51 MB |
The WAF is not the memory story — CRS costs about 50 MB and does not grow under
load. Git subprocesses are: index-pack and gc scale with what is being
pushed, and that number climbs with repo size.
bootstrap.sh provisions a 2 GB swap file with vm.swappiness = 10 as ballast,
because GCE images ship with none and an OOM kill mid-push is the failure this
prevents. Check it:
free -m
swapon --show
Swap should be near-idle. If free -m shows sustained swap use, that is the
signal to move to e2-medium (pulumi config set gitea:machineType e2-medium),
not to enlarge the swap file. Things that will push you over:
- repositories in the multi-GB range, or many concurrent clones
- switching from SQLite to a PostgreSQL container on the same host
- adding a Gitea Actions runner to this VM (don't — see migrate-to-gitea-scm.md)
Resilience
make backup # object lands in the backups bucket
gcloud compute instances reset gitea-vm --zone <zone>
# after it comes back: data intact, cert valid, nftables and fail2ban up
Common operations
Roll back to a previous image
Every build pushes :$SHORT_SHA alongside :prod. Find the digest in the build
log, then:
make ssh
sudo podman tag <region>-docker.pkg.dev/<project>/gitea/gitea:<sha> \
<region>-docker.pkg.dev/<project>/gitea/gitea:prod
sudo systemctl restart gitea
For a durable rollback, re-point :prod in Artifact Registry instead — a host-
local retag is undone by the next podman auto-update.
Restore from a dump
gcloud storage cp gs://<project>-gitea-backups/dumps/gitea-<stamp>.zip .
A gitea dump archive contains the repositories, the SQLite database, custom
files, and config. Restore is documented upstream at
https://docs.gitea.com/administration/backup-and-restore; the short version is
to stop gitea.service, unpack over /var/lib/gitea, fix ownership to
1000:1000, and start it again. Do this once as a drill before you trust it.
Caddy cannot get a certificate
The most likely cause is the ACME plugin failing to reach the GCE metadata
server for Application Default Credentials. vm/bootstrap.sh probes this at
first boot and caches the answer:
cat /etc/gitea/caddy-network # "bridge" or "host"
To re-probe, delete that file and run sudo systemctl start gitea-config-sync.
If the bridge cannot reach 169.254.169.254, the file will say host and Caddy
is switched to the host network, reaching Gitea over 127.0.0.1:3000 instead.
Both paths are supported; only the rendering differs.
Also verify the zone-scoped grant:
gcloud dns managed-zones get-iam-policy <zone-name>
Changing the podman firewall driver
/etc/containers/containers.conf.d/10-gitea.conf sets
firewall_driver = "nftables". Changing it on a live host leaves conflicting
rules behind — stop the containers and reboot rather than reloading.
A config change did not take effect
gitea-config-sync.service only restarts services when a rendered file actually
changed. To see what it did:
sudo journalctl -u gitea-config-sync -n 100 --no-pager
If the render was skipped, the message will say the Gitea secrets were
unavailable — check gcloud secrets versions list gitea-internal-token.
Deferred hardening
- Rootless podman under a dedicated user. Needs
loginctl enable-linger,--usertimers, and subuid mapping; on a single-tenant VM the isolation gain is small, which is why it was not done up front. - PostgreSQL instead of SQLite. Add a
postgres.containerquadlet plus its volume and password secret, then follow Gitea's documented dump/restore migration. Worth doing well before SQLite write contention shows up. - Wildcard certificate for
*.gitea.jasonmross.dev— trivial now that DNS-01 works. - Ops Agent for metrics dashboards and log-based alerts.