Gitea on GCE: podman quadlets, Pulumi, Cloud Build

Self-hosted Gitea on a single e2-small AlmaLinux 10 VM in us-east1,
serving gitea.jasonmross.dev.

Runtime is podman quadlets (systemd .container/.network/.volume units).
Both images are built on Debian 13: Gitea from a GPG-verified release
binary, and Caddy from an xcaddy build carrying the Google Cloud DNS
provider (ACME DNS-01) and the Coraza WAF with the OWASP CRS embedded.

Infrastructure is a Pulumi program in Go against a GCS state backend.
Cloud Build handles CI: a push trigger for images, one for infra, and a
weekly scheduled rebuild. Everything Cloud Build touches is 2nd gen.

Notable design decisions, each documented where it lives:

- Quadlets track a floating :prod tag. AutoUpdate=registry compares
  digests for a tag, so a digest-pinned image silently disables
  auto-updates.
- Git transport and LFS bypass the WAF. With the bypass removed, a plain
  git push returns 403 -- packfiles trip CRS reliably.
- gitea:wafMode drives both SecRuleEngine and whether the fail2ban jail
  acting on WAF verdicts exists. Banning on detections that were never
  blocks would turn a tuning false positive into an nftables ban.
- fail2ban bans at the nftables prerouting hook. Published container
  ports are DNAT'd and never traverse INPUT, where the stock actions
  install their rules.
- The DNS zone, backup bucket, and Gitea signing secrets are not
  Pulumi-owned, so pulumi destroy cannot take them with it.
- The podman subnet is pinned because it is what Gitea's
  REVERSE_PROXY_TRUSTED_PROXIES names.

Three update layers: dnf5-automatic for the OS, podman-auto-update with
health-gated rollback for containers, and a weekly image rebuild that
gives the second layer something to pull.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-18 21:45:33 -05:00
co-authored by Claude Opus 5
commit c0382d5d31
50 changed files with 4449 additions and 0 deletions
+616
View File
@@ -0,0 +1,616 @@
#!/usr/bin/env bash
#
# Gitea VM bootstrap. Set as the GCE `startup-script` metadata value by Pulumi,
# and re-run by gitea-config-sync.service with --sync-only after a config push.
#
# MUST be idempotent: GCE runs the startup script on every boot.
#
# (no args) full run -- packages, disk, SELinux, firewall, units, config
# --sync-only re-pull vm/ from GCS, re-render templates, restart what changed
#
set -euo pipefail
readonly STATE_DIR=/opt/gitea-config
readonly RENDER_DIR=/etc/containers/systemd
readonly LOG_TAG=gitea-bootstrap
log() { echo "[${LOG_TAG}] $*" >&2; }
die() { echo "[${LOG_TAG}] FATAL: $*" >&2; exit 1; }
warn() { echo "[${LOG_TAG}] WARN: $*" >&2; }
MODE=full
[[ "${1:-}" == "--sync-only" ]] && MODE=sync
# ---------------------------------------------------------------------------
# Instance metadata (populated by Pulumi)
# ---------------------------------------------------------------------------
meta() {
curl -fsS -H 'Metadata-Flavor: Google' \
"http://169.254.169.254/computeMetadata/v1/instance/attributes/$1" 2>/dev/null || true
}
CONFIG_BUCKET=$(meta config-bucket)
BACKUP_BUCKET=$(meta backup-bucket)
GCP_PROJECT=$(meta gcp-project)
AR_HOST=$(meta ar-host)
IMAGE_GITEA=$(meta image-gitea)
IMAGE_CADDY=$(meta image-caddy)
DOMAIN=$(meta domain)
ACME_EMAIL=$(meta acme-email)
APP_NAME=$(meta app-name)
PODMAN_SUBNET=$(meta podman-subnet)
PODMAN_GATEWAY=$(meta podman-gateway)
REQUIRE_SIGNIN_VIEW=$(meta require-signin-view)
WAF_MODE=$(meta waf-mode)
DATA_DISK_DEVICE=$(meta data-disk-device)
[[ -n "${CONFIG_BUCKET}" ]] || die "config-bucket metadata is missing; nothing to sync from"
[[ -n "${DOMAIN}" ]] || die "domain metadata is missing"
: "${APP_NAME:=Gitea}"
: "${REQUIRE_SIGNIN_VIEW:=false}"
# DetectionOnly is the safe default: run it, read what it flags, add
# exclusions, then switch to On. See docs/waf.md.
: "${WAF_MODE:=DetectionOnly}"
case "${WAF_MODE}" in
On|DetectionOnly|Off) ;;
*) die "waf-mode must be On, DetectionOnly, or Off (got: ${WAF_MODE})" ;;
esac
: "${DATA_DISK_DEVICE:=/dev/disk/by-id/google-gitea-data}"
export GCP_PROJECT AR_HOST IMAGE_GITEA IMAGE_CADDY DOMAIN ACME_EMAIL APP_NAME
export PODMAN_SUBNET PODMAN_GATEWAY REQUIRE_SIGNIN_VIEW WAF_MODE
# ---------------------------------------------------------------------------
# Packages
# ---------------------------------------------------------------------------
install_packages() {
log "installing packages"
dnf -y install \
podman container-selinux \
nftables \
jq gettext \
policycoreutils-python-utils \
xfsprogs
# fail2ban lives in EPEL on RHEL-family distros. The exact package set has
# shifted between EPEL releases, so probe rather than assume -- this is the
# one dependency most likely to be named differently on EPEL 10.
if ! rpm -q epel-release >/dev/null 2>&1; then
dnf -y install epel-release || warn "epel-release unavailable; fail2ban will be skipped"
fi
if dnf -y install fail2ban fail2ban-server 2>/dev/null; then
# The systemd journal backend needs the Python bindings; without them
# fail2ban silently falls back and matches nothing.
dnf -y install python3-systemd || warn "python3-systemd missing; the systemd backend may not work"
else
warn "fail2ban not installable from configured repos -- skipping fail2ban setup"
fi
}
# ---------------------------------------------------------------------------
# Data disk
# ---------------------------------------------------------------------------
setup_data_disk() {
log "configuring data disk ${DATA_DISK_DEVICE}"
[[ -e "${DATA_DISK_DEVICE}" ]] || die "data disk ${DATA_DISK_DEVICE} not present"
if ! blkid "${DATA_DISK_DEVICE}" >/dev/null 2>&1; then
log "disk is unformatted -- creating XFS filesystem"
mkfs.xfs -q "${DATA_DISK_DEVICE}"
fi
local uuid
uuid=$(blkid -s UUID -o value "${DATA_DISK_DEVICE}")
[[ -n "${uuid}" ]] || die "could not read UUID from ${DATA_DISK_DEVICE}"
mkdir -p /var/lib/gitea
# By UUID, never by device path: GCE can reorder /dev/sdX across reboots.
if ! grep -q "UUID=${uuid}" /etc/fstab; then
log "adding fstab entry for ${uuid}"
printf 'UUID=%s /var/lib/gitea xfs defaults,nofail,x-systemd.device-timeout=30 0 2\n' \
"${uuid}" >> /etc/fstab
fi
systemctl daemon-reload
mountpoint -q /var/lib/gitea || mount /var/lib/gitea
mountpoint -q /var/lib/gitea || die "/var/lib/gitea failed to mount"
# Set the SELinux label persistently ONCE, rather than putting :Z on the
# quadlet's volume line. :Z would force a recursive relabel of the entire
# repository tree on every container start.
if ! semanage fcontext -l 2>/dev/null | grep -q '^/var/lib/gitea(/\.\*)?'; then
log "setting persistent SELinux fcontext on /var/lib/gitea"
semanage fcontext -a -t container_file_t '/var/lib/gitea(/.*)?' || \
warn "semanage fcontext failed; check SELinux state"
fi
restorecon -RF /var/lib/gitea || warn "restorecon failed"
# The mount point itself must belong to the container's uid, not just the
# subdirectories: Gitea creates GITEA_CUSTOM (/var/lib/gitea/custom) at
# startup, and a root-owned 0755 mount point makes that mkdir fail with a
# bare "permission denied" that reads like an SELinux problem.
chown 1000:1000 /var/lib/gitea
chmod 0750 /var/lib/gitea
install -d -o 1000 -g 1000 -m 0750 \
/var/lib/gitea/data /var/lib/gitea/log /var/lib/gitea/custom
}
# ---------------------------------------------------------------------------
# Swap
# ---------------------------------------------------------------------------
# GCE instances ship with no swap. On e2-small (2 GB) that is a real risk: a
# steady-state Gitea + Caddy/Coraza pair measures ~155 MB, but git subprocesses
# spawned during a push (git-receive-pack, index-pack, gc) push the total past
# 400 MB on a modest repo and scale with repo size. Without swap, the OOM killer
# picks a victim mid-push.
#
# This is ballast, not working memory -- hence the low swappiness. If the box is
# swapping steadily, the answer is a bigger machine type, not more swap.
setup_swap() {
local swapfile=/swapfile size_mb=2048
if swapon --show=NAME --noheadings 2>/dev/null | grep -qx "${swapfile}"; then
log "swap already active"
else
if [[ ! -f "${swapfile}" ]]; then
log "creating ${size_mb}MB swap file"
# dd, not fallocate: a fallocated file can carry unwritten extents
# that mkswap accepts and the kernel then refuses to swap to.
dd if=/dev/zero of="${swapfile}" bs=1M count="${size_mb}" status=none
chmod 0600 "${swapfile}"
mkswap "${swapfile}" >/dev/null
fi
swapon "${swapfile}" || warn "swapon failed"
fi
grep -q "^${swapfile} " /etc/fstab \
|| printf '%s none swap sw 0 0\n' "${swapfile}" >> /etc/fstab
echo 'vm.swappiness = 10' > /etc/sysctl.d/90-gitea-swappiness.conf
sysctl -q -p /etc/sysctl.d/90-gitea-swappiness.conf || warn "could not apply swappiness"
}
# ---------------------------------------------------------------------------
# Podman / netavark
# ---------------------------------------------------------------------------
configure_podman() {
log "configuring podman firewall driver"
mkdir -p /etc/containers/containers.conf.d
# Netavark keeps its rules in a dedicated `netavark` nftables table, which is
# what makes coexistence with our own table workable. Changing this with
# containers running leaves conflicting rules behind -- it is set here,
# before anything starts, and a reboot is the documented fix if it is ever
# changed on a live host.
cat > /etc/containers/containers.conf.d/10-gitea.conf <<'EOF'
[network]
firewall_driver = "nftables"
EOF
}
# ---------------------------------------------------------------------------
# Host firewall
# ---------------------------------------------------------------------------
# Runs on every invocation so a pushed vm/nftables/gitea.nft change applies
# without waiting for a reboot.
setup_nftables() {
log "installing nftables ruleset"
# firewalld and a hand-managed ruleset will fight. Pick one.
systemctl disable --now firewalld >/dev/null 2>&1 || true
systemctl mask firewalld >/dev/null 2>&1 || true
install -m 0600 "${STATE_DIR}/nftables/gitea.nft" /etc/sysconfig/nftables.conf
nft -c -f /etc/sysconfig/nftables.conf || die "nftables ruleset failed validation"
systemctl enable --now nftables
systemctl reload nftables
# If this fails, the ruleset flushed something it should not have.
nft list table inet gitea_filter >/dev/null || die "gitea_filter table missing after reload"
}
# ---------------------------------------------------------------------------
# Artifact Registry credentials for root podman
# ---------------------------------------------------------------------------
install_ar_auth() {
log "installing Artifact Registry auth refresher"
cat > /usr/local/bin/gitea-ar-auth <<'EOF'
#!/usr/bin/env bash
# Writes a docker-format auth file for Artifact Registry using the VM service
# account's metadata token.
#
# podman-auto-update.service runs as root with no interactive gcloud session, so
# it needs a credential sitting on disk. This is the #1 reason auto-update
# quietly stops working on GCE.
set -euo pipefail
AR_HOST=$(curl -fsS -H 'Metadata-Flavor: Google' \
http://169.254.169.254/computeMetadata/v1/instance/attributes/ar-host)
TOKEN=$(curl -fsS -H 'Metadata-Flavor: Google' \
http://169.254.169.254/computeMetadata/v1/instance/service-accounts/default/token \
| jq -r .access_token)
[[ -n "${TOKEN}" && "${TOKEN}" != "null" ]] || { echo "no access token from metadata server" >&2; exit 1; }
AUTH=$(printf 'oauth2accesstoken:%s' "${TOKEN}" | base64 -w0)
umask 077
tmp=$(mktemp /etc/containers/.ar-auth.XXXXXX)
jq -n --arg host "${AR_HOST}" --arg auth "${AUTH}" \
'{auths: {($host): {auth: $auth}}}' > "${tmp}"
chmod 0600 "${tmp}"
mv "${tmp}" /etc/containers/ar-auth.json
EOF
chmod 0755 /usr/local/bin/gitea-ar-auth
/usr/local/bin/gitea-ar-auth || warn "initial AR auth refresh failed"
}
# ---------------------------------------------------------------------------
# Helper scripts
# ---------------------------------------------------------------------------
install_helpers() {
log "installing helper scripts"
cat > /usr/local/bin/gitea-backup <<EOF
#!/usr/bin/env bash
# Streams a portable \`gitea dump\` straight to Cloud Storage.
#
# Complements the PD snapshot policy rather than replacing it: a dump restores
# onto any host, a snapshot only restores this disk.
set -euo pipefail
BUCKET="${BACKUP_BUCKET}"
[[ -n "\${BUCKET}" ]] || { echo "no backup bucket configured" >&2; exit 0; }
stamp=\$(date -u +%Y%m%dT%H%M%SZ)
podman exec -u 1000 gitea gitea dump -c /etc/gitea/app.ini -t /tmp -f - \\
| gcloud storage cp - "gs://\${BUCKET}/dumps/gitea-\${stamp}.zip"
echo "backup complete: gs://\${BUCKET}/dumps/gitea-\${stamp}.zip"
EOF
chmod 0755 /usr/local/bin/gitea-backup
cat > /usr/local/bin/gitea-reboot-if-needed <<'EOF'
#!/usr/bin/env bash
# Reboots only when the package layer says a reboot is genuinely required.
#
# `needs-restarting -r` exits 0 for "no reboot needed" and 1 for "reboot
# needed". Anything else -- most likely 127 because the dnf5 plugin is packaged
# differently on this release -- means we do not KNOW, and "do not know" must
# never mean "reboot the Gitea host every Sunday".
set -uo pipefail
dnf needs-restarting -r >/dev/null 2>&1
rc=$?
case "${rc}" in
0) echo "no reboot required" ; exit 0 ;;
1) echo "reboot required by pending updates -- rebooting" ; systemctl reboot ;;
*) echo "needs-restarting returned ${rc} (plugin missing?) -- NOT rebooting" >&2
echo "install the dnf needs-restarting plugin, or this check is inert" >&2
exit 0 ;;
esac
EOF
chmod 0755 /usr/local/bin/gitea-reboot-if-needed
}
# ---------------------------------------------------------------------------
# Config sync + render
# ---------------------------------------------------------------------------
sync_config() {
log "syncing configuration from gs://${CONFIG_BUCKET}/vm/"
mkdir -p "${STATE_DIR}"
gcloud storage rsync --recursive --delete-unmatched-destination-objects \
"gs://${CONFIG_BUCKET}/vm" "${STATE_DIR}" \
|| die "config sync failed"
}
# Reads a Gitea secret from Secret Manager. If it has no version yet, generates
# one -- but ONLY if a Gitea image is available locally to generate it with.
#
# INTERNAL_TOKEN must be a valid Gitea-issued JWT, so this cannot be a random
# string from Pulumi. Normally scripts/bootstrap.sh has already populated these
# before the first `pulumi up`; this is the safety net.
fetch_or_create_secret() {
local name="$1" value=""
if value=$(gcloud secrets versions access latest --secret="${name}" --project="${GCP_PROJECT}" 2>/dev/null); then
printf '%s' "${value}"
return 0
fi
if ! podman image exists "${IMAGE_GITEA}" 2>/dev/null; then
return 1
fi
local key
case "${name}" in
*secret-key) key=SECRET_KEY ;;
*internal-token) key=INTERNAL_TOKEN ;;
*oauth2-jwt-secret) key=JWT_SECRET ;;
*lfs-jwt-secret) key=LFS_JWT_SECRET ;;
*) return 1 ;;
esac
log "generating missing secret ${name}"
value=$(podman run --rm "${IMAGE_GITEA}" generate secret "${key}") || return 1
printf '%s' "${value}" \
| gcloud secrets versions add "${name}" --project="${GCP_PROJECT}" --data-file=- >/dev/null || return 1
printf '%s' "${value}"
}
# Renders src -> dst only if the content actually differs, and reports whether
# it changed. Keeps config-sync from restarting healthy services for no reason.
render() {
local src="$1" dst="$2" owner="$3" mode="$4" vars="$5"
local tmp
tmp=$(mktemp)
envsubst "${vars}" < "${src}" > "${tmp}"
if [[ -f "${dst}" ]] && cmp -s "${tmp}" "${dst}"; then
rm -f "${tmp}"
return 1
fi
install -o "${owner%:*}" -g "${owner#*:}" -m "${mode}" "${tmp}" "${dst}"
rm -f "${tmp}"
log "rendered ${dst}"
return 0
}
# Decides whether Caddy can live on the podman bridge or needs the host network.
#
# The googleclouddns ACME plugin authenticates via Application Default
# Credentials, which on GCE means reaching the metadata server at
# 169.254.169.254. If a container on our bridge cannot reach it, DNS-01 issuance
# fails at certificate time -- long after this script has reported success -- so
# the check happens here, up front, and the answer is cached.
probe_caddy_network() {
local cache=/etc/gitea/caddy-network
if [[ -f "${cache}" ]]; then
cat "${cache}"
return 0
fi
if ! podman image exists "${IMAGE_GITEA}" 2>/dev/null; then
# Cannot probe yet (first boot, before the first image build). Assume the
# bridge and re-probe on the next config-sync.
echo bridge
return 0
fi
# The network normally does not exist yet: on a full boot this runs before
# the quadlet units start. Creating it here with the same arguments quadlet
# uses keeps the probe honest -- otherwise `podman run --network gitea`
# fails and the probe wrongly concludes the metadata server is unreachable.
# stdout is discarded because this function's stdout IS the return value.
podman network create --ignore \
--subnet "${PODMAN_SUBNET}" --gateway "${PODMAN_GATEWAY}" gitea >/dev/null 2>&1 || true
local mode=host
if podman run --rm --network gitea --entrypoint curl "${IMAGE_GITEA}" \
-fsS -m 10 -H 'Metadata-Flavor: Google' \
http://169.254.169.254/computeMetadata/v1/instance/service-accounts/default/token \
>/dev/null 2>&1; then
mode=bridge
else
warn "metadata server unreachable from the podman bridge -- Caddy will use the host network"
fi
mkdir -p /etc/gitea
echo "${mode}" > "${cache}"
echo "${mode}"
}
render_all() {
log "rendering configuration"
mkdir -p /etc/gitea /etc/caddy "${RENDER_DIR}"
local secret_key internal_token oauth2_jwt lfs_jwt
if ! secret_key=$(fetch_or_create_secret gitea-secret-key) \
|| ! internal_token=$(fetch_or_create_secret gitea-internal-token) \
|| ! oauth2_jwt=$(fetch_or_create_secret gitea-oauth2-jwt-secret) \
|| ! lfs_jwt=$(fetch_or_create_secret gitea-lfs-jwt-secret); then
# Deliberately NOT a failure: on first boot the image does not exist yet
# and the secrets may not be populated. Writing an app.ini with empty
# SECRET_KEY/INTERNAL_TOKEN would be far worse than doing nothing --
# Gitea would come up with broken sessions and tokens.
warn "Gitea secrets unavailable -- skipping config render (will retry on the next sync)"
return 0
fi
local caddy_mode caddy_network caddy_publish caddy_sysctl gitea_upstream
caddy_mode=$(probe_caddy_network)
if [[ "${caddy_mode}" == "host" ]]; then
caddy_network="host"
caddy_publish="# Network=host: ports are bound directly, publishing would be invalid."
caddy_sysctl="# Network=host: podman rejects net.* sysctls; set on the host instead."
gitea_upstream="127.0.0.1:3000"
# Container-local sysctls are unavailable on the host network, so allow
# unprivileged binds to 80/443 host-wide. Narrower than granting the
# container CAP_NET_BIND_SERVICE.
echo 'net.ipv4.ip_unprivileged_port_start = 80' > /etc/sysctl.d/90-gitea-caddy.conf
sysctl -q -p /etc/sysctl.d/90-gitea-caddy.conf || warn "could not apply unprivileged port sysctl"
else
caddy_network="gitea.network"
caddy_publish=$'PublishPort=80:80\nPublishPort=443:443\nPublishPort=443:443/udp'
caddy_sysctl="Sysctl=net.ipv4.ip_unprivileged_port_start=0"
gitea_upstream="gitea:3000"
rm -f /etc/sysctl.d/90-gitea-caddy.conf
fi
# Trust both the bridge CIDR and loopback so this value stays correct in
# either Caddy networking mode. Rootful podman SNATs host-loopback traffic
# to the bridge gateway, so the CIDR covers the host-network case too.
local trusted_proxies="${PODMAN_SUBNET},127.0.0.1/32"
local changed=0
export GITEA_SECRET_KEY="${secret_key}" \
GITEA_INTERNAL_TOKEN="${internal_token}" \
GITEA_OAUTH2_JWT_SECRET="${oauth2_jwt}" \
GITEA_LFS_JWT_SECRET="${lfs_jwt}" \
TRUSTED_PROXIES="${trusted_proxies}" \
GITEA_UPSTREAM="${gitea_upstream}" \
CADDY_NETWORK="${caddy_network}" \
CADDY_PUBLISH_PORTS="${caddy_publish}" \
CADDY_SYSCTL="${caddy_sysctl}"
# app.ini is 0400 owned by uid 1000: it holds SECRET_KEY and INTERNAL_TOKEN,
# and the container runs as that uid and must be able to read it.
render "${STATE_DIR}/config/app.ini.tmpl" /etc/gitea/app.ini 1000:1000 0400 \
'${APP_NAME} ${DOMAIN} ${GITEA_SECRET_KEY} ${GITEA_INTERNAL_TOKEN} ${GITEA_OAUTH2_JWT_SECRET} ${GITEA_LFS_JWT_SECRET} ${TRUSTED_PROXIES} ${REQUIRE_SIGNIN_VIEW}' \
&& changed=1
render "${STATE_DIR}/config/Caddyfile.tmpl" /etc/caddy/Caddyfile root:root 0644 \
'${DOMAIN} ${ACME_EMAIL} ${GITEA_UPSTREAM} ${WAF_MODE}' \
&& changed=1
local unit
for unit in gitea.network gitea.container caddy.container caddy-data.volume caddy-config.volume; do
render "${STATE_DIR}/quadlets/${unit}" "${RENDER_DIR}/${unit}" root:root 0644 \
'${IMAGE_GITEA} ${IMAGE_CADDY} ${PODMAN_SUBNET} ${PODMAN_GATEWAY} ${GCP_PROJECT} ${CADDY_NETWORK} ${CADDY_PUBLISH_PORTS} ${CADDY_SYSCTL}' \
&& changed=1
done
for unit in "${STATE_DIR}"/systemd/*; do
[[ -f "${unit}" ]] || continue
local base; base=$(basename "${unit}")
if ! cmp -s "${unit}" "/etc/systemd/system/${base}"; then
install -m 0644 "${unit}" "/etc/systemd/system/${base}"
log "installed unit ${base}"
changed=1
fi
done
unset GITEA_SECRET_KEY GITEA_INTERNAL_TOKEN GITEA_OAUTH2_JWT_SECRET GITEA_LFS_JWT_SECRET
systemctl daemon-reload
if (( changed )); then
log "configuration changed -- restarting services"
systemctl restart gitea.service || warn "gitea did not restart cleanly"
systemctl restart caddy.service || warn "caddy did not restart cleanly"
else
log "configuration unchanged"
fi
}
# ---------------------------------------------------------------------------
# fail2ban
# ---------------------------------------------------------------------------
# Runs on every invocation, not just full boots: the jail's enabled flag is
# derived from WAF_MODE, and that has to be re-applied whenever the mode changes.
setup_fail2ban() {
command -v fail2ban-server >/dev/null 2>&1 || { warn "fail2ban not installed; skipping"; return 0; }
log "configuring fail2ban"
install -m 0644 "${STATE_DIR}/fail2ban/action.d/nft-prerouting.conf" /etc/fail2ban/action.d/
install -m 0644 "${STATE_DIR}/fail2ban/filter.d/gitea.conf" /etc/fail2ban/filter.d/
install -m 0644 "${STATE_DIR}/fail2ban/filter.d/caddy-coraza.conf" /etc/fail2ban/filter.d/
# One config value drives both halves: the WAF only blocks in On, and only
# then is it safe to escalate a WAF verdict into an nftables ban. In
# DetectionOnly the jail is inert so tuning cannot lock anyone out.
local coraza_enabled=false
[[ "${WAF_MODE}" == "On" ]] && coraza_enabled=true
export CORAZA_JAIL_ENABLED="${coraza_enabled}"
log "coraza fail2ban jail enabled=${coraza_enabled} (waf-mode=${WAF_MODE})"
envsubst '${PODMAN_SUBNET} ${CORAZA_JAIL_ENABLED}' \
< "${STATE_DIR}/fail2ban/jail.d/gitea.local" > /etc/fail2ban/jail.d/gitea.local
chmod 0644 /etc/fail2ban/jail.d/gitea.local
systemctl enable --now fail2ban
systemctl reload fail2ban || systemctl restart fail2ban
}
# ---------------------------------------------------------------------------
# Automatic updates
# ---------------------------------------------------------------------------
setup_auto_updates() {
log "configuring automatic updates"
install -m 0644 "${STATE_DIR}/dnf/automatic.conf" /etc/dnf/automatic.conf
# AlmaLinux 10 ships dnf5, where the unit is dnf5-automatic.timer -- but the
# package providing it has moved around between releases, so resolve it
# instead of hardcoding a name that may not exist.
local timer=""
for candidate in dnf5-automatic.timer dnf-automatic.timer; do
if systemctl list-unit-files "${candidate}" >/dev/null 2>&1 \
&& systemctl cat "${candidate}" >/dev/null 2>&1; then
timer="${candidate}"; break
fi
done
if [[ -z "${timer}" ]]; then
log "no dnf automatic timer present -- installing provider"
dnf -y install "$(dnf -q provides '*/dnf5-automatic.timer' 2>/dev/null | awk 'NR==1{print $1}')" \
|| dnf -y install dnf-automatic \
|| warn "could not install a dnf-automatic provider"
for candidate in dnf5-automatic.timer dnf-automatic.timer; do
systemctl cat "${candidate}" >/dev/null 2>&1 && { timer="${candidate}"; break; }
done
fi
[[ -n "${timer}" ]] && systemctl enable --now "${timer}" || warn "no dnf automatic timer enabled"
systemctl enable --now podman-auto-update.timer
}
enable_units() {
log "enabling units"
systemctl daemon-reload
systemctl enable --now gitea-ar-auth.timer
systemctl enable --now gitea-backup.timer
systemctl enable --now gitea-reboot-window.timer
# Quadlet-generated units are not "enabled" in the usual sense -- the
# [Install] section is honoured by the generator at daemon-reload time.
systemctl start gitea.service || warn "gitea not started yet (expected before the first image build)"
systemctl start caddy.service || warn "caddy not started yet (expected before the first image build)"
}
# ---------------------------------------------------------------------------
# Main
# ---------------------------------------------------------------------------
main() {
log "starting (mode=${MODE})"
if [[ "${MODE}" == "full" ]]; then
install_packages
setup_data_disk
setup_swap
configure_podman
install_helpers
fi
sync_config
# Make the synced copy the canonical one, so gitea-config-sync.service always
# runs the version that matches the config in the bucket.
#
# Under --sync-only this file IS the script bash is currently reading. GNU
# install truncates in place and bash reads scripts incrementally, so a
# naive copy can rewrite the interpreter's input mid-execution -- exactly in
# the case this mechanism exists for (a vm/bootstrap.sh change). Skip when
# identical, and otherwise replace via atomic rename onto a fresh inode so
# the running process keeps reading the old one.
if ! cmp -s "${STATE_DIR}/bootstrap.sh" /usr/local/sbin/gitea-bootstrap; then
install -m 0755 "${STATE_DIR}/bootstrap.sh" /usr/local/sbin/.gitea-bootstrap.new
mv -f /usr/local/sbin/.gitea-bootstrap.new /usr/local/sbin/gitea-bootstrap
log "updated /usr/local/sbin/gitea-bootstrap"
fi
if [[ "${MODE}" == "full" ]]; then
install_ar_auth
fi
# setup_nftables and setup_fail2ban run in BOTH modes, deliberately.
#
# They apply configuration that lives in the vm/ tree, so gating them on a
# full boot would mean a pushed change never takes effect until the next
# reboot. The fail2ban case is the dangerous one: flipping gitea:wafMode to
# On re-renders the Caddyfile through render_all and Caddy starts issuing
# 403s, but the jail that acts on them would stay enabled=false -- a WAF
# that blocks and a ban that never happens, with nothing in the logs to say
# so. Both functions are idempotent and self-validating (`nft -c` before
# load, `command -v fail2ban-server` before touching fail2ban).
setup_nftables
setup_fail2ban
render_all
if [[ "${MODE}" == "full" ]]; then
setup_auto_updates
enable_units
fi
log "done"
}
main "$@"
+111
View File
@@ -0,0 +1,111 @@
# Rendered by vm/bootstrap.sh -> /etc/caddy/Caddyfile
#
# Two compile-time plugins are doing the work here:
#
# googleclouddns -- ACME DNS-01, so issuance and renewal never need inbound 80
# coraza_waf -- OWASP Coraza with the Core Rule Set embedded in the binary
{
email ${ACME_EMAIL}
admin 127.0.0.1:2019
# Required by coraza-caddy: Caddy has no built-in ordering for a third-party
# directive, and the WAF must run before anything that could act on the
# request. Applies inside handle blocks too.
order coraza_waf first
}
${DOMAIN} {
tls {
dns googleclouddns {
# Application Default Credentials come from the GCE metadata server.
# The VM service account holds roles/dns.admin scoped to this zone only.
gcp_project {env.GCP_PROJECT}
}
# Only used for propagation checks. If issuance stalls waiting for
# propagation, this is the knob to turn.
resolvers 8.8.8.8 8.8.4.4
}
encode zstd gzip
# Governs the git/LFS branch below. The WAF branch has its own, much smaller,
# SecRequestBodyLimit -- they apply to different routes and are not a mismatch
# to be "fixed".
request_body {
max_size 512MB
}
# ---------------------------------------------------------------------
# Git transport and LFS: deliberately NOT behind the WAF.
#
# Packfiles are binary and trip CRS's SQLi/XSS rules constantly, and
# coraza.conf-recommended's SecRequestBodyLimit would reject any push
# larger than ~12MB outright. Running CRS here does not harden anything;
# it just breaks git. Authentication still applies -- Gitea does that.
# ---------------------------------------------------------------------
@gittransport path_regexp ^/[^/]+/[^/]+/(?:info/refs|git-upload-pack|git-receive-pack|HEAD|objects/.*|info/lfs(?:/.*)?)$
handle @gittransport {
reverse_proxy ${GITEA_UPSTREAM} {
transport http {
read_timeout 900s
write_timeout 900s
}
}
}
# ---------------------------------------------------------------------
# Everything else -- web UI and API -- goes through the WAF.
# ---------------------------------------------------------------------
handle {
coraza_waf {
load_owasp_crs
directives `
Include @coraza.conf-recommended
Include @crs-setup.conf.example
Include @owasp_crs/*.conf
# DetectionOnly logs what it would have blocked without blocking.
# The fail2ban jail is enabled ONLY when this is On -- banning on
# detections that were never blocks would turn a false positive
# into an nftables ban, which is worse than the 403 it avoided.
SecRuleEngine ${WAF_MODE}
# Inspecting responses on a git host costs CPU and catches nothing
# worth catching.
SecResponseBodyAccess Off
# Truncate-and-inspect rather than reject: a large but legitimate
# attachment upload should not 413 because the WAF gave up.
SecRequestBodyLimitAction ProcessPartial
# RelevantOnly, not On. Auditing every request would pour the full
# request volume into journald and then into Cloud Logging.
SecAuditEngine RelevantOnly
# The default relevant-status is ^(?:5|4(?!04)), which audits every
# 401 -- and on a public git host unauthenticated API and web probes
# produce those constantly. That noise would bury the would-be blocks
# this log exists to surface. Narrowed to real WAF refusals and
# server errors; rule-triggered transactions are still audited on
# severity regardless of status, which is what keeps DetectionOnly
# useful.
SecAuditLogRelevantStatus "^(?:5[0-9]{2}|403)$"
SecAuditLog /dev/stdout
SecAuditLogFormat json
SecAuditLogParts ABIJDEFHZ
`
}
reverse_proxy ${GITEA_UPSTREAM} {
transport http {
read_timeout 900s
write_timeout 900s
}
}
}
log {
output stderr
format json
}
}
+107
View File
@@ -0,0 +1,107 @@
; Rendered by vm/bootstrap.sh -> /etc/gitea/app.ini (owner 1000:1000, mode 0400).
;
; This file is FULLY MANAGED. Gitea's GITEA__section__KEY environment support
; comes from the upstream image's `environment-to-ini` entrypoint helper, which
; does not exist on our Debian base -- so configuration happens here, on the
; host, and the file is bind-mounted read-only.
;
; INSTALL_LOCK=true means the web installer is never reachable. Editing Gitea
; settings through the UI that map to app.ini will NOT persist; change the
; template in the repo and let the config-sync pipeline re-render it.
APP_NAME = ${APP_NAME}
RUN_USER = git
RUN_MODE = prod
WORK_PATH = /var/lib/gitea
[server]
PROTOCOL = http
HTTP_ADDR = 0.0.0.0
HTTP_PORT = 3000
DOMAIN = ${DOMAIN}
ROOT_URL = https://${DOMAIN}/
APP_DATA_PATH = /var/lib/gitea/data
DISABLE_SSH = false
; Built-in Go SSH server -- no sshd inside the container, and the host's sshd
; keeps port 22 for OS Login / IAP admin access.
START_SSH_SERVER = true
BUILTIN_SSH_SERVER_USER = git
SSH_DOMAIN = ${DOMAIN}
SSH_LISTEN_HOST = 0.0.0.0
SSH_LISTEN_PORT = 2222
; Advertised in clone URLs; must match the published host port.
SSH_PORT = 2222
LFS_START_SERVER = true
LFS_JWT_SECRET = ${GITEA_LFS_JWT_SECRET}
OFFLINE_MODE = true
[database]
DB_TYPE = sqlite3
PATH = /var/lib/gitea/data/gitea.db
; WAL is what makes SQLite tolerable under concurrent reads.
SQLITE_JOURNAL_MODE = WAL
SQLITE_TIMEOUT = 500
[repository]
ROOT = /var/lib/gitea/data/gitea-repositories
DEFAULT_BRANCH = main
DEFAULT_PRIVATE = private
DISABLE_HTTP_GIT = false
[repository.upload]
TEMP_PATH = /var/lib/gitea/data/tmp/uploads
[lfs]
PATH = /var/lib/gitea/data/lfs
[security]
INSTALL_LOCK = true
SECRET_KEY = ${GITEA_SECRET_KEY}
INTERNAL_TOKEN = ${GITEA_INTERNAL_TOKEN}
; Without these two, Gitea sees Caddy's address as the client for every request
; and fail2ban ends up banning the reverse proxy, locking everyone out.
REVERSE_PROXY_TRUSTED_PROXIES = ${TRUSTED_PROXIES}
REVERSE_PROXY_LIMIT = 1
PASSWORD_HASH_ALGO = argon2
[oauth2]
JWT_SECRET = ${GITEA_OAUTH2_JWT_SECRET}
[service]
DISABLE_REGISTRATION = true
REQUIRE_SIGNIN_VIEW = ${REQUIRE_SIGNIN_VIEW}
REGISTER_EMAIL_CONFIRM = false
ENABLE_NOTIFY_MAIL = false
ALLOW_ONLY_EXTERNAL_REGISTRATION = false
ENABLE_CAPTCHA = false
DEFAULT_KEEP_EMAIL_PRIVATE = true
DEFAULT_ALLOW_CREATE_ORGANIZATION = true
DEFAULT_ENABLE_TIMETRACKING = true
[session]
PROVIDER = file
PROVIDER_CONFIG = /var/lib/gitea/data/sessions
COOKIE_SECURE = true
[mailer]
ENABLED = false
[log]
; console -> journald -> Cloud Logging, and it is what fail2ban's systemd
; backend reads via CONTAINER_NAME=gitea. Do not switch to file logging without
; updating vm/fail2ban/jail.d/gitea.local.
MODE = console
LEVEL = info
ROOT_PATH = /var/lib/gitea/log
[actions]
; Enabled now so the eventual migration off GitHub Cloud Build triggers onto
; Gitea Actions does not need a config change + restart.
ENABLED = true
DEFAULT_ACTIONS_URL = github
[cron.update_checker]
ENABLED = false
[ui]
DEFAULT_THEME = gitea-auto
+27
View File
@@ -0,0 +1,27 @@
# Installed by vm/bootstrap.sh as /etc/dnf/automatic.conf.
# Consumed by dnf5-automatic.timer (dnf-automatic.timer on a dnf4 system).
[commands]
# `default` follows the distro's notion of what an upgrade is; switch to
# `security` if you want to narrow the blast radius of unattended patching.
upgrade_type = default
# Spread the herd. Every AlmaLinux box on the internet firing at 03:30 is how
# mirrors fall over.
random_sleep = 3600
network_online_timeout = 60
download_updates = yes
apply_updates = yes
# Never reboot from here. gitea-reboot-window.timer owns that decision, so
# restarts land in a known window with a fresh disk snapshot behind them.
reboot = never
[emitters]
# stdio -> journald -> Cloud Logging. No mail configured on this host.
emit_via = stdio
system_name = gitea-vm
[base]
debuglevel = 1
+48
View File
@@ -0,0 +1,48 @@
# fail2ban action: drop banned sources at the nftables PREROUTING hook.
#
# Why this exists instead of the stock `nftables` action:
#
# Traffic to a published container port (443, 2222) is DNAT'd in prerouting and
# then traverses FORWARD -- it never reaches the INPUT hook. fail2ban's default
# actions install their rules on INPUT, so on a podman host they drop nothing
# at all while `fail2ban-client status` cheerfully reports active bans. It fails
# silently, which is the worst way for a security control to fail.
#
# Hooking at prerouting with priority -300 (raw) puts us ahead of both the DNAT
# that redirects the packet and netavark's own chains, so the drop applies to
# host-terminated and container-bound traffic alike.
#
# Verify with a real ban drill -- see docs/runbook.md.
[Definition]
actionstart = nft add table inet f2b-prerouting
nft add chain inet f2b-prerouting preroute '{ type filter hook prerouting priority -300 ; policy accept ; }'
nft add set inet f2b-prerouting f2b-<name>-v4 '{ type ipv4_addr ; flags interval ; }'
nft add set inet f2b-prerouting f2b-<name>-v6 '{ type ipv6_addr ; flags interval ; }'
nft add rule inet f2b-prerouting preroute ip saddr @f2b-<name>-v4 drop
nft add rule inet f2b-prerouting preroute ip6 saddr @f2b-<name>-v6 drop
# Flush rather than delete: the preroute rules still reference these sets, so a
# delete would be refused. Emptying them drops every ban, which is what stopping
# the jail means.
actionstop = nft flush set inet f2b-prerouting f2b-<name>-v4 2>/dev/null || true
nft flush set inet f2b-prerouting f2b-<name>-v6 2>/dev/null || true
# If something flushed our table out from under us, actionstart runs again.
actioncheck = nft list set inet f2b-prerouting f2b-<name>-v4 >/dev/null 2>&1
# Family is derived from the address itself rather than a fail2ban tag, so this
# stays correct across fail2ban versions that disagree about <family>.
actionban = case "<ip>" in \
*:*) nft add element inet f2b-prerouting f2b-<name>-v6 '{ <ip> }' ;; \
*) nft add element inet f2b-prerouting f2b-<name>-v4 '{ <ip> }' ;; \
esac
actionunban = case "<ip>" in \
*:*) nft delete element inet f2b-prerouting f2b-<name>-v6 '{ <ip> }' 2>/dev/null || true ;; \
*) nft delete element inet f2b-prerouting f2b-<name>-v4 '{ <ip> }' 2>/dev/null || true ;; \
esac
[Init]
name = default
+31
View File
@@ -0,0 +1,31 @@
# Matches Coraza's BLOCKING DECISION on the Caddy log (LogDriver=journald).
#
# Written from real output, not documentation. Coraza emits two shapes:
#
# Warning -- one line per individual CRS rule that matched. A single
# one of these means almost nothing; a legitimate issue
# comment containing SQL will produce several.
# Access denied -- emitted once, by the anomaly-score threshold rule (949110),
# when the accumulated score actually crosses the limit and
# the request is refused.
#
# Only the second is matched. Banning on individual rule hits would ban users
# for pasting a code snippet.
#
# Sample line this is built against:
# {"level":"warn",...,"logger":"http.handlers.waf","msg":"[client \"203.0.113.9\"]
# Coraza: Access denied (phase 2). Inbound Anomaly Score Exceeded (Total Score: 23)
# [file \"@owasp_crs/REQUEST-949-BLOCKING-EVALUATION.conf\"] ... [id \"949110\"] ..."}
#
# NOTE: in DetectionOnly mode Coraza never emits "Access denied" -- nothing is
# refused -- so this filter matches nothing and the jail is inert. That is
# intentional and is why vm/bootstrap.sh only enables the jail when
# SecRuleEngine is On.
[Definition]
failregex = ^.*"logger":"http\.handlers\.waf".*\[client \\"<HOST>\\"\] Coraza: Access denied
ignoreregex =
datepattern = {^LN-BEG}
+29
View File
@@ -0,0 +1,29 @@
# Matches Gitea authentication failures on stdout (LogDriver=journald).
#
# The patterns below were taken from real log output, not from documentation:
#
# [W] Failed authentication attempt for someuser from 203.0.113.77:0: user does not exist
# [W] Failed authentication attempt from 198.51.100.9:48236
#
# The first is the web UI / API; the second is the built-in SSH server. They
# share an auth path, which is why one jail covers both ports rather than
# guessing at separate per-transport formats.
#
# <HOST> resolves to the real client ONLY because app.ini sets
# REVERSE_PROXY_TRUSTED_PROXIES to the podman subnet. Without it, Gitea logs
# Caddy's address and every ban would hit the reverse proxy.
#
# Deliberately NOT matched: "Failed connection from <HOST> with error: EOF".
# Port scanners and health probes produce it constantly and it says nothing
# about credentials.
[Definition]
failregex = ^.*Failed authentication attempt for .+? from <HOST>
^.*Failed authentication attempt from <HOST>
^.*[Ii]nvalid credentials from <HOST>
^.*Attempted OAuth2 login .*? from <HOST>
ignoreregex =
datepattern = {^LN-BEG}
+53
View File
@@ -0,0 +1,53 @@
# Installed by vm/bootstrap.sh as /etc/fail2ban/jail.d/gitea.local
[DEFAULT]
# Podman's journald log driver tags container output with CONTAINER_NAME, so
# there are no log files to bind-mount, rotate, or keep in sync.
backend = systemd
# Every ban goes through the prerouting action. The stock nftables/iptables
# actions install INPUT rules, which do not see DNAT'd container traffic at all.
banaction = nft-prerouting
banaction_allports = nft-prerouting
bantime = 1h
findtime = 10m
maxretry = 5
# Never lock out loopback, the podman bridge, or the IAP range -- IAP is the
# only way back into this box.
ignoreip = 127.0.0.1/8 ::1 ${PODMAN_SUBNET} 35.235.240.0/20
[sshd]
enabled = true
port = 22
# Belt-and-braces: port 22 already only accepts the IAP range, at both the VPC
# firewall and nftables.
maxretry = 5
[gitea]
enabled = true
filter = gitea
# One jail across all three ports rather than separate web/ssh jails: Gitea
# serves them from one process with shared auth logging, and a bad actor should
# lose every door at once, not just the one they knocked on.
port = 80,443,2222
journalmatch = CONTAINER_NAME=gitea
maxretry = 5
bantime = 1h
[caddy-coraza]
# Enabled ONLY when gitea:wafMode is "On". In DetectionOnly the WAF reports what
# it would have blocked but lets it through, and escalating those reports into
# nftables bans would turn a tuning false positive into a locked-out user --
# strictly worse than the 403 that DetectionOnly was chosen to avoid.
# vm/bootstrap.sh derives this from the same config value as SecRuleEngine.
enabled = ${CORAZA_JAIL_ENABLED}
filter = caddy-coraza
port = 80,443
journalmatch = CONTAINER_NAME=caddy
# Deliberately looser than the auth jail. A WAF verdict is a weaker signal than
# a failed password, so it takes sustained hostile traffic to earn a ban.
maxretry = 10
findtime = 10m
bantime = 2h
+54
View File
@@ -0,0 +1,54 @@
#!/usr/sbin/nft -f
#
# Host firewall. Installed by vm/bootstrap.sh as /etc/sysconfig/nftables.conf.
#
# CRITICAL: this file must NEVER contain `flush ruleset`. The stock nftables
# config ships with one, and it would wipe podman/netavark's NAT and forward
# rules -- silently breaking all container networking on every reload. We touch
# exactly one table and nothing else.
#
# Equally deliberate: there is NO forward chain here. Netavark manages forward
# and nat rules in its own `netavark` table. A drop-policy forward chain in this
# table would drop container traffic regardless of what netavark allows, because
# a packet is dropped if any chain drops it.
#
# Traffic to published container ports (443, 2222) is DNAT'd in prerouting and
# traverses forward, never input -- so the input rules below only actually
# govern host-terminated traffic. Blocking abusive clients from container ports
# is fail2ban's job, via the prerouting-hook action in
# vm/fail2ban/action.d/nft-prerouting.conf.
# Create-then-delete makes this reload idempotent: the bare `table` line is a
# no-op if it already exists, so the delete can never fail on a fresh boot.
table inet gitea_filter
delete table inet gitea_filter
table inet gitea_filter {
chain input {
type filter hook input priority filter; policy drop;
iif "lo" accept comment "loopback"
ct state established,related accept
ct state invalid drop
meta l4proto icmp accept comment "IPv4 ICMP"
meta l4proto ipv6-icmp accept comment "IPv6 ICMP / NDP"
# DHCP renewal from the GCE metadata server.
udp sport 67 udp dport 68 accept
# Admin SSH: IAP TCP forwarding range only. There is no other path in --
# the VPC firewall enforces the same restriction as the outer layer.
ip saddr 35.235.240.0/20 tcp dport 22 accept comment "IAP SSH"
# Only reached when Caddy runs with Network=host (the metadata-server
# fallback path). Harmless otherwise: on the bridge these are DNAT'd
# before input and never match here.
tcp dport { 80, 443 } accept comment "HTTP/HTTPS"
udp dport 443 accept comment "HTTP/3"
tcp dport 2222 accept comment "git over SSH"
# Rate-limited logging so a scan cannot fill the disk via journald.
limit rate 5/minute burst 10 packets log prefix "nft-drop-in: " level info
}
}
+6
View File
@@ -0,0 +1,6 @@
[Unit]
Description=Caddy autosaved JSON config storage
[Volume]
VolumeName=caddy-config
Label=app=gitea
+9
View File
@@ -0,0 +1,9 @@
# ACME account keys and issued certificates. Losing this means re-issuing certs
# (and burning Let's Encrypt rate limit), so it is a named volume rather than
# a tmpfs or an anonymous mount.
[Unit]
Description=Caddy certificate and ACME account storage
[Volume]
VolumeName=caddy-data
Label=app=gitea
+62
View File
@@ -0,0 +1,62 @@
# Rendered by vm/bootstrap.sh into /etc/containers/systemd/
#
# ${CADDY_NETWORK} is chosen by bootstrap.sh at run time, not by hand:
# the googleclouddns ACME plugin authenticates via Application Default
# Credentials, which on GCE means reaching the metadata server at
# 169.254.169.254. bootstrap.sh probes whether a container on the gitea bridge
# can do that and falls back to Network=host if it cannot. See
# gitea-probe-metadata in bootstrap.sh and docs/runbook.md.
[Unit]
Description=Caddy (TLS termination, ACME DNS-01 via Google Cloud DNS)
Documentation=https://caddyserver.com/docs/
After=gitea.service gitea-ar-auth.service network-online.target
Wants=gitea.service gitea-ar-auth.service
[Container]
ContainerName=caddy
Image=${IMAGE_CADDY}
AutoUpdate=registry
# Registry auth for Artifact Registry. Two settings, because two different
# code paths need it: PodmanArgs covers `podman run`'s pull, and the
# io.containers.autoupdate.authfile label is what `podman auto-update` reads
# when it checks the registry digest. (There is no AuthFile= key in the
# [Container] group -- that one only exists for .image and .build units.)
PodmanArgs=--authfile=/etc/containers/ar-auth.json
Label=io.containers.autoupdate.authfile=/etc/containers/ar-auth.json
Network=${CADDY_NETWORK}
LogDriver=journald
${CADDY_PUBLISH_PORTS}
Volume=/etc/caddy/Caddyfile:/etc/caddy/Caddyfile:ro,Z
Volume=caddy-data.volume:/data
Volume=caddy-config.volume:/config
User=1000:1000
# Caddy binds 80/443 as a non-root user. On the bridge this is a container-local
# sysctl; podman REJECTS net.* sysctls when Network=host, so bootstrap.sh emits
# nothing here in that mode and sets the equivalent host sysctl instead. File
# capabilities are not an option -- NoNewPrivileges blocks them.
${CADDY_SYSCTL}
# The plugin reads Application Default Credentials; the project must be explicit
# because the metadata server's project and the DNS zone's project need not match.
Environment=GCP_PROJECT=${GCP_PROJECT}
HealthCmd=curl -fsS http://127.0.0.1:2019/config/
HealthInterval=30s
HealthTimeout=5s
HealthStartPeriod=30s
HealthRetries=3
Notify=healthy
NoNewPrivileges=true
[Service]
Restart=always
RestartSec=30
StartLimitIntervalSec=0
TimeoutStartSec=300
[Install]
WantedBy=multi-user.target
+69
View File
@@ -0,0 +1,69 @@
# Rendered by vm/bootstrap.sh into /etc/containers/systemd/
[Unit]
Description=Gitea
Documentation=https://docs.gitea.com/
# Requires= (not just After=) on the mount: without it podman happily creates an
# empty /var/lib/gitea and Gitea initialises a FRESH install on top of the
# unmounted path, which looks like total data loss.
Requires=var-lib-gitea.mount
# Wants= (not just After=): the credential file has to be written before the
# first pull, and the refresh timer alone would not guarantee that at boot.
Wants=gitea-ar-auth.service
After=var-lib-gitea.mount gitea-ar-auth.service network-online.target
[Container]
ContainerName=gitea
Image=${IMAGE_GITEA}
# Floating tag, deliberately. AutoUpdate=registry compares the local digest
# against the registry's digest FOR A TAG; a digest-pinned image would give it
# nothing to poll and the feature would be silently dead.
AutoUpdate=registry
# Registry auth for Artifact Registry. Two settings, because two different
# code paths need it: PodmanArgs covers `podman run`'s pull, and the
# io.containers.autoupdate.authfile label is what `podman auto-update` reads
# when it checks the registry digest. (There is no AuthFile= key in the
# [Container] group -- that one only exists for .image and .build units.)
PodmanArgs=--authfile=/etc/containers/ar-auth.json
Label=io.containers.autoupdate.authfile=/etc/containers/ar-auth.json
Network=gitea.network
LogDriver=journald
# Git over SSH, public. HTTP is loopback-only: Caddy is the only thing that
# should reach it, and binding it to 127.0.0.1 keeps the reverse-proxy wiring
# identical whether Caddy runs on the bridge or on the host network.
PublishPort=2222:2222
PublishPort=127.0.0.1:3000:3000
# No :z/:Z on the data volume -- bootstrap.sh sets a persistent SELinux fcontext
# for it instead. A relabel flag here would force a recursive relabel of the
# entire repository tree on every single container start.
Volume=/var/lib/gitea:/var/lib/gitea
Volume=/etc/gitea/app.ini:/etc/gitea/app.ini:ro,Z
User=1000:1000
Environment=GITEA_WORK_DIR=/var/lib/gitea
# Notify=healthy is what arms auto-update rollback. Rollback only fires when the
# restarted unit fails to START; without this a container that starts and is
# broken would never roll back.
HealthCmd=curl -fsS http://127.0.0.1:3000/api/healthz
HealthInterval=30s
HealthTimeout=5s
HealthStartPeriod=60s
HealthRetries=3
Notify=healthy
NoNewPrivileges=true
DropCapability=ALL
[Service]
Restart=always
# On first boot the :prod image does not exist yet (it is built by the first
# Cloud Build run). StartLimitIntervalSec=0 lets the unit retry indefinitely
# instead of hitting the start limit and parking in `failed` forever.
RestartSec=30
StartLimitIntervalSec=0
TimeoutStartSec=300
[Install]
WantedBy=multi-user.target
+15
View File
@@ -0,0 +1,15 @@
# Rendered by vm/bootstrap.sh into /etc/containers/systemd/
#
# The subnet is PINNED deliberately. Left to podman it is drawn from the default
# pool and is not stable across a network recreate -- and this CIDR is exactly
# what Gitea's REVERSE_PROXY_TRUSTED_PROXIES has to name. An unpinned subnet
# means a network recreate silently reverts every fail2ban ban to targeting
# Caddy instead of the real client.
[Unit]
Description=Gitea container network
[Network]
NetworkName=gitea
Subnet=${PODMAN_SUBNET}
Gateway=${PODMAN_GATEWAY}
Label=app=gitea
+20
View File
@@ -0,0 +1,20 @@
[Unit]
Description=Refresh Artifact Registry credentials for root podman
Documentation=file:///usr/local/sbin/gitea-bootstrap
After=network-online.target
Wants=network-online.target
# Ordered before auto-update rather than merely wanted by it: a stale token here
# is the single most common reason podman-auto-update silently stops working.
Before=podman-auto-update.service
[Service]
Type=oneshot
RemainAfterExit=no
ExecStart=/usr/local/bin/gitea-ar-auth
# The token is short-lived; a transient metadata-server hiccup should not leave
# the file missing until the next timer tick.
Restart=on-failure
RestartSec=10
[Install]
WantedBy=multi-user.target
+12
View File
@@ -0,0 +1,12 @@
[Unit]
Description=Periodically refresh Artifact Registry credentials
[Timer]
# Metadata-server access tokens last ~1h. Refreshing every 30m keeps a valid
# credential on disk at all times, including whenever auto-update fires.
OnBootSec=1min
OnUnitActiveSec=30min
AccuracySec=1min
[Install]
WantedBy=timers.target
+12
View File
@@ -0,0 +1,12 @@
[Unit]
Description=Gitea dump to Cloud Storage
After=gitea.service
Requires=gitea.service
[Service]
Type=oneshot
ExecStart=/usr/local/bin/gitea-backup
# A dump of a large instance is not fast, and it should never be killed halfway.
TimeoutStartSec=3600
Nice=10
IOSchedulingClass=idle
+10
View File
@@ -0,0 +1,10 @@
[Unit]
Description=Nightly Gitea dump to Cloud Storage
[Timer]
OnCalendar=*-*-* 02:30:00 UTC
RandomizedDelaySec=15min
Persistent=true
[Install]
WantedBy=timers.target
+16
View File
@@ -0,0 +1,16 @@
[Unit]
Description=Re-sync Gitea VM configuration from GCS and apply it
Documentation=file:///usr/local/sbin/gitea-bootstrap
After=network-online.target
Wants=network-online.target
[Service]
Type=oneshot
RemainAfterExit=no
# --sync-only skips package installation and disk setup; it re-pulls the vm/
# tree, re-renders the templates, and restarts only what actually changed.
ExecStart=/usr/local/sbin/gitea-bootstrap --sync-only
TimeoutStartSec=600
[Install]
WantedBy=multi-user.target
+10
View File
@@ -0,0 +1,10 @@
[Unit]
Description=Reboot into a maintenance window if unattended updates require it
Documentation=man:needs-restarting(1)
[Service]
Type=oneshot
# Reboots ONLY when the package layer says a reboot is actually required.
# dnf5-automatic is configured with reboot=never precisely so this unit owns
# the decision and restarts land in a predictable window.
ExecStart=/usr/local/bin/gitea-reboot-if-needed
+12
View File
@@ -0,0 +1,12 @@
[Unit]
Description=Weekly maintenance reboot window
[Timer]
# Sunday 05:00 UTC -- an hour after the weekly image rebuild has pushed :prod,
# so a reboot picks up both OS and container updates in one outage.
OnCalendar=Sun *-*-* 05:00:00 UTC
RandomizedDelaySec=30min
Persistent=false
[Install]
WantedBy=timers.target