First deploy: fix Pulumi, image build, container DNS and DNS-01; Dependabot updates #2

Merged
JMR-dev merged 11 commits from deploy-prep into main 2026-10-10 09:04:19 +00:00
11 Commits
Author SHA1 Message Date
JMR-devandClaude Opus 5.5 f0de95ee4e README: day-to-day make targets need ADC
The Makefile derives PROJECT and ZONE from `pulumi config get`, which
reads the stack from the GCS backend and so needs Application Default
Credentials. Without them the lookup fails silently and every gcloud
command runs with `--project=`. The passphrase is not needed for
plaintext config values.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 15:42:48 +07:00
JMR-devandClaude Opus 5.5 f362d26ed7 Let the VM list DNS zones so Caddy can present DNS-01 challenges
With DNS resolution fixed, issuance failed at the challenge:

  presenting for challenge: adding temporary record for zone
  "gitea.jasonmross.dev.": googleapi: Error 403: Forbidden

Testing with the VM service account's own token: managedZones/main and
its rrsets return 200, but managedZones (list) returns 403. The
googleclouddns plugin resolves the domain to a zone by listing the
project's managed zones, and listing is a project-level permission that
the zone-scoped dns.admin binding cannot grant.

Grant roles/dns.reader on the project. It adds read access only, in a
project that holds this single zone; every write stays zone-scoped. A
custom role with just dns.managedZones.list was the alternative, but
managing it would need iam.roleAdmin on the cb-infra Pulumi runner,
which widens a far more powerful identity to narrow a read-only one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 15:37:12 +07:00
JMR-devandClaude Opus 5.5 be965b58cf runbook: note that Caddy's certificates live on the boot disk
The caddy-data volume is a podman named volume, so it sits under
/var/lib/containers on the boot disk rather than the separately managed
data disk. An instance replacement re-registers the ACME account and
re-issues, and enough of those in a week hits Let's Encrypt's
duplicate-certificate limit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 15:21:56 +07:00
JMR-devandClaude Opus 5.5 b4fad783e4 nftables: let containers reach aardvark-dns
Caddy could not obtain a certificate on first boot: every request to
acme-v02.api.letsencrypt.org timed out. Containers could reach
1.1.1.1:443 by address but resolved no names at all.

Container DNS goes to aardvark-dns on the bridge gateway (10.89.10.1:53).
That traffic terminates on the host, so it takes the input hook, not
forward. Netavark accepts it in its own table, but gitea_filter's input
chain has policy drop, and a packet must be accepted by every base chain
on the hook. Our drop won.

gitea_filter now accepts tcp/udp 53 from the podman subnet to its
gateway. Both come from instance metadata, so the ruleset is rendered
with envsubst, as the fail2ban jail already is. It is not interface-based
because netavark's bridge name (podman1) is not pinned.

setup_nftables also validates the rendered file before installing it.
Previously it installed first and validated second, so a ruleset that
failed to parse stayed in /etc/sysconfig and would fail nftables.service
on the next boot, leaving the host with no gitea_filter table.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 15:09:49 +07:00
JMR-devandClaude Opus 5.5 9adfe9fc84 image build: enable BuildKit
The first image build failed in build-gitea:

  the --chmod option requires BuildKit

Both Dockerfiles declare `# syntax=docker/dockerfile:1` and use
`COPY --chmod`, but gcr.io/cloud-builders/docker runs the legacy builder
unless DOCKER_BUILDKIT=1 is set. The image ships the buildx plugin, so
setting it on the two build steps is all that is needed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 15:01:26 +07:00
JMR-devandClaude Opus 5.5 529dc87373 Give manual image builds a source bucket cb-image can read
The first `make build` failed before any step ran:

  INVALID_ARGUMENT: could not resolve source: cb-image@... does not
  have storage.objects.get access to ... gitea-496920_cloudbuild/source/...

`gcloud builds submit` uploads the source tarball to <project>_cloudbuild
and the build, running as the user-specified cb-image@, must read it
back. Nothing grants that. Binding on that bucket is not an option: gcloud
creates it on the first submit, after `pulumi up` has already run.
Project-wide objectViewer would also open the backup, config and state
buckets.

Pulumi now owns <project>-gitea-build-source, readable by cb-image@ and
nothing else, with a 7-day delete rule since each tarball is read once.
`make build` stages there via --gcs-source-staging-dir. Triggered builds
fetch source through the GitHub connection and are unaffected.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 14:50:58 +07:00
JMR-devandClaude Opus 5.5 da62266766 make build: supply SHORT_SHA to manual builds
image.yaml tags each image :$SHORT_SHA as its audit trail and rollback
target. Cloud Build populates SHORT_SHA only for triggered builds; for
`gcloud builds submit` it substitutes an empty string, so the tag becomes
`<image>:` and docker build fails with an invalid reference format. That
is the very first build in the README's setup sequence.

Pass the current commit's short hash explicitly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 14:49:25 +07:00
JMR-devandClaude Opus 5.5 0acf3b7863 Configure the prod stack for gitea-496920
Fills in the values scripts/bootstrap.sh and the existing project
provide: the project id, the Cloud DNS zone resource name (main, holding
gitea.jasonmross.dev), and the cb-infra Pulumi runner account.

encryptionsalt is from `pulumi stack init` with the passphrase already in
Secret Manager (pulumi-config-passphrase), so Cloud Build's infra trigger
can open the stack with the same key.

githubAppInstallationId stays "0" for now: GitHubConfigured() is then
false and the Cloud Build triggers are skipped until the GitHub App is
installed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 14:41:58 +07:00
JMR-devandClaude Opus 5.5 f0e7a3b66f Keep the ACME contact email in Secret Manager
gitea:acmeEmail sat in Pulumi.prod.yaml, which this public repository
publishes, and was then copied into instance metadata. It now lives in a
gitea-acme-email secret instead, read by the VM when it renders the
Caddyfile.

- scripts/bootstrap.sh creates the secret empty and prints how to set
  it, the same as github-pat: the address is chosen, not generated.
- Pulumi grants the VM secretAccessor on it and nothing more. It is kept
  out of secrets.Names, whose members also get secretVersionAdder and are
  mapped to `gitea generate secret` by vm/bootstrap.sh.
- The gitea:acmeEmail config key and the acme-email metadata entry are
  gone.
- vm/bootstrap.sh renders the whole `email` directive. If the secret is
  unreadable it renders a comment instead and warns: Caddy still issues
  certificates under an account with no contact address, whereas an
  empty `email` would fail to parse and leave nothing serving TLS. Same
  directive-or-comment pattern as CADDY_PUBLISH_PORTS and CADDY_SYSCTL.

README setup gains the secret step, plus two that were missing: ADC
login (Pulumi's GCS backend and provider do not use the gcloud login),
and exporting PULUMI_CONFIG_PASSPHRASE from Secret Manager before
`stack init`. Without the latter, init prompts for a new passphrase and
the stack is encrypted with a key Cloud Build's infra trigger never sees.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 14:39:44 +07:00
JMR-devandClaude Opus 5.5 367b188eae Build the Pulumi binary before running Pulumi
Pulumi.yaml sets runtime.options.binary to ./gitea-infra, which tells the
Go language host to execute that file instead of compiling the program.
Nothing produced it: `make check` ran `go build ./...`, which discards
output when building multiple packages, and the infra Cloud Build step
went straight to `pulumi up`. Both local preview/up and the first infra
trigger run would fail before planning anything.

`make check` (and therefore preview/up) now builds it with -o, and the
infra pipeline builds it at the top of the pulumi step. `go vet ./...`
still type-checks every package.

Also corrects the Makefile's ZONE fallback from <region>-a to <region>-b.
It applies whenever `pulumi config get` cannot read the stack, and
us-east1 has no -a zone, so make ssh/build would target a zone that does
not exist. -b matches the default in infra/pkg/config.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 14:29:46 +07:00
JMR-devandClaude Opus 5.5 d05908133c Fix Dependabot alerts: upgrade grpc and otel, pin go1.26.9
Eight open Dependabot alerts, all transitive through the Pulumi SDK:

  grpc  < 1.82.2 / <= 1.83.0   (#3, #4 high; #5 medium) -> v1.83.2
  otel/sdk, otlptrace, otlptracegrpc  <= 1.44.0  (#6-8)  -> v1.45.0
  otel/sdk/log, otlplog/otlploggrpc   <  0.21.0  (#9-10) -> v0.21.0

This supersedes Dependabot PR #1, which bumps grpc only and predates the
otel alerts.

otel/log v0.21 changed its API, so the otelslog bridge has to move with
it: v0.18.0 no longer compiles against it. v0.20.1 is the release built
for that otel line.

govulncheck then reported ten reachable standard-library and x/net
vulnerabilities disclosed since the last pin (net/http HTTP/2 and CONNECT
handling, net/textproto, crypto/tls ECH, os on Windows), all fixed in
go1.26.9 and golang.org/x/net v0.60.0. Raising the toolchain floor is the
same remedy as before; x/net v0.60.0 requires go 1.26, which lifts the
go directive from 1.25.11 to 1.26.0.

The Pulumi SDK and pulumi-gcp direct dependencies are deliberately left
where they are, so this does not also change provider behaviour ahead of
the first deploy.

govulncheck now reports no reachable vulnerabilities. GO-2026-5932
(x/crypto/openpgp, no fix available, not called) remains, as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 14:29:06 +07:00