Author SHA1 Message Date
JMR-devandClaude Opus 5.5 f672019c91 Add the GitHub to Gitea migration scripts
These moved every non-fork JMR-dev repository from GitHub into this Gitea
instance and made GitHub a push mirror. They are kept for re-runs and
for rotating the mirror PAT, which expires and fails silently.

- migrate.py    full one-time migration (code, issues, PRs, releases,
                wiki, LFS); skips anything already on Gitea
- verify.py     compares branch/tag shas, issue/PR/release counts, and
                the archived and visibility flags; read-only
- repoint.py    re-points local clones' JMR-dev remotes at Gitea,
                following GitHub renames; dry run unless --apply
- pushmirror.py sync-on-commit push mirrors to GitHub, created only when
                every GitHub branch and tag already matches Gitea, so
                the first force-push/prune cannot delete anything

Tokens come from the environment, sourced from Secret Manager. None
appear in these files.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 18:16:18 +07:00
JMR-devandClaude Opus 5.5 f553a87ce6 README: day-to-day make targets need ADC
The Makefile derives PROJECT and ZONE from `pulumi config get`, which
reads the stack from the GCS backend and so needs Application Default
Credentials. Without them the lookup fails silently and every gcloud
command runs with `--project=`. The passphrase is not needed for
plaintext config values.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 04:04:18 -05:00
JMR-devandClaude Opus 5.5 69d586bfcc Let the VM list DNS zones so Caddy can present DNS-01 challenges
With DNS resolution fixed, issuance failed at the challenge:

  presenting for challenge: adding temporary record for zone
  "gitea.jasonmross.dev.": googleapi: Error 403: Forbidden

Testing with the VM service account's own token: managedZones/main and
its rrsets return 200, but managedZones (list) returns 403. The
googleclouddns plugin resolves the domain to a zone by listing the
project's managed zones, and listing is a project-level permission that
the zone-scoped dns.admin binding cannot grant.

Grant roles/dns.reader on the project. It adds read access only, in a
project that holds this single zone; every write stays zone-scoped. A
custom role with just dns.managedZones.list was the alternative, but
managing it would need iam.roleAdmin on the cb-infra Pulumi runner,
which widens a far more powerful identity to narrow a read-only one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 04:04:18 -05:00
JMR-devandClaude Opus 5.5 25dda00eaf runbook: note that Caddy's certificates live on the boot disk
The caddy-data volume is a podman named volume, so it sits under
/var/lib/containers on the boot disk rather than the separately managed
data disk. An instance replacement re-registers the ACME account and
re-issues, and enough of those in a week hits Let's Encrypt's
duplicate-certificate limit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 04:04:18 -05:00
JMR-devandClaude Opus 5.5 0959065d82 nftables: let containers reach aardvark-dns
Caddy could not obtain a certificate on first boot: every request to
acme-v02.api.letsencrypt.org timed out. Containers could reach
1.1.1.1:443 by address but resolved no names at all.

Container DNS goes to aardvark-dns on the bridge gateway (10.89.10.1:53).
That traffic terminates on the host, so it takes the input hook, not
forward. Netavark accepts it in its own table, but gitea_filter's input
chain has policy drop, and a packet must be accepted by every base chain
on the hook. Our drop won.

gitea_filter now accepts tcp/udp 53 from the podman subnet to its
gateway. Both come from instance metadata, so the ruleset is rendered
with envsubst, as the fail2ban jail already is. It is not interface-based
because netavark's bridge name (podman1) is not pinned.

setup_nftables also validates the rendered file before installing it.
Previously it installed first and validated second, so a ruleset that
failed to parse stayed in /etc/sysconfig and would fail nftables.service
on the next boot, leaving the host with no gitea_filter table.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 04:04:18 -05:00
JMR-devandClaude Opus 5.5 3d38240b2f image build: enable BuildKit
The first image build failed in build-gitea:

  the --chmod option requires BuildKit

Both Dockerfiles declare `# syntax=docker/dockerfile:1` and use
`COPY --chmod`, but gcr.io/cloud-builders/docker runs the legacy builder
unless DOCKER_BUILDKIT=1 is set. The image ships the buildx plugin, so
setting it on the two build steps is all that is needed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 04:04:18 -05:00
JMR-devandClaude Opus 5.5 a1669d6be1 Give manual image builds a source bucket cb-image can read
The first `make build` failed before any step ran:

  INVALID_ARGUMENT: could not resolve source: cb-image@... does not
  have storage.objects.get access to ... gitea-496920_cloudbuild/source/...

`gcloud builds submit` uploads the source tarball to <project>_cloudbuild
and the build, running as the user-specified cb-image@, must read it
back. Nothing grants that. Binding on that bucket is not an option: gcloud
creates it on the first submit, after `pulumi up` has already run.
Project-wide objectViewer would also open the backup, config and state
buckets.

Pulumi now owns <project>-gitea-build-source, readable by cb-image@ and
nothing else, with a 7-day delete rule since each tarball is read once.
`make build` stages there via --gcs-source-staging-dir. Triggered builds
fetch source through the GitHub connection and are unaffected.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 04:04:18 -05:00
JMR-devandClaude Opus 5.5 d48f2f5e0c make build: supply SHORT_SHA to manual builds
image.yaml tags each image :$SHORT_SHA as its audit trail and rollback
target. Cloud Build populates SHORT_SHA only for triggered builds; for
`gcloud builds submit` it substitutes an empty string, so the tag becomes
`<image>:` and docker build fails with an invalid reference format. That
is the very first build in the README's setup sequence.

Pass the current commit's short hash explicitly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 04:04:18 -05:00
JMR-devandClaude Opus 5.5 e70a5beac4 Configure the prod stack for gitea-496920
Fills in the values scripts/bootstrap.sh and the existing project
provide: the project id, the Cloud DNS zone resource name (main, holding
gitea.jasonmross.dev), and the cb-infra Pulumi runner account.

encryptionsalt is from `pulumi stack init` with the passphrase already in
Secret Manager (pulumi-config-passphrase), so Cloud Build's infra trigger
can open the stack with the same key.

githubAppInstallationId stays "0" for now: GitHubConfigured() is then
false and the Cloud Build triggers are skipped until the GitHub App is
installed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 04:04:18 -05:00
JMR-devandClaude Opus 5.5 973c1bbcd3 Keep the ACME contact email in Secret Manager
gitea:acmeEmail sat in Pulumi.prod.yaml, which this public repository
publishes, and was then copied into instance metadata. It now lives in a
gitea-acme-email secret instead, read by the VM when it renders the
Caddyfile.

- scripts/bootstrap.sh creates the secret empty and prints how to set
  it, the same as github-pat: the address is chosen, not generated.
- Pulumi grants the VM secretAccessor on it and nothing more. It is kept
  out of secrets.Names, whose members also get secretVersionAdder and are
  mapped to `gitea generate secret` by vm/bootstrap.sh.
- The gitea:acmeEmail config key and the acme-email metadata entry are
  gone.
- vm/bootstrap.sh renders the whole `email` directive. If the secret is
  unreadable it renders a comment instead and warns: Caddy still issues
  certificates under an account with no contact address, whereas an
  empty `email` would fail to parse and leave nothing serving TLS. Same
  directive-or-comment pattern as CADDY_PUBLISH_PORTS and CADDY_SYSCTL.

README setup gains the secret step, plus two that were missing: ADC
login (Pulumi's GCS backend and provider do not use the gcloud login),
and exporting PULUMI_CONFIG_PASSPHRASE from Secret Manager before
`stack init`. Without the latter, init prompts for a new passphrase and
the stack is encrypted with a key Cloud Build's infra trigger never sees.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 04:04:18 -05:00
JMR-devandClaude Opus 5.5 31524e9ddb Build the Pulumi binary before running Pulumi
Pulumi.yaml sets runtime.options.binary to ./gitea-infra, which tells the
Go language host to execute that file instead of compiling the program.
Nothing produced it: `make check` ran `go build ./...`, which discards
output when building multiple packages, and the infra Cloud Build step
went straight to `pulumi up`. Both local preview/up and the first infra
trigger run would fail before planning anything.

`make check` (and therefore preview/up) now builds it with -o, and the
infra pipeline builds it at the top of the pulumi step. `go vet ./...`
still type-checks every package.

Also corrects the Makefile's ZONE fallback from <region>-a to <region>-b.
It applies whenever `pulumi config get` cannot read the stack, and
us-east1 has no -a zone, so make ssh/build would target a zone that does
not exist. -b matches the default in infra/pkg/config.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 04:04:18 -05:00
JMR-devandClaude Opus 5.5 b85df9ee7a Fix Dependabot alerts: upgrade grpc and otel, pin go1.26.9
Eight open Dependabot alerts, all transitive through the Pulumi SDK:

  grpc  < 1.82.2 / <= 1.83.0   (#3, #4 high; #5 medium) -> v1.83.2
  otel/sdk, otlptrace, otlptracegrpc  <= 1.44.0  (#6-8)  -> v1.45.0
  otel/sdk/log, otlplog/otlploggrpc   <  0.21.0  (#9-10) -> v0.21.0

This supersedes Dependabot PR #1, which bumps grpc only and predates the
otel alerts.

otel/log v0.21 changed its API, so the otelslog bridge has to move with
it: v0.18.0 no longer compiles against it. v0.20.1 is the release built
for that otel line.

govulncheck then reported ten reachable standard-library and x/net
vulnerabilities disclosed since the last pin (net/http HTTP/2 and CONNECT
handling, net/textproto, crypto/tls ECH, os on Windows), all fixed in
go1.26.9 and golang.org/x/net v0.60.0. Raising the toolchain floor is the
same remedy as before; x/net v0.60.0 requires go 1.26, which lifts the
go directive from 1.25.11 to 1.26.0.

The Pulumi SDK and pulumi-gcp direct dependencies are deliberately left
where they are, so this does not also change provider behaviour ahead of
the first deploy.

govulncheck now reports no reachable vulnerabilities. GO-2026-5932
(x/crypto/openpgp, no fix available, not called) remains, as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-10 04:04:18 -05:00
5 changed files with 506 additions and 0 deletions
+70
View File
@@ -0,0 +1,70 @@
# GitHub → Gitea migration
These scripts moved every non-fork repository owned by `JMR-dev` from GitHub into this
Gitea instance and made Gitea the primary. GitHub is now a push mirror: Gitea pushes
every commit to it. They are kept here for re-runs and for rotating the mirror token.
All of them are standard-library Python. They read tokens from the environment, never
from arguments or files. `verify.py`, `pushmirror.py` and `repoint.py` also read GitHub
through the local `gh` CLI.
| Script | What it does | Writes to |
|---|---|---|
| `migrate.py` | Full one-time migration: code, issues, PRs, releases, labels, milestones, wiki, LFS. Archives on Gitea whatever is archived on GitHub. | Gitea only |
| `verify.py` | Compares every branch and tag sha, the issue/PR/release counts, and the archived flag and visibility. | nothing |
| `repoint.py` | Re-points local clones' `JMR-dev` GitHub remotes at Gitea, following renames. Dry run unless `--apply`. | local `.git/config` |
| `pushmirror.py` | Creates a sync-on-commit push mirror to GitHub for each active repository. | Gitea, then GitHub through the mirror |
The input is a snapshot of the repository list:
```bash
gh repo list JMR-dev --limit 1000 \
--json name,visibility,isFork,isArchived,diskUsage,description,defaultBranchRef > repos.json
```
## Tokens
Every token lives in Secret Manager in the Gitea project and is passed in by environment:
```bash
export GITEA_TOKEN=$(gcloud secrets versions access latest --secret=gitea-migration-token --project=<project>)
export GITHUB_TOKEN=$(gcloud secrets versions access latest --secret=github-migration-pat --project=<project>) # migrate.py
export MIRROR_PAT=$(gcloud secrets versions access latest --secret=github-mirror-pat --project=<project>) # pushmirror.py
```
- `GITEA_TOKEN` needs `write:repository` and `read:issue`.
- `GITHUB_TOKEN` for migrating should be a **read-only** fine-grained PAT (Contents, Metadata,
Issues, Pull requests). Read-only cannot see draft releases; copy those by hand.
- `MIRROR_PAT` needs Contents and Workflows **read & write**. Without Workflows, GitHub rejects
any push that touches `.github/workflows/`.
## Why they are safe to re-run
- `migrate.py` skips any repository that already exists on Gitea. It never deletes one in
order to retry.
- `pushmirror.py` skips repositories that already have a GitHub push mirror. It also refuses
to create one unless every GitHub branch and tag already matches Gitea. A push mirror
force-pushes and prunes, so this check is what guarantees the first sync cannot overwrite
or delete anything on GitHub.
- Archived repositories never get a push mirror, because GitHub rejects pushes to them.
## Rotating the mirror PAT
Each mirror stores its own copy of the PAT, and an expired PAT fails silently: the only sign
is the error on the repository's *Settings → Mirror* page. To rotate:
1. Store the new token as a new version of `github-mirror-pat`.
2. Delete each repository's GitHub mirror with
`DELETE /api/v1/repos/JMR-dev/<repo>/push_mirrors/<remote_name>`.
3. Re-run `pushmirror.py`.
The refs check runs again in step 3. Anything pushed to GitHub directly in the meantime
shows up as `blocked` rather than being overwritten.
## Behaviour worth knowing
- New commits and branches reach GitHub within seconds.
- A bare branch delete does not trigger a sync. It reaches GitHub at the next push that carries
commits, or at the 8-hour interval.
- Branches created on GitHub, such as Dependabot's, are pruned by the next sync.
- GitHub Actions `on: push` workflows run for mirrored pushes.
+125
View File
@@ -0,0 +1,125 @@
#!/usr/bin/env python3
"""Copy GitHub repositories into Gitea as full, one-time migrations.
Non-destructive by construction:
* GitHub is only ever read (clone + REST API through Gitea's downloader).
* A repository that already exists in Gitea is never deleted or overwritten;
it is skipped, so the script is safe to re-run after a partial failure.
* The only Gitea-side change to an existing repository is setting the archived
flag on a repository that is archived on GitHub.
Usage:
GITEA_TOKEN=... [GITHUB_TOKEN=...] migrate.py repos.json results.jsonl [name ...]
repos.json is `gh repo list OWNER --json name,visibility,isFork,isArchived,
diskUsage,description,defaultBranchRef`. Forks are always skipped. With names,
only those repositories are processed, in the order given; otherwise every
non-fork, smallest first.
"""
import json
import os
import sys
import time
import urllib.error
import urllib.request
GITEA = os.environ.get("GITEA_URL", "https://gitea.jasonmross.dev").rstrip("/")
OWNER = os.environ.get("OWNER", "JMR-dev")
GITEA_TOKEN = os.environ["GITEA_TOKEN"]
GITHUB_TOKEN = os.environ.get("GITHUB_TOKEN", "")
def gitea(method, path, body=None, timeout=60):
req = urllib.request.Request(
f"{GITEA}/api/v1{path}",
method=method,
data=None if body is None else json.dumps(body).encode(),
headers={"Authorization": f"token {GITEA_TOKEN}", "Content-Type": "application/json"},
)
try:
with urllib.request.urlopen(req, timeout=timeout) as r:
raw = r.read()
return r.status, (json.loads(raw) if raw else None)
except urllib.error.HTTPError as e:
raw = e.read()
try:
return e.code, json.loads(raw)
except ValueError:
return e.code, {"message": raw.decode(errors="replace")[:500]}
def migrate_one(repo):
name = repo["name"]
status, existing = gitea("GET", f"/repos/{OWNER}/{name}")
if status == 200:
result = {"result": "exists-skipped"}
if repo["isArchived"] and not existing.get("archived"):
s, _ = gitea("PATCH", f"/repos/{OWNER}/{name}", {"archived": True})
result["archived_set"] = s == 200
return result
if status != 404:
return {"result": "error", "stage": "lookup", "status": status, "detail": existing}
body = {
"clone_addr": f"https://github.com/{OWNER}/{name}.git",
"service": "github",
"repo_owner": OWNER,
"repo_name": name,
"private": repo["visibility"] != "PUBLIC",
"description": (repo.get("description") or "")[:2048],
"mirror": False,
"wiki": True,
"issues": True,
"labels": True,
"milestones": True,
"pull_requests": True,
"releases": True,
"lfs": True,
}
if GITHUB_TOKEN:
body["auth_token"] = GITHUB_TOKEN
started = time.time()
# The API migrates synchronously; a repository with hundreds of PRs takes
# many minutes, mostly in Gitea's rate-limit-aware GitHub downloader.
status, resp = gitea("POST", "/repos/migrate", body, timeout=4 * 3600)
elapsed = round(time.time() - started, 1)
if status != 201:
return {"result": "error", "stage": "migrate", "status": status, "seconds": elapsed,
"detail": (resp or {}).get("message", resp)}
result = {"result": "migrated", "seconds": elapsed, "private": resp.get("private")}
if repo["isArchived"]:
s, _ = gitea("PATCH", f"/repos/{OWNER}/{name}", {"archived": True})
result["archived_set"] = s == 200
return result
def main():
repos_file, results_file, *names = sys.argv[1:]
repos = [r for r in json.load(open(repos_file)) if not r["isFork"]]
if names:
by_name = {r["name"]: r for r in repos}
missing = [n for n in names if n not in by_name]
if missing:
sys.exit(f"not in {repos_file} (or a fork): {', '.join(missing)}")
repos = [by_name[n] for n in names]
else:
repos.sort(key=lambda r: r["diskUsage"])
print(f"{len(repos)} repositories; github token: {'yes' if GITHUB_TOKEN else 'NO'}", flush=True)
with open(results_file, "a") as out:
for i, repo in enumerate(repos, 1):
print(f"[{i}/{len(repos)}] {repo['name']} ({repo['diskUsage']} KB) ...", flush=True)
try:
result = migrate_one(repo)
except Exception as e: # keep going; one bad repository must not stop the batch
result = {"result": "error", "stage": "exception", "detail": repr(e)}
result = {"name": repo["name"], "ts": time.strftime("%H:%M:%S"), **result}
out.write(json.dumps(result) + "\n")
out.flush()
print(f" -> {result['result']} {json.dumps({k: v for k, v in result.items() if k not in ('name', 'result', 'ts')})}", flush=True)
if __name__ == "__main__":
main()
+117
View File
@@ -0,0 +1,117 @@
#!/usr/bin/env python3
"""Create Gitea -> GitHub push mirrors that sync on every commit.
Gitea's push mirror force-pushes and prunes, so it would delete or overwrite
anything that exists only on GitHub. Guard: a mirror is created only when every
GitHub branch and tag already exists on Gitea at the same commit. Then the first
sync cannot remove or rewrite anything on GitHub. Repositories with an existing
GitHub push mirror are skipped, so re-runs are safe.
Usage: GITEA_TOKEN=... MIRROR_PAT=... pushmirror.py repos.json results.jsonl [name ...]
Archived (on GitHub) repositories and forks are always skipped: GitHub rejects
pushes to archived repositories.
"""
import json
import os
import subprocess
import sys
import time
import urllib.error
import urllib.request
GITEA = "https://gitea.jasonmross.dev"
OWNER = "JMR-dev"
GITEA_TOKEN = os.environ["GITEA_TOKEN"]
MIRROR_PAT = os.environ["MIRROR_PAT"]
INTERVAL = os.environ.get("MIRROR_INTERVAL", "8h") # backstop; sync_on_commit does the real work
def gitea(method, path, body=None):
req = urllib.request.Request(f"{GITEA}/api/v1{path}", method=method,
data=None if body is None else json.dumps(body).encode(),
headers={"Authorization": f"token {GITEA_TOKEN}", "Content-Type": "application/json"})
try:
with urllib.request.urlopen(req, timeout=120) as r:
raw = r.read()
return r.status, (json.loads(raw) if raw else None)
except urllib.error.HTTPError as e:
raw = e.read()
try:
return e.code, json.loads(raw)
except ValueError:
return e.code, {"message": raw.decode(errors="replace")[:300]}
def gitea_all(path):
items, page = [], 1
while True:
_, batch = gitea("GET", f"{path}?limit=50&page={page}")
batch = batch or []
items += batch
if len(batch) < 50:
return items
page += 1
def gh_refs(name, kind):
out = subprocess.run(["gh", "api", "--paginate", f"repos/{OWNER}/{name}/{kind}"],
check=True, capture_output=True, text=True).stdout
items = json.loads(out.replace("]\n[", ",").replace("][", ",")) if out.strip() else []
return {i["name"]: i["commit"]["sha"] for i in items}
def refs_problems(name):
problems = []
for kind, gitea_key in (("branches", "id"), ("tags", "sha")):
gh = gh_refs(name, kind)
gt = {i["name"]: i["commit"][gitea_key] for i in gitea_all(f"/repos/{OWNER}/{name}/{kind}")}
for ref, sha in gh.items():
if gt.get(ref) != sha:
problems.append(f"{kind[:-1] if kind != 'branches' else 'branch'} {ref}: github={sha[:10]} gitea={(gt.get(ref) or 'missing')[:10]}")
return problems
def mirror_one(name):
status, mirrors = gitea("GET", f"/repos/{OWNER}/{name}/push_mirrors")
if status != 200:
return {"result": "error", "stage": "list", "status": status, "detail": mirrors}
if any("github.com" in (m.get("remote_address") or "") for m in mirrors or []):
return {"result": "exists-skipped"}
problems = refs_problems(name)
if problems:
return {"result": "blocked", "problems": problems}
status, resp = gitea("POST", f"/repos/{OWNER}/{name}/push_mirrors", {
"remote_address": f"https://github.com/{OWNER}/{name}.git",
"remote_username": OWNER,
"remote_password": MIRROR_PAT,
"interval": INTERVAL,
"sync_on_commit": True,
})
if status not in (200, 201):
return {"result": "error", "stage": "create", "status": status, "detail": (resp or {}).get("message", resp)}
gitea("POST", f"/repos/{OWNER}/{name}/push_mirrors-sync")
return {"result": "created", "remote_name": resp.get("remote_name")}
def main():
repos_file, results_file, *names = sys.argv[1:]
repos = [r for r in json.load(open(repos_file)) if not r["isFork"] and not r["isArchived"]]
if names:
repos = [r for r in repos if r["name"] in set(names)]
print(f"{len(repos)} repositories", flush=True)
with open(results_file, "a") as out:
for i, r in enumerate(repos, 1):
try:
res = mirror_one(r["name"])
except Exception as e:
res = {"result": "error", "stage": "exception", "detail": repr(e)}
res = {"name": r["name"], "ts": time.strftime("%H:%M:%S"), **res}
out.write(json.dumps(res) + "\n")
out.flush()
print(f"[{i}/{len(repos)}] {r['name']}: {res['result']} {json.dumps({k: v for k, v in res.items() if k not in ('name', 'ts', 'result')})}", flush=True)
if __name__ == "__main__":
main()
+85
View File
@@ -0,0 +1,85 @@
#!/usr/bin/env python3
"""Re-point local clones' remotes from GitHub (JMR-dev) to Gitea.
Only remotes whose URL is a JMR-dev GitHub repository that exists on Gitea are
changed. Third-party remotes (upstreams, other owners) and JMR-dev repos that
were not migrated (forks) are left untouched and reported. GitHub renames are
followed via the API redirect. Every change is appended to a TSV so it can be
reverted with `git remote set-url <remote> <old-url>`.
Usage: GITEA_TOKEN=... repoint.py <dirs-file> <workspace-root> <changes.tsv> [--apply]
Without --apply it only prints the plan.
"""
import json
import os
import re
import subprocess
import sys
import urllib.error
import urllib.request
OWNER = "JMR-dev"
GITEA = "https://gitea.jasonmross.dev"
GITEA_SSH = "ssh://git@gitea.jasonmross.dev:2222"
GH_URL = re.compile(r"^(?:git@github\.com:|ssh://git@github\.com/|https://github\.com/)([^/]+)/(.+?)(?:\.git)?/?$")
def git(cwd, *args):
return subprocess.run(["git", *args], cwd=cwd, capture_output=True, text=True)
def gitea_has(name):
req = urllib.request.Request(f"{GITEA}/api/v1/repos/{OWNER}/{name}",
headers={"Authorization": f"token {os.environ['GITEA_TOKEN']}"})
try:
with urllib.request.urlopen(req, timeout=30) as r:
return json.load(r)["name"]
except urllib.error.HTTPError as e:
if e.code == 404:
return None
raise
def canonical(name):
"""Follow a GitHub rename: the API redirects old names to the new repo."""
r = subprocess.run(["gh", "api", f"repos/{OWNER}/{name}", "--jq", ".name"], capture_output=True, text=True)
return r.stdout.strip() or name
def main():
dirs_file, root, changes_file, *flags = sys.argv[1:]
apply = "--apply" in flags
out = open(changes_file, "a") if apply else None
for rel in open(dirs_file).read().split():
d = os.path.normpath(os.path.join(root, rel))
remotes = git(d, "remote").stdout.split()
for remote in remotes:
url = git(d, "remote", "get-url", remote).stdout.strip()
m = GH_URL.match(url)
if not m:
print(f"SKIP {rel:38} {remote:9} not a GitHub URL: {url}")
continue
owner, name = m.groups()
if owner != OWNER:
print(f"SKIP {rel:38} {remote:9} third-party ({owner}/{name})")
continue
target = gitea_has(canonical(name))
if not target:
print(f"SKIP {rel:38} {remote:9} {owner}/{name} is not on Gitea (fork, not migrated)")
continue
new = f"{GITEA_SSH}/{OWNER}/{target}.git"
if url == new:
print(f"OK {rel:38} {remote:9} already {new}")
continue
print(f"{'CHANGE' if apply else 'PLAN '} {rel:38} {remote:9} {url} -> {new}")
if apply:
r = git(d, "remote", "set-url", remote, new)
if r.returncode:
print(f" !! set-url failed: {r.stderr.strip()}")
continue
out.write(f"{d}\t{remote}\t{url}\t{new}\n")
out.flush()
if __name__ == "__main__":
main()
+109
View File
@@ -0,0 +1,109 @@
#!/usr/bin/env python3
"""Compare migrated repositories between GitHub and Gitea. Read-only on both.
Checks, per repository: every branch and tag points at the same commit; issue,
pull request and release counts match; archived flag and visibility match.
Usage: GITEA_TOKEN=... verify.py repos.json [name ...]
GitHub is read through the local `gh` CLI.
"""
import json
import os
import subprocess
import sys
import urllib.error
import urllib.request
GITEA = os.environ.get("GITEA_URL", "https://gitea.jasonmross.dev").rstrip("/")
OWNER = os.environ.get("OWNER", "JMR-dev")
GITEA_TOKEN = os.environ["GITEA_TOKEN"]
def gh_api(path):
out = subprocess.run(["gh", "api", "--paginate", path], check=True, capture_output=True, text=True).stdout
# --paginate concatenates JSON arrays as `][`; join them into one.
return json.loads(out.replace("]\n[", ",").replace("][", ",")) if out.strip() else []
def gh_counts(name):
q = ('query($o:String!,$n:String!){repository(owner:$o,name:$n){'
'issues{totalCount} pullRequests{totalCount} releases{totalCount}}}')
out = subprocess.run(["gh", "api", "graphql", "-f", f"query={q}", "-f", f"o={OWNER}", "-f", f"n={name}"],
check=True, capture_output=True, text=True).stdout
r = json.loads(out)["data"]["repository"]
return {"issues": r["issues"]["totalCount"], "pulls": r["pullRequests"]["totalCount"],
"releases": r["releases"]["totalCount"]}
def gitea_get(path):
req = urllib.request.Request(f"{GITEA}/api/v1{path}", headers={"Authorization": f"token {GITEA_TOKEN}"})
with urllib.request.urlopen(req, timeout=60) as r:
return json.loads(r.read()), r.headers.get("X-Total-Count")
def gitea_all(path):
items, page = [], 1
sep = "&" if "?" in path else "?"
while True:
batch, _ = gitea_get(f"{path}{sep}limit=50&page={page}")
batch = batch or [] # an empty repository returns null, not []
items += batch
if len(batch) < 50:
return items
page += 1
def gitea_count(path):
_, total = gitea_get(path + ("&" if "?" in path else "?") + "limit=1")
return int(total or 0)
def verify(repo):
name = repo["name"]
problems = []
try:
g, _ = gitea_get(f"/repos/{OWNER}/{name}")
except urllib.error.HTTPError as e:
return {"name": name, "ok": False, "problems": [f"not in gitea ({e.code})"]}
if g["archived"] != repo["isArchived"]:
problems.append(f"archived gitea={g['archived']} github={repo['isArchived']}")
if g["private"] != (repo["visibility"] != "PUBLIC"):
problems.append(f"private gitea={g['private']} github={repo['visibility']}")
gh_branches = {b["name"]: b["commit"]["sha"] for b in gh_api(f"repos/{OWNER}/{name}/branches")}
gt_branches = {b["name"]: b["commit"]["id"] for b in gitea_all(f"/repos/{OWNER}/{name}/branches")}
gh_tags = {t["name"]: t["commit"]["sha"] for t in gh_api(f"repos/{OWNER}/{name}/tags")}
gt_tags = {t["name"]: t["commit"]["sha"] for t in gitea_all(f"/repos/{OWNER}/{name}/tags")}
for kind, a, b in (("branch", gh_branches, gt_branches), ("tag", gh_tags, gt_tags)):
for ref in sorted(set(a) | set(b)):
if a.get(ref) != b.get(ref):
problems.append(f"{kind} {ref}: github={(a.get(ref) or 'missing')[:10]} gitea={(b.get(ref) or 'missing')[:10]}")
gh = gh_counts(name)
gt = {"issues": gitea_count(f"/repos/{OWNER}/{name}/issues?state=all&type=issues"),
"pulls": gitea_count(f"/repos/{OWNER}/{name}/issues?state=all&type=pulls"),
"releases": gitea_count(f"/repos/{OWNER}/{name}/releases")}
for k in gh:
if gh[k] != gt[k]:
problems.append(f"{k}: github={gh[k]} gitea={gt[k]}")
return {"name": name, "ok": not problems, "branches": len(gt_branches), "tags": len(gt_tags),
**{f"gitea_{k}": v for k, v in gt.items()}, "problems": problems}
def main():
repos_file, *names = sys.argv[1:]
repos = [r for r in json.load(open(repos_file)) if not r["isFork"]]
if names:
repos = [r for r in repos if r["name"] in set(names)]
for repo in repos:
try:
res = verify(repo)
except Exception as e:
res = {"name": repo["name"], "ok": False, "problems": [f"verify error: {e!r}"]}
print(json.dumps(res), flush=True)
if __name__ == "__main__":
main()