homelab/docs/authentik-migration.md
Russell Seymour 1250c9cef6
Some checks are pending
deploy / deploy (push) Waiting to run
Initial checkin of code for managing homelab
2026-08-24 20:52:36 +01:00

15 KiB

Migrating Authentik from Unraid to k3s

A one-time, hand-run cutover: move the Authentik database off Unraid onto the cluster's CloudNativePG instance, bring Authentik up on k3s from src/authentik/, verify it, then change DNS.

This is deliberately not automated. It runs once, most of its steps are irreversible in the wrong order, and the verification between them is a judgement call rather than a task result — the sort of thing a playbook makes harder rather than easier. The repo half is automated: everything from "deploy Authentik" onwards is playbooks/k3s.yml, and re-running it later is an ordinary converge.

Read the whole thing before starting. The dump is taken with Authentik stopped, so there's a service outage from step 5 to step 11 — budget an hour and do it when nobody needs to log in.


What moves, and what doesn't

Where it lives now How it gets across
Users, groups, applications, providers, flows, tokens, certificates Postgres pg_dump | psql (step 7)
Uploaded icons and flow backgrounds /media on disk scp + kubectl cp (step 9)
Session cookies signed by AUTHENTIK_SECRET_KEY carry the key over (step 2)
Cache / task queue in-process nothing — recent Authentik has no Redis, and the chart has no Redis dependency

Authentik stores its certificates (including the signing keys OIDC/SAML providers use) in the database, not in a certs/ directory, so a successful database restore carries them. That is what lets integrations like Shelfarr's OIDC (src/arr/) keep working without being re-registered.

The one piece with no automatic path is anything you mounted into the Unraid container by hand — custom templates, a certs/ directory you populated yourself. Step 1 is where you find out whether you have any.


1. Gather the facts

Everything below needs these. Run on the Unraid box (web terminal, or SSH).

# Which containers make up the current deployment, and what image/tag.
docker ps -a --format '{{.Names}}\t{{.Image}}\t{{.Status}}' | grep -i authentik

Write down the image tag — that is the Authentik version, and step 3 depends on it.

# Every mount, so you can see where /media lives and spot anything unexpected.
docker inspect authentik-server \
  --format '{{range .Mounts}}{{.Source}} -> {{.Destination}}{{println}}{{end}}'

# The database connection it currently uses, and the secret key.
docker inspect authentik-server \
  --format '{{range .Config.Env}}{{println .}}{{end}}' | grep -E 'POSTGRES|SECRET_KEY'

That last command tells you the thing this document can't assume: which Postgres the Unraid Authentik talks to. Two possibilities, and they change step 7:

  • AUTHENTIK_POSTGRESQL__HOST=shared-postgres (or 192.168.50.1) — it uses the shared instance this repo already manages, src/shared/postgres/. The superuser password is in Vault at homelab/shared/postgres.

  • anything else (authentik-postgresql, postgres, a container name from Authentik's own compose bundle) — the Unraid UI deployment brought its own Postgres container. You'll need its credentials, which are in that container's own environment:

    docker inspect <that-container> \
      --format '{{range .Config.Env}}{{println .}}{{end}}' | grep POSTGRES
    

    and it is probably not published on a host port, which step 7 handles.

Also note AUTHENTIK_SECRET_KEY from the output above — step 2 needs it verbatim.


2. Populate Vault

vault kv put kv/homelab/authentik \
  AUTHENTIK_SECRET_KEY='<the value from step 1>' \
  AUTHENTIK_POSTGRESQL__PASSWORD='<a new strong password>'

Two different rules for those two values, and it's worth being clear about why:

  • AUTHENTIK_SECRET_KEY must be the existing one. It signs session cookies. A fresh value doesn't break anything permanently, but it invalidates every active session at the moment DNS flips — so instead of a silent cutover, everyone gets logged out and the migration announces itself. Copy it across.
  • AUTHENTIK_POSTGRESQL__PASSWORD can be new. You are creating a new role on a different Postgres; nothing in the dump references it. Generate a fresh one.

If you're also using cert-manager's DNS-01 solver (step 4), populate its credentials now too:

vault kv put kv/homelab/k3s-cert-manager \
  CLOUDFLARE_API_TOKEN='<token with Zone:DNS:Edit on the zone>'

3. Confirm the repo's values

Three values are guesses until you check them. All are flagged in-file.

src/authentik/common/vars.yml

  • AUTHENTIK_VERSION — set this to the image tag from step 1, not to whatever is newest. Authentik runs Django migrations on startup against whatever schema it finds, and they are one-way: a newer Authentik pointed at an older dump upgrades the schema silently, and if the cutover then has to be rolled back, the Unraid instance can no longer read its own database. Migrate first, upgrade second, in a separate commit.
  • AUTHENTIK_HOST / AUTHENTIK_URL — the hostname you'll be moving. Must match what's there today, or every OIDC/SAML redirect URI registered in Authentik breaks.

build/config/ansible/inventory/group_vars/k3s_cluster.yml

  • k3s_cert_manager_solver — empty by default, and roles/k3s_cert_manager refuses to run until it's set. The commented Cloudflare example in that file is the common case; for anything else take the stanza from cert-manager's DNS-01 docs.

While you're testing, point k3s_cert_manager_acme_server at Let's Encrypt staging. Production allows five failed validations per hostname per hour, and exhausting it means waiting rather than retrying.


4. Deploy cert-manager and confirm the issuer works

cd build/config/ansible
ansible-playbook playbooks/k3s.yml --tags cert-manager

Then check it actually registered — a ClusterIssuer is created whether or not its solver is valid, so "the resource exists" proves nothing:

$env:KUBECONFIG = "$PWD\local\k3s\homelab-utils.kubeconfig"
kubectl get clusterissuer letsencrypt -o jsonpath='{.status.conditions[*].message}'

You want The ACME account was registered with the ACME server. Anything else, fix it here — not after Authentik is deployed and waiting on a certificate.

Worth knowing: with DNS-01, the certificate can be issued before DNS points at the cluster. The challenge is a TXT record, not a request to your web server, so cert-manager never needs to be reachable. That's why TLS is sorted out in this step and not after the flip.


5. Freeze the Unraid instance

Stop writes before dumping. Skipping this gives you a dump that's missing whatever happened during it — usually invisible, occasionally a user account that no longer exists.

docker stop authentik-server authentik-worker

Leave its Postgres running — step 7 reads from it. The outage starts here.


6. Create the role and database on CNPG

The restore needs the authentik role to exist first so the dump's objects land with the right ownership.

$env:KUBECONFIG = "$PWD\local\k3s\homelab-utils.kubeconfig"
kubectl -n shared-postgres exec -i shared-postgres-1 -- psql -v ON_ERROR_STOP=1 <<'SQL'
CREATE ROLE authentik LOGIN PASSWORD 'the-password-from-step-2';
CREATE DATABASE authentik OWNER authentik;
SQL

kubectl exec into a CNPG pod lands you as the postgres superuser over the local socket, so bare psql needs no credentials.

You could instead let roles/k3s_app create these — its tasks are idempotent and will run in step 8 regardless — but doing it by hand here keeps the database creation and the restore adjacent, and avoids deploying Authentik against an empty database first (which would have it run initial migrations and build a schema the restore then has to fight).


7. Dump and restore

Run the whole thing inside the CNPG pod, piping pg_dump straight into psql. Three reasons this is better than dumping to a file on Windows:

  • pg_dump must be at least the version of the server being dumped, and the CNPG pod ships Postgres 18 client tools — dumping the Unraid Postgres 13 from there is the supported direction. Your workstation probably has no pg_dump at all.
  • PowerShell re-encodes bytes passing through > and |, which silently corrupts a dump file. Nothing on the Windows filesystem means nothing to corrupt.
  • No intermediate file holding the entire user database in cleartext.

If Authentik used the shared Postgres (published on 192.168.50.1:5432)

$env:KUBECONFIG = "$PWD\local\k3s\homelab-utils.kubeconfig"
$srcPw = Read-Host 'Unraid postgres superuser password'

kubectl -n shared-postgres exec -i shared-postgres-1 -- `
  env "PGPASSWORD=$srcPw" bash -c @'
set -euo pipefail
{ echo "SET ROLE authentik;"
  pg_dump -h 192.168.50.1 -p 5432 -U postgres -d authentik \
          --no-owner --no-privileges
} | psql -v ON_ERROR_STOP=1 -d authentik
'@

SET ROLE authentik before the dump stream, combined with --no-owner, is what gets ownership right: the session creates every object as authentik, so the application role owns its own schema. Without it, everything lands owned by postgres and Authentik can read but not migrate.

ON_ERROR_STOP=1 matters — without it psql reports success having skipped every statement that failed.

If it used its own Postgres container (not published)

Publish it temporarily, then use the command above with the container's own credentials and port:

# on Unraid
docker run -d --rm --name pgbridge -p 55432:5432 \
  --network container:<authentik-postgres-container> alpine/socat \
  TCP-LISTEN:5432,fork TCP:127.0.0.1:5432

then -h 192.168.50.1 -p 55432 -U <its user> in the pg_dump above, and docker stop pgbridge afterwards. Simpler alternative if you'd rather not: docker exec <container> pg_dump ... > /mnt/user/appdata/authentik.sql on Unraid, scp it to your workstation, and kubectl exec -i ... -- psql -d authentik < authentik.sql — but a plain-SQL file through PowerShell redirection is exactly the encoding hazard noted above, so run that last command from WSL.

Check the restore

kubectl -n shared-postgres exec -i shared-postgres-1 -- `
  psql -d authentik -c '\dt' -c 'SELECT count(*) FROM authentik_core_user;'

You should see a long table list and a user count matching what the old instance had. An empty table list means the restore did nothing.


8. Deploy Authentik on k3s

cd build/config/ansible
ansible-playbook playbooks/k3s.yml --tags apps

A green run means the manifests landed and the database was provisioned — not that the workload came up. Watch it actually start:

kubectl -n authentik get pods -w

The first start takes several minutes on a Pi: pulling the image, then the worker checking migrations. Because the deployed version matches the dumped version (step 3), it should find nothing to migrate.

If pods stay Pending, check the PVC — local-path uses WaitForFirstConsumer, so a Pending PVC alongside a Pending pod is normal only until the scheduler picks a node:

kubectl -n authentik describe pvc authentik-media
kubectl -n authentik logs -l app.kubernetes.io/component=worker --tail=50

9. Copy /media across

Two hops rather than a pipe, again to keep binary data out of a PowerShell pipeline:

# from the media path you noted in step 1
scp -r root@192.168.50.1:/mnt/user/appdata/authentik/media ./local/authentik-media

$pod = kubectl -n authentik get pod -l app.kubernetes.io/component=server `
  -o jsonpath='{.items[0].metadata.name}'
kubectl -n authentik cp ./local/authentik-media/. "${pod}:/media"

local/ is already gitignored, so the staging copy won't be committed.

Skip this entirely if you never uploaded custom icons or backgrounds — Authentik falls back to its built-in assets, and the directory is empty on a stock install.


10. Verify before touching DNS

The hostname still resolves to Unraid, so test by overriding resolution rather than changing it. --resolve sends the request to Traefik with the right SNI and Host header:

curl.exe -sv --resolve auth.turtlesystems.uk:443:192.168.50.80 `
  https://auth.turtlesystems.uk/-/health/ready/

Check, in order:

  1. The TLS certificate is the Let's Encrypt one, not Traefik's self-signed default (kubectl -n authentik get certificate → READY True). If you used the staging server in step 3, it will be untrusted — expected; switch to production and re-run --tags cert-manager once satisfied.
  2. /-/health/ready/ returns 204.
  3. Add a hosts-file entry pointing the name at 192.168.50.80 and log in through a browser with a real user account. This is the check that matters — it exercises the restored password hashes, the flows, and the session cookie signed with the carried-over secret key.
  4. Open an existing application from the Authentik dashboard and confirm the OIDC/SAML round-trip still works (Shelfarr is the one this repo knows about — see src/arr/).

Remove the hosts-file entry afterwards.


11. Flip DNS

Point auth.turtlesystems.uk at 192.168.50.80 (Traefik's MetalLB address).

Lower the record's TTL a few hours beforehand if you can — it shortens the window in which a rollback is still invisible to clients.


Rollback

Viable until users start making changes in the new instance; after that, rolling back loses whatever they did.

  1. Point DNS back at Unraid.
  2. docker start authentik-server authentik-worker.

The old database was only ever read, never written, so it's exactly as it was at step 5. Nothing to restore.

Do not run both instances at once against their separate databases — they diverge immediately, and there is no way to merge them afterwards. That is the same "never declare an app on two platforms" rule from CLAUDE.md, and it is why src/authentik/ansible/unraid/ is a README rather than a compose file.


After it has settled

Once you're confident, in this order:

  1. Delete the Unraid containers. A UI action, not state: absent — this repo never deployed them, so it has nothing to tear down. Keep the appdata directory a while longer.
  2. Drop the old database, if it was on the shared Postgres:
    docker exec -it shared-postgres psql -U postgres -c 'DROP DATABASE authentik;'
    
    Take a copy first if you want one; nothing else does.
  3. Upgrade Authentik, now as an ordinary change: bump AUTHENTIK_VERSION in src/authentik/common/vars.yml, read the release notes for anything between your version and the target, and ansible-playbook playbooks/k3s.yml --tags apps -e only_apps=authentik. Do this as its own commit, so a bad upgrade is one git revert away from the version you know worked.

Backups, afterwards

Worth stating plainly: the cluster's CNPG has no backups configured (CLAUDE.md → "Key decisions" — HA replication only). Before this migration, Authentik's database sat on Unraid, inside whatever covers /mnt/user/appdata. After it, the only redundancy is a streaming replica on another Pi, which protects against a dead SD card and not against a bad migration, a dropped table, or a mistake in this document.

The gap this leaves is real and this cutover widens it. CNPG's answer is a ScheduledBackup to object storage or an NFS volume on the NAS; that's not built here yet, and it's the obvious next piece of work after Authentik is settled.