Building a GitOps VM Provisioning Platform on Proxmox — Part 2: Vault, Secrets & the JWT Trust Chain

Building a GitOps VM Provisioning Platform on Proxmox — Part 2: Vault, Secrets & the JWT Trust Chain

In Part 1 I made a claim I now have to back up: no credential in this pipeline is stored in GitLab. Not the Proxmox API token, not the SSH keys — nothing.

That's not quite the whole truth. Exactly one value lives in GitLab's CI variables, and it's the address of the Vault server. That's it. A URL. Everything else — the actual secrets — lives in Vault and is fetched at runtime by CI jobs that prove their identity cryptographically without ever holding a long-lived token.

This is the post where that machinery gets built. It's the conceptual heart of the whole project, so I'm going to take my time on the part that matters most: how Vault knows a request from GitLab CI is genuinely from my project and authorized to read secrets. If you've ever set up Vault JWT auth by copy-pasting commands and not fully understood what bound_claims was doing, this is for you.

One assumption before we start: Vault is already installed and running. As I noted in Part 1, vault-01 is the one VM I stood up by hand — a small VM cloned from the template, Vault installed from HashiCorp's apt repo, vault operator init run once and the unseal keys saved offline. Everything below configures that already-running server; it doesn't cover the install itself (HashiCorp's docs and the repo's SETUP.md have those commands). When I say "my Vault server," that's the box I mean.

Three things to build:

  1. Store the secrets (and give Vault a real TLS certificate)
  2. Write policies that scope what each part of the pipeline can read
  3. Configure JWT auth so CI jobs can authenticate with no stored token

TLS first, because self-signed is a trap

Vault holds every secret in this system. Talking to it over anything but a properly verified TLS connection is a non-starter — a self-signed cert you click past is a self-signed cert an attacker can impersonate. So Vault gets a real Let's Encrypt certificate.

The wrinkle: my Vault server isn't publicly reachable. Its DNS name resolves to an internal 192.168.1.x address. That rules out the HTTP-01 challenge (Let's Encrypt can't reach an internal host to verify it). The answer is the DNS-01 challenge, which proves domain control by setting a TXT record instead of serving a file — and that works regardless of whether the host is reachable from the internet.

I use acme-dns — a tiny DNS server purpose-built for exactly this. Rather than handing my main DNS provider's API credentials to every host that needs a cert, acme-dns handles only the _acme-challenge TXT records, and each client gets credentials scoped to a single subdomain. If a client is compromised, the blast radius is one challenge record, not my entire DNS zone.

The one-time DNS setup (in my public DNS provider) is three records:

Name Type Value
_acme-challenge.vault.example.com CNAME <uuid>.acmedns.example.com
acmedns.example.com A <acme-dns server IP>
acmedns.example.com NS acmedns.example.com

That NS record is the one everyone forgets — it delegates the acmedns.example.com subdomain to the acme-dns server itself, so that when Let's Encrypt follows the CNAME and queries for the TXT record, the authoritative answer comes from acme-dns. Without the NS delegation, the CNAME points into a void and the challenge fails with an error that does not obviously say "you're missing an NS record." I lost an evening to this.

Issuing the cert with acme.sh, on the Vault host:

# Register with the acme-dns server — returns username, password, subdomain, fulldomain
curl -s -X POST http://acmedns.example.com:8181/register -H "Content-Type: application/json"
# (save those four values to ~/.acme.sh/acmedns.json keyed by the domain)

curl https://get.acme.sh | sh -s email=you@example.com

export ACMEDNS_BASE_URL="http://acmedns.example.com:8181"
acme.sh --issue --force -d vault.example.com --dns dns_acmedns --server letsencrypt

Then install the cert and wire up a deploy hook that restarts Vault on renewal. The gotcha here: the deploy hook doesn't receive cert paths as environment variables the way you might expect — --install-cert handles the copy via its flags, and the reload command just fixes ownership and restarts:

sudo tee /opt/vault/tls/deploy.sh << 'EOF'
#!/bin/bash
set -e
chown vault:vault /opt/vault/tls/vault-fullchain.crt /opt/vault/tls/vault.key
chmod 640 /opt/vault/tls/vault-fullchain.crt /opt/vault/tls/vault.key
systemctl restart vault
EOF
sudo chmod +x /opt/vault/tls/deploy.sh

sudo -i
export ACMEDNS_BASE_URL="http://acmedns.example.com:8181"
acme.sh --install-cert -d vault.example.com \
  --fullchain-file /opt/vault/tls/vault-fullchain.crt \
  --key-file       /opt/vault/tls/vault.key \
  --reloadcmd      "/opt/vault/tls/deploy.sh"

Two more details that cost me time:

  • Use the fullchain, not the leaf cert. Some clients — including the Go HTTP client Terraform's Vault provider uses — reject a leaf-only cert because they can't build the trust chain. Point Vault's tls_cert_file at vault-fullchain.crt.
  • Binding to 443 as a non-root user. Vault runs as the vault user, and non-root can't bind ports below 1024. Rather than run Vault as root, grant the binary the capability:
# /etc/systemd/system/vault.service.d/override.conf
[Service]
AmbientCapabilities=CAP_NET_BIND_SERVICE

Now Vault serves HTTPS on 443 with a trusted cert, as an unprivileged user. VAULT_ADDR=https://vault.example.com — no port, no -tls-skip-verify, no warnings.

Storing the secrets

With TLS sorted, the secrets go in. KV v2 mounted at secret/. First the Proxmox token from Part 1:

vault kv put secret/proxmox \
  api_token="terraform@pam!gitops=xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"

Then an SSH keypair for Ansible. Generate a dedicated one — this key exists only to let the pipeline configure VMs, and keeping it separate from any personal key means you can rotate it without touching anything else:

ssh-keygen -t ed25519 -f ansible_key -C ansible

vault kv put secret/ansible \
  ssh_pubkey="$(cat ansible_key.pub)" \
  ssh_private_key="$(cat ansible_key)"

The public key gets injected into new VMs via cloud-init (so Ansible can connect); the private key is read by the Ansible stage to actually make the connection. Both halves live in Vault; neither ever touches GitLab or the repo.

Finally, the public keys for the human users the pipeline will create on each VM:

vault kv put secret/users/admin    ssh_pubkey="ssh-ed25519 AAAA... admin"
vault kv put secret/users/deployer ssh_pubkey="ssh-ed25519 AAAA... deployer"

Policies: the second trust boundary

In Part 1 the Proxmox token was the first least-privilege boundary. Vault policies are the second, and they're where the design gets deliberate.

The pipeline has two stages that touch secrets: Terraform (provisioning) and Ansible (configuration). They need different secrets, so they get different policies. Terraform never needs to read a user's SSH key; Ansible never needs the Proxmox token. Encoding that separation means a compromise or bug in one stage can't reach the other's secrets.

vault policy write terraform-policy - << 'EOF'
path "secret/data/proxmox" { capabilities = ["read"] }
path "secret/data/ansible" { capabilities = ["read"] }
EOF

vault policy write ansible-policy - << 'EOF'
path "secret/data/ansible"  { capabilities = ["read"] }
path "secret/data/users/*"  { capabilities = ["read"] }
EOF

Terraform reads the Proxmox token (to talk to the API) and the Ansible public key (to inject via cloud-init). Ansible reads the Ansible private key (to connect) and the user public keys (to install). The overlap on secret/ansible is intentional — both need it, for different halves of the keypair. Everything else is separated. Note the secret/data/ prefix: that's a KV v2 quirk — the API path for a secret at secret/proxmox is actually secret/data/proxmox.

The main event: JWT auth with zero stored tokens

Here's the problem this solves. A CI job needs to read from Vault. The naive approach is to store a Vault token in a GitLab CI variable and have jobs use it. But that token is long-lived, sits in GitLab's settings, and anyone with project access (or a leak of that variable) can use it forever. That's exactly the static-credential problem I set out to avoid.

The solution is JWT authentication, and the mental model that makes it click is two separate questions Vault asks about every request:

  1. Is this token genuine? (Did GitLab really issue it?)
  2. Is this token authorized? (Is it from my project, on my branch?)

Different mechanisms answer each. Get both, and you have authentication with no stored secret.

How GitLab issues the identity

When a CI job runs, GitLab can mint a JWT (JSON Web Token) scoped to that specific job. You request one by declaring it in the job:

id_tokens:
  VAULT_ID_TOKEN:
    aud: "${VAULT_ADDR}"

GitLab generates a fresh token for the job, signs it with its private key, and embeds claims describing the job — including which project and branch it's running from:

{
  "iss": "https://gitlab.com",
  "project_path": "your-username/proxmox-gitops",
  "ref": "main",
  "ref_type": "branch",
  "aud": "https://vault.example.com"
}

This token exists only for the life of the job. Nobody stored it. GitLab created it on the fly and it expires when the job ends.

Question 1: Is it genuine?

You configure Vault's JWT auth to trust GitLab's signing key:

vault auth enable jwt

vault write auth/jwt/config \
  jwks_url="https://gitlab.com/-/jwks" \
  bound_issuer="https://gitlab.com"

That jwks_url is GitLab's public key endpoint. GitLab signs every JWT with a private key that never leaves its servers, and publishes the matching public key at that URL. When a job presents a JWT, Vault fetches GitLab's public key and verifies the signature.

This is asymmetric cryptography doing its job: a signature made with GitLab's private key can only be verified with GitLab's public key, and only GitLab has the private key. If the token were forged or tampered with — even one character of one claim changed — the signature check fails. Passing it proves the token was issued by GitLab and hasn't been altered.

But — and this is the insight most Vault tutorials skip — a valid signature only proves it came from GitLab. It does not prove it came from your project. Any GitLab project, anywhere on gitlab.com, can mint a validly-signed JWT. Signature verification alone would let a stranger's repo authenticate to your Vault. That's what the second question is for.

Question 2: Is it authorized?

This is bound_claims, and it's the part that actually protects you:

vault write auth/jwt/role/gitlab-terraform \
  role_type="jwt" \
  user_claim="sub" \
  bound_claims_type="glob" \
  bound_claims='{"project_path":"your-username/proxmox-gitops","ref":"main","ref_type":"branch"}' \
  policies="terraform-policy" \
  ttl="1h"

vault write auth/jwt/role/gitlab-ansible \
  role_type="jwt" \
  user_claim="sub" \
  bound_claims_type="glob" \
  bound_claims='{"project_path":"your-username/proxmox-gitops","ref":"main","ref_type":"branch"}' \
  policies="ansible-policy" \
  ttl="1h"

After the signature checks out, Vault compares the claims inside the JWT against the bound_claims in the role. All of them must match. The JWT says project_path: your-username/proxmox-gitops — does it match the role? The JWT says ref: main — does it match?

And here's why it can't be faked: those claims are part of the signed payload. GitLab put them there and signed over them. An attacker on a different project can absolutely get GitLab to sign a JWT — but it'll carry their project's path, because GitLab fills in the claims based on where the job actually runs. They can't change project_path to yours without invalidating the signature, and they can't re-sign because they don't have GitLab's private key.

So the full gate looks like this:

JWT presented to Vault
   │
   ├── fetch GitLab's public key from /-/jwks
   ├── verify signature ─────────────────── fail → rejected (forged/tampered)
   ├── check iss == https://gitlab.com ──── fail → rejected
   ├── check aud == https://vault.example.com ── fail → rejected (wrong audience)
   ├── check project_path == yours ──────── fail → rejected (another project)
   ├── check ref == main ────────────────── fail → rejected (feature branch/fork)
   │
   └── all pass → issue a token with terraform-policy, 1-hour TTL

Two roles, two policies, same bound_claims. The Terraform stage authenticates as gitlab-terraform and gets terraform-policy; the Ansible stage authenticates as gitlab-ansible and gets ansible-policy. The trust boundary from earlier is now enforced at authentication time.

What the job actually does with it

The exchange itself is a single API call in the job's before_script:

export VAULT_TOKEN=$(curl -s \
  --request POST \
  --header "Content-Type: application/json" \
  --data "{\"role\":\"gitlab-terraform\",\"jwt\":\"${VAULT_ID_TOKEN}\"}" \
  "${VAULT_ADDR}/v1/auth/jwt/login" \
  | grep -o '"client_token":"[^"]*"' | cut -d'"' -f4)

The job hands Vault its GitLab-issued JWT, Vault runs the whole gate above, and hands back a scoped token good for one hour. That token can read only what its policy allows, and it evaporates when the job ends. At no point did a long-lived credential exist anywhere.

Why the aud field matters too

One easy-to-miss detail: the aud (audience) claim, which you set when declaring the token (aud: "${VAULT_ADDR}") and enforce via bound_issuer/audience checks. It scopes the token to your Vault specifically. Without it, a JWT minted for some other service that also trusts GitLab could theoretically be replayed against your Vault. Setting the audience to your Vault's address means a token minted for anything else won't be accepted here. It's defense in depth — the project and branch checks already do the heavy lifting, but there's no reason to leave the audience unbound.

Where this leaves us

Vault now holds every secret behind a trusted TLS endpoint, scoped into two policies that mirror the pipeline's two privileged stages, reachable by CI jobs that authenticate with a cryptographically-verified, short-lived, project-bound identity and no stored token anywhere.

The single value that lives in GitLab is VAULT_ADDR. A URL. If it leaks, someone learns my Vault's hostname — which does them no good without a validly-signed JWT from my exact project on my exact branch.

That's the security spine of the whole project, and everything from here hangs off it.

In Part 3 we get to Terraform: the reusable VM module, the environment that calls it, and how the Vault provider reads those secrets we just stored — plus storing Terraform state in GitLab's backend so the pipeline is stateless between runs.

Full code for the series is at github.com/jevellangelo/proxmox-gitops.


One operational note I'll expand on later: this Vault is manually unsealed. After a restart it needs vault operator unseal run with the unseal keys before it'll serve anything — which means if my Vault host reboots, the pipeline fails until I intervene. Auto-unseal is on the improvement list; for a homelab where I control the reboots, manual is an acceptable tradeoff.

GitHub - jevellangelo/proxmox-gitops: GitOps VM provisioning for Proxmox with Terraform, Ansible, Vault, and GitLab CI
GitOps VM provisioning for Proxmox with Terraform, Ansible, Vault, and GitLab CI - jevellangelo/proxmox-gitops
Building a GitOps VM Provisioning Platform on Proxmox — Part 3: Terraform & the GitLab State Backend
Part 2 got Vault holding every secret behind a trusted TLS endpoint, with CI jobs able to authenticate using nothing but a short-lived GitLab-issued JWT. Now we put that to work. This is the Terraform layer — the code that actually turns “I want a VM” into a running machine on
Building a GitOps VM Provisioning Platform on Proxmox — Part 1: Architecture & Proxmox Prep
I got tired of clicking through the Proxmox UI every time I wanted a VM. Not because clicking is hard — because clicking isn’t reproducible. I’d spin up a VM, install Docker, add my SSH key, set up a user, and a month later I’d have no record of what I

Part 1