SPIRE workload identity on Kubernetes

14 min read · Advanced


This guide sets up SPIFFE identities for Edera pods with a standard SPIRE server and the SPIRE Kubernetes registrar (spire-controller-manager). At the end, each Edera pod gets its registrar-written identity in one or both of two ways:

  • Workload API. The pod’s containers fetch SVIDs from a SPIRE agent socket.
  • Mounted identity. Edera writes the SVID, its key, and the trust bundle into the container as files, for workloads that cannot speak the Workload API.

The same ClusterSPIFFEID serves both. Nothing is registered per workload.

For what SPIFFE and SPIRE are, and how Edera runs a SPIRE agent in each zone, see Workload identity with SPIFFE and SPIRE.

⚠️
SPIRE support is alpha. Evaluate it before you rely on it as a trust boundary.

How it works

Every Edera pod runs in its own zone, and each zone runs its own SPIRE agent. The agent does not attest with a Kubernetes token. It presents evidence: a statement about the zone that the Edera daemon on the node signs with a node key.

Each node gets its node key from SPIRE. A small DaemonSet, the node enroller, proves which node it runs on with its service account token, and SPIRE issues it an identity for that node. The enroller writes that identity’s key where the daemon reads it, and renews it before it expires.

The SPIRE server’s Edera plugin checks that the evidence is signed by the node the pod runs on, checks the pod against the Kubernetes API, and names the agent after the pod:

spiffe://<trust domain>/spire/agent/edera-hypervisor/<node name>/pod/<pod uid>

The registrar renders each Edera pod’s entries under that agent, so a zone can only ever hold identities for its own pod.

Prerequisites

  • A Kubernetes cluster with Edera installed on the nodes that run Edera pods. See the installation guides.
  • The API server runs with the Node authorizer and the NodeRestriction admission plugin. Most distributions turn both on by default.
  • kubectl, helm, and docker (or crane) on your workstation.
  • Shell access to each Edera node, to configure the daemon.
  • Outbound access from the cluster to ghcr.io.

This guide was tested with these versions:

ComponentVersion
spire-crds chart0.6.1
spire chart0.30.2
Edera spire-controller-manager0.7.1-edera.0 (ghcr.io/edera-dev/spire-controller-manager@sha256:4b0bb917ff67b00fc8df87d05660b59e9d4d70d1e6f00f125727896047f3568d)
Edera attestor pluginsv0.0.8 (ghcr.io/edera-dev/edera-attestor-plugins@sha256:7294b4a58c29f71034afe5a182fd59e66d6e12c3d4b475248aa9d75eae3b632f)
edera-node-enroller chart1.2.0

The registrar image is Edera’s fork of spire-controller-manager. The standard upstream spire charts and CRDs can be used with it unchanged.

Set these variables for the rest of the guide:

export TRUST_DOMAIN=example.org      # your SPIFFE trust domain
export CLUSTER_NAME=my-cluster       # the registrar's cluster name
export NAMESPACE=spire-system
export REGISTRAR_CLASS=spire         # the registrar's class name
export SPIRE_SERVER=spire-server.${NAMESPACE}:443   # the SPIRE server that enrollers reach, as host:port
export PLUGINS_VERSION=v0.0.8
export PLUGINS_IMAGE=ghcr.io/edera-dev/edera-attestor-plugins:${PLUGINS_VERSION}

Step 1: Check the RuntimeClass

The registrar decides which parent an entry gets from the pod’s node selector, so every Edera pod needs runtime: edera in its node selector. The standard edera RuntimeClass adds it. If you have not applied it and labeled the Edera nodes, follow Apply the Edera RuntimeClass and Label nodes for Edera workloads.

Check that the RuntimeClass has the node selector:

kubectl get runtimeclass edera -o jsonpath='{.scheduling.nodeSelector}'
# {"runtime":"edera"}

A pod’s node selector cannot change after the pod is scheduled, which is why the registrar uses it rather than a label or an annotation.

Step 2: Install SPIRE

Create the namespace, and install the standard CRDs:

kubectl create namespace "$NAMESPACE"
helm upgrade --install spire-crds spire-crds \
  --repo https://spiffe.github.io/helm-charts-hardened --version 0.6.1 \
  --namespace "$NAMESPACE" --wait

The chart requires the checksum of the server plugin binary. Take it from the plugin image:

docker pull "$PLUGINS_IMAGE"
id=$(docker create --entrypoint /bin/true "$PLUGINS_IMAGE")
docker cp "$id:/var/lib/edera/protect/attestor-plugins/edera-server-nodeattestor" .
docker rm "$id"
export PLUGIN_SHA256=$(sha256sum edera-server-nodeattestor | cut -d' ' -f1)

Write the chart values. The server loads the Edera plugin twice: as edera-hypervisor, which attests zones, and as edera-node, which enrolls nodes.

cat > spire-values.yaml <<EOF
global:
  spire:
    trustDomain: ${TRUST_DOMAIN}
    clusterName: ${CLUSTER_NAME}
spire-server:
  controllerManager:
    enabled: true
    className: ${REGISTRAR_CLASS}
    image:
      registry: ghcr.io
      repository: edera-dev/spire-controller-manager
      tag: "0.7.1-edera.0"
    parentIDTemplate: '{{ if eq (index .PodSpec.NodeSelector "runtime") "edera" }}spiffe://{{ .TrustDomain }}/spire/agent/edera-hypervisor/{{ .NodeMeta.Name }}/pod/{{ .PodMeta.UID }}{{ else }}spiffe://{{ .TrustDomain }}/spire/agent/k8s_psat/{{ .ClusterName }}/{{ .NodeMeta.UID }}{{ end }}'
  customPlugins:
    nodeAttestor:
      edera-hypervisor:
        plugin_cmd: /var/lib/edera/protect/attestor-plugins/edera-server-nodeattestor
        plugin_checksum: ${PLUGIN_SHA256}
        image:
          registry: ghcr.io
          repository: edera-dev/edera-attestor-plugins
          tag: "${PLUGINS_VERSION}"
        plugin_data:
          mode: evidence
          node_identity: spire
      edera-node:
        plugin_cmd: /var/lib/edera/protect/attestor-plugins/edera-server-nodeattestor
        plugin_checksum: ${PLUGIN_SHA256}
        image:
          registry: ghcr.io
          repository: edera-dev/edera-attestor-plugins
          tag: "${PLUGINS_VERSION}"
        plugin_data:
          mode: enrollment
          service_account_allow_list: ["${NAMESPACE}:edera-node-enroller"]
EOF

The parent ID template puts an Edera pod’s entries under that pod’s Edera agent, and every other pod keeps the chart’s normal node agent.

Leave the chart’s PSAT node attestor enabled, which is its default. It also grants the SPIRE server’s service account what the Edera plugin needs: get on pods and nodes, and create on token reviews.

Keep spire-agent and spiffe-csi-driver enabled if non-Edera pods in the cluster need identities. Edera pods use neither.

Install:

helm upgrade --install spire spire \
  --repo https://spiffe.github.io/helm-charts-hardened --version 0.30.2 \
  --namespace "$NAMESPACE" --values spire-values.yaml --wait --timeout 10m

Step 3: Let each zone fetch mounted identities

Edera writes a mounted identity through the SPIRE agent’s Delegated Identity API, which needs a delegate identity for the zone’s own agent. One cluster-wide ClusterSPIFFEID issues it for every pod:

kubectl apply -f - <<EOF
apiVersion: spire.spiffe.io/v1alpha1
kind: ClusterSPIFFEID
metadata:
  name: edera-zone-agent
  annotations:
    spire-controller-manager.edera.dev/additive: "true"
spec:
  className: ${REGISTRAR_CLASS}
  spiffeIDTemplate: "spiffe://{{ .TrustDomain }}/edera/pod/{{ .PodMeta.UID }}/zone-agent"
  workloadSelectorTemplates:
    - "edera:principal:zone-agent"
EOF

Keep the annotation. The chart’s default ClusterSPIFFEID is a fallback, and any other ClusterSPIFFEID that matches a pod suppresses fallbacks. Without the annotation, this object takes away every pod’s default identity. The annotation only works with the Edera registrar image.

Skip this step if no workload uses mounted identities.

Step 4: Point the Edera daemon on each node at the server

Get the server’s address:

kubectl -n "$NAMESPACE" get service spire-server -o jsonpath='{.spec.clusterIP}'

On each Edera node, add this to /var/lib/edera/protect/daemon.toml:

[features]
spire-v0 = "enabled"
object-capabilities-v0 = "enabled"

[spire]
trust-domain = "example.org"          # the same as TRUST_DOMAIN
server-address = "10.96.0.123"        # the spire-server ClusterIP
server-port = 443
trust-bundle-path = "/var/lib/edera/protect/spire/trust/server-bundle.pem"
node-key-path = "/var/lib/edera/protect/spire/node/node.key"
node-certificate-path = "/var/lib/edera/protect/spire/node/node.crt"
# evidence-lifetime-seconds = 300     # default 300, from 1 to 86400

Then restart the daemon:

sudo systemctl restart protect-daemon

The bundle and the node key do not exist yet. The chart in the next step creates both. Until then, the daemon starts each Edera pod’s zone but holds its SPIRE agent back.

See the daemon.toml reference for every [spire] option.

Step 5: Enroll the nodes

Install the node enroller chart into the SPIRE namespace:

helm upgrade --install edera-node-enroller \
  oci://ghcr.io/edera-dev/charts/edera-node-enroller --version 1.2.0 \
  --namespace "$NAMESPACE" \
  --set trustDomain="$TRUST_DOMAIN" \
  --set serverAddress="$SPIRE_SERVER" \
  --set image.tag="$PLUGINS_VERSION"

The chart runs a DaemonSet on the nodes labeled runtime: edera. Each pod does two things:

  1. Its init container copies the server’s bundle to the node. It waits until the SPIRE server has published its bundle to the spire-bundle ConfigMap, writes it as the daemon’s trust-bundle-path, and exits. It holds no token and no key.
  2. The enroller enrolls the node. It writes node.key and node.crt, and renews them at half their lifetime. The daemon picks up each new pair with no restart. New nodes enroll as soon as the DaemonSet reaches them.

Expected Result: every Edera node is listed as an enrolled agent.

kubectl -n "$NAMESPACE" logs daemonset/edera-node-enroller
kubectl -n "$NAMESPACE" exec spire-server-0 -c spire-server -- \
  /opt/spire/bin/spire-server agent list
# ... spiffe://example.org/spire/agent/edera-node/<node uid>

Chart values

ValueDefaultMeaning
trustDomainRequired. The same as TRUST_DOMAIN.
serverAddressspire-server.<namespace>:443The SPIRE server, as host:port.
serviceAccount.nameedera-node-enrollerMust be the entry in the edera-node attestor’s service_account_allow_list (step 2).
bundle.configMap, bundle.keyspire-bundle, bundle.spiffeThe chart’s bundle ConfigMap. Change them if you set global.spire.bundleConfigMap, or bundlePublisher.k8sConfigMap.format: pem (key bundle.crt).
hostPaths.nodeKeyDirectory, hostPaths.trustDirectory/var/lib/edera/protect/spire/node, /var/lib/edera/protect/spire/trustMust match the paths in daemon.toml (step 4).
nodeSelector, tolerationsruntime: edera, noneWhere the enroller runs. Keep it to Edera nodes.

Properties to keep

  • The service account is only for the enroller, and only it is named in service_account_allow_list. Never add it to the chart’s PSAT allow list (nodeAttestor.k8sPSAT.serviceAccountAllowList). If you do, the enroller can take the node’s standard agent ID, and its key can fetch every non-Edera pod identity on the node.
  • The token automount is off, and the projected token’s audience is edera-node. It is not the standard agents’ spire-server, so a token for one attestor never passes the other.
  • The enroller container mounts the trust directory read-only, so it cannot change the bundle that the daemon and every new zone trust. Only the init container writes it. No other host path is mounted, and no port is exposed.
  • No container has a Linux capability, and the pod sets no runtimeClassName. The server refuses an enroller pod that sets one.

Step 6: Run a workload

Ask for an identity with a pod annotation. Use either or both:

AnnotationEffect
dev.edera/spire-access: "true"The containers get the SPIRE agent’s Workload API socket at /shared/zone/sockets/agent.sock.
dev.edera/spire-identity: "true"Edera writes the pod’s SVIDs into each container at /run/spire/identity/: svid.0.pem, svid.0.key, and bundle.<trust domain>.pem, and keeps them fresh.
apiVersion: v1
kind: Pod
metadata:
  name: hello-spiffe
  annotations:
    dev.edera/spire-access: "true"
    dev.edera/spire-identity: "true"
spec:
  runtimeClassName: edera
  containers:
    - name: app
      image: ghcr.io/spiffe/spire-agent:1.15.3
      command: ["/opt/spire/bin/spire-agent", "api", "watch",
                "-socketPath", "/shared/zone/sockets/agent.sock"]

With the chart’s default ClusterSPIFFEID, this pod gets spiffe://example.org/ns/default/sa/default by either path.

For a SPIFFE library, set SPIFFE_ENDPOINT_SOCKET=unix:///shared/zone/sockets/agent.sock.

Verification

The pod’s zone attests as its own agent:

kubectl -n "$NAMESPACE" exec spire-server-0 -c spire-server -- \
  /opt/spire/bin/spire-server agent list
# ... spiffe://example.org/spire/agent/edera-hypervisor/<node>/pod/<pod uid>

The registrar wrote the pod’s entry under that agent:

POD_UID=$(kubectl get pod hello-spiffe -o jsonpath='{.metadata.uid}')
kubectl -n "$NAMESPACE" exec spire-server-0 -c spire-server -- \
  /opt/spire/bin/spire-server entry show \
  -parentID "spiffe://$TRUST_DOMAIN/spire/agent/edera-hypervisor/<node>/pod/$POD_UID"

The workload receives it both ways:

kubectl exec hello-spiffe -- /opt/spire/bin/spire-agent api fetch x509 \
  -socketPath /shared/zone/sockets/agent.sock
kubectl logs hello-spiffe          # the api watch output

For a container with a shell, cat /run/spire/identity/svid.0.pem shows the same SPIFFE ID in the mounted certificate.

Scoping identities

Write ClusterSPIFFEID objects as for any SPIRE deployment. The k8s: selectors below have exactly the meaning the upstream Kubernetes workload attestor gives them, so existing ClusterSPIFFEID objects work for Edera pods unchanged.

SelectorMeaning
k8s:pod-uid, k8s:ns, k8s:pod-nameThe pod
k8s:container-nameThe container
k8s:node-nameThe node
k8s:container-imageThe container’s image and imageID from the pod status. Under CRI-O, only the image form.
k8s:saThe pod’s service account
k8s:pod-image-count, k8s:pod-init-image-countHow many containers and init containers the pod spec lists. Left out once the pod has run an ephemeral container, so an entry that needs them stops matching.

Two upstream selectors have an Edera variant instead, and some are not available:

Instead ofUseDifference
k8s:pod-image, k8s:pod-init-imageedera:k8s-pod-imageThe images of every container the pod has run, including init and ephemeral containers, and ones that exited
k8s:pod-label, k8s:pod-owner, k8s:pod-owner-uid, k8s:ns-labelNot availableAn entry that needs them matches nothing

This example gives one identity to the api container of app: payments pods, only under the pod’s own service account. It never matches a pod that has run a container its spec does not list, such as an ephemeral debug container:

apiVersion: spire.spiffe.io/v1alpha1
kind: ClusterSPIFFEID
metadata:
  name: payments-api
spec:
  className: spire
  podSelector:
    matchLabels:
      app: payments
  spiffeIDTemplate: "spiffe://{{ .TrustDomain }}/payments/api"
  workloadSelectorTemplates:
    - "k8s:container-name:api"
    - "k8s:sa:{{ .PodSpec.ServiceAccountName }}"
    - "k8s:pod-image-count:{{ len .PodSpec.Containers }}"

The edera: selectors also work, for entries written by hand.

Operations

  • Node key renewal. Automatic. Each node key is an agent SVID with the server’s agent SVID lifetime, one hour by default. The enroller renews it at half that. That lifetime is also how long a stolen node key stays usable, so keep the chart’s agent SVID TTL at one hour or less.
  • Removing a node. Delete its Node object. The server then refuses its evidence at once, because the node key names the node’s UID. An enroller cannot renew for a node that no longer exists.
  • SPIRE CA rotation. Nothing to do. The chart rotates the CA daily. trust-bundle-path is only the first trust anchor. The daemon and the enroller each read the server’s current bundle over TLS that the bundle they already hold authenticates. The daemon reads it every 5 minutes and at each zone launch, and the enroller every 5 minutes and at each enrollment. SPIRE publishes each new CA hours before it signs with it. If a node is off for longer than the CA lifetime, one day by default, delete its enroller pod. The new pod’s init container copies the current bundle from the ConfigMap.
  • Evidence lifetime. The server refuses evidence valid for longer than the plugin’s max_evidence_lifetime_seconds (default 86400). Lower it in plugin_data to enforce a shorter limit than the daemons use.
  • Server outside the cluster. The Edera plugin must read pods and nodes from the Kubernetes API. Run a nested SPIRE server in the cluster with the plugin, and make the outside server its upstream authority.
  • Zones that are not a pod’s. A node key from SPIRE can sign only for a pod, because only the pod gives the node’s name. Zones started outside Kubernetes need a node certificate from your own CA, the plugin’s node_identity = "certificate" mode.

Security notes

  • Who controls node identity. SPIRE’s CA issues the node keys. Anyone who can run a pod, or mint a token, as the enroller’s service account can enroll the node that pod runs on. Treat the rights in $NAMESPACE as node credentials, as upstream SPIRE already requires for its own agent.
  • The bundle ConfigMap is a trust anchor. Whoever can change spire-bundle in $NAMESPACE decides what each node trusts at its enroller’s next pod start, as for the standard SPIRE agent. The same namespace rights can already enroll nodes. For an out-of-band check, leave out the init container, and install server-bundle.pem on each node yourself.
  • A zone cannot enroll. The plugin refuses an enroller token from any pod that sets a runtime class. A workload in a zone never gets a node key, even if it runs with the enroller’s service account, and whatever the Edera RuntimeClass is named.
  • A node key signs for its own node only. It can claim any Edera pod on its node, and no pod on any other node or in any other cluster. It grants nothing else: the enroller’s agent has no selectors and no entries.
  • spiffe://<trust domain>/edera/ is reserved. It holds each pod zone’s delegate identity, which can fetch every identity in that zone. Only the ClusterSPIFFEID from step 3 may render IDs under it.
  • Parents. Keep entries for Edera pods under the pod’s agent, which the registrar does. If you write a node alias by hand, select on edera-hypervisor:node-id together with edera-hypervisor:zone-id or edera-hypervisor:k8s-pod-uid. An alias that spans many zones gives its identities to every zone under it.

Troubleshooting

Read the SPIRE server log first:

kubectl -n "$NAMESPACE" logs spire-server-0 -c spire-server
SymptomCause
node enrollment refused: ... is not an allowed enroller service accountThe enroller’s service account is not in service_account_allow_list.
node enrollment refused: ... must run on the default runtimeThe enroller pod sets runtimeClassName. Remove it.
node enrollment refused: the token is not for audience "edera-node"The chart’s token.audience differs from the plugin’s audience.
node enrollment refused: the token was issued on node ...The node was replaced under the same name, and a token from the old node was used. Delete the old enroller pod.
The enroller logs the server's certificate is not for spiffe://.../spire/serverThe enroller’s trust domain differs from the chart’s trustDomain.
The enroller pod stays in InitThe server has not published its bundle to spire-bundle, or the ConfigMap has another name or format. See chart values.
evidence did not verify: node SVID ... is not an enrolled nodeThe node key is not from the enroller. Check node-key-path and node-certificate-path.
statement is for trust domain "a", not "b"The daemon’s trust-domain differs from the chart’s trustDomain.
statement is valid for ..., longer than ...evidence-lifetime-seconds is above the plugin’s max_evidence_lifetime_seconds.
pod not confirmed: ... could not read nodeThe server’s service account cannot read nodes. Keep the chart’s PSAT node attestor enabled.
The daemon logs waiting for the SPIRE node keyThe node has not enrolled. Check the enroller on that node.
The agent is listed, but the pod has no entry under itThe pod’s node selector lacks runtime: edera, so its entry went under a node agent. Check the RuntimeClass (step 1).
Pods lost their default identityThe step 3 object is missing its annotation, or the registrar is not the Edera image.
No mounted identity, but the Workload API worksThe step 3 object is missing, or its className differs from the registrar’s.
The daemon does not startCheck journalctl -u protect-daemon. server-address without node-key-path, node-certificate-path, and trust-bundle-path stops it.
The daemon logs starting without a SPIRE node keyThe node key or its certificate is missing, cannot be read, is readable by others, or does not match. The daemon runs, and uses the pair once a good one appears.
Entries written by hand disappearaddEntryIDPrefix is off, so the registrar deletes every entry that no ClusterSPIFFEID declares. Keep the chart default, which turns it on.

See also

Last updated on