Skip to content

etcd Backup & Restore

Kubernetes relies on etcd for storing critical cluster data, making regular backups essential for disaster recovery. This section explains how to perform etcd backups using native tools and Kubernetes-native methods, and how to restore them in case of failure.


Backup Methods

1. Using etcdctl (Native Tool)

The etcdctl CLI tool allows direct interaction with etcd. This method is ideal for on-premises clusters.

Prerequisites: - Access to etcd endpoints (e.g., https://<etcd-ip>:2379). - etcdctl installed and configured with proper SSL/TLS certificates.

Backup Command:

ETCDCTL_API=3 etcdctl \
  --endpoints=https://<etcd-ip>:2379 \
  --cacert=/etc/etcd/ca.pem \
  --cert=/etc/etcd/kubernetes.pem \
  --key=/etc/etcd/kubernetes.key \
  snapshot save /var/lib/etcd/backup.db
- Replace <etcd-ip> with your etcd node's IP. - Ensure the backup path exists and has sufficient permissions.

Automate with CronJob (Optional):

apiVersion: batch/v1
kind: CronJob
metadata:
  name: etcd-backup
spec:
  schedule: "0 2 * * *"
  jobTemplate:
    spec:
      template:
        spec:
          containers:
          - name: backup
            image: etcd:3.5.0
            command:
            - /bin/sh
            - -c
            - |
              ETCDCTL_API=3 etcdctl --endpoints=https://<etcd-ip>:2379 \
                --cacert=/etc/etcd/ca.pem \
                --cert=/etc/etcd/kubernetes.pem \
                --key=/etc/etcd/kubernetes.key \
                snapshot save /backup/etcd.db
            volumeMounts:
            - name: backup-volume
              mountPath: /backup
          volumes:
          - name: backup-volume
            emptyDir: {}


2. Using Kubernetes Backup CRD (Operator-Managed)

Kubernetes distributions like Rancher or managed services (e.g., GKE, EKS) often provide a Backup custom resource for etcd backups. Note that the API used depends on the specific operator or service.

Example Backup YAML:

apiVersion: etcd.example.com/v1beta1
kind: Backup
metadata:
  name: etcd-backup
  namespace: etcd-system
spec:
  etcdEndpoints:
  - https://<etcd-ip>:2379
  storagePath: /var/lib/etcd/backups/backup.db
  snapshot: true
- Apply the YAML: kubectl apply -f backup.yaml. - The operator handles the backup process, storing the snapshot in the specified path.


Restore Methods

1. Restore from etcdctl Snapshot

  1. Stop etcd service (on the node where etcd runs):
    systemctl stop etcd
    
  2. Replace etcd data directory with the backup:
    mv /var/lib/etcd/ /var/lib/etcd-old/
    cp -r /backup/etcd.db /var/lib/etcd/
    
  3. Restart etcd:
    systemctl start etcd
    

Note: Ensure the backup was created with the same etcd version. Mismatched versions can cause corruption.


2. Restore via Kubernetes Backup CRD

  1. Create a Restore resource specifying the backup:
    apiVersion: etcd.example.com/v1beta1
    kind: Restore
    metadata:
      name: etcd-restore
      namespace: etcd-system
    spec:
      etcdEndpoints:
      - https://<etcd-ip>:2379
      storagePath: /var/lib/etcd/backups/backup.db
    
  2. Apply the YAML: kubectl apply -f restore.yaml.
  3. Monitor the restore progress using the operator's metrics or logs.

Key Takeaways

  • Regular backups are critical for etcd, as data loss can lead to cluster downtime.
  • Use etcdctl for direct control or Kubernetes Backup CRD for managed environments.
  • Verify backups by testing restores in a staging environment.
  • Always validate etcd version compatibility before restoring to avoid corruption.