etcd Backup & Restore
Kubernetes relies on etcd for storing critical cluster data, making regular backups essential for disaster recovery. This section explains how to perform etcd backups using native tools and Kubernetes-native methods, and how to restore them in case of failure.
Backup Methods¶
1. Using etcdctl (Native Tool)¶
The etcdctl CLI tool allows direct interaction with etcd. This method is ideal for on-premises clusters.
Prerequisites:
- Access to etcd endpoints (e.g., https://<etcd-ip>:2379).
- etcdctl installed and configured with proper SSL/TLS certificates.
Backup Command:
ETCDCTL_API=3 etcdctl \
--endpoints=https://<etcd-ip>:2379 \
--cacert=/etc/etcd/ca.pem \
--cert=/etc/etcd/kubernetes.pem \
--key=/etc/etcd/kubernetes.key \
snapshot save /var/lib/etcd/backup.db
<etcd-ip> with your etcd node's IP.
- Ensure the backup path exists and has sufficient permissions.
Automate with CronJob (Optional):
apiVersion: batch/v1
kind: CronJob
metadata:
name: etcd-backup
spec:
schedule: "0 2 * * *"
jobTemplate:
spec:
template:
spec:
containers:
- name: backup
image: etcd:3.5.0
command:
- /bin/sh
- -c
- |
ETCDCTL_API=3 etcdctl --endpoints=https://<etcd-ip>:2379 \
--cacert=/etc/etcd/ca.pem \
--cert=/etc/etcd/kubernetes.pem \
--key=/etc/etcd/kubernetes.key \
snapshot save /backup/etcd.db
volumeMounts:
- name: backup-volume
mountPath: /backup
volumes:
- name: backup-volume
emptyDir: {}
2. Using Kubernetes Backup CRD (Operator-Managed)¶
Kubernetes distributions like Rancher or managed services (e.g., GKE, EKS) often provide a Backup custom resource for etcd backups. Note that the API used depends on the specific operator or service.
Example Backup YAML:
apiVersion: etcd.example.com/v1beta1
kind: Backup
metadata:
name: etcd-backup
namespace: etcd-system
spec:
etcdEndpoints:
- https://<etcd-ip>:2379
storagePath: /var/lib/etcd/backups/backup.db
snapshot: true
kubectl apply -f backup.yaml.
- The operator handles the backup process, storing the snapshot in the specified path.
Restore Methods¶
1. Restore from etcdctl Snapshot¶
- Stop etcd service (on the node where etcd runs):
- Replace etcd data directory with the backup:
- Restart etcd:
Note: Ensure the backup was created with the same etcd version. Mismatched versions can cause corruption.
2. Restore via Kubernetes Backup CRD¶
- Create a
Restoreresource specifying the backup: - Apply the YAML:
kubectl apply -f restore.yaml. - Monitor the restore progress using the operator's metrics or logs.
Key Takeaways¶
- Regular backups are critical for etcd, as data loss can lead to cluster downtime.
- Use
etcdctlfor direct control or Kubernetes Backup CRD for managed environments. - Verify backups by testing restores in a staging environment.
- Always validate etcd version compatibility before restoring to avoid corruption.