Troubleshooting
Kubernetes environments often face challenges that require deep troubleshooting skills. This section explores common issues like node taints, resource contention, and network policies, along with practical strategies to diagnose and resolve them.
Node Taints and Scheduling Conflicts¶
Node taints prevent pods from scheduling on specific nodes unless they have matching tolerations. Misconfigured taints can lead to pods being evicted or failing to deploy.
Diagnosis:
Use kubectl describe node <node-name> to inspect taints. Look for entries like node-role.kubernetes.io/worker:NoSchedule.
Mitigation:
1. Remove or adjust taints:
Example: A deployment fails with Taints invalid. Add tolerations to the pod spec or adjust node taints to align with the workload requirements.
Resource Contention and Eviction Pressure¶
Resource contention occurs when nodes exceed CPU/memory limits, leading to pod evictions. This is often due to misconfigured resource requests/limits or insufficient cluster sizing.
Diagnosis:
Check node resource usage:
kubectl describe pod <pod-name> to see eviction events:Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Warning Evicted 2m kubelet Running pod was evicted due to memory pressure
Mitigation:
1. Set proper resource limits:
spec:
containers:
- name: my-app
resources:
requests:
memory: "256Mi"
cpu: "500m"
limits:
memory: "512Mi"
cpu: "1000m"
3. Prioritize workloads: Use Kubernetes Quality of Service (QoS) classes to ensure critical pods are not evicted.
Example: A database pod is evicted due to memory pressure. Increase its memory limit and ensure the node has sufficient capacity.
Network Policies and Connectivity Issues¶
Network policies restrict traffic between pods, and misconfigurations can block essential communication (e.g., between services or external endpoints).
Diagnosis:
List active network policies:
Test connectivity using
curl or telnet:Mitigation:
1. Adjust policy rules:
3. Verify DNS resolution: Ensure services are reachable via DNS names (e.g.,
my-service.namespace.svc.cluster.local).
Example: A service fails to connect to a backend. Check the network policy to ensure ingress rules allow traffic from the service's namespace.
Key takeaways¶
- Node taints require balancing tolerations and taint configurations to ensure proper scheduling.
- Resource contention demands careful monitoring and proper resource allocation to prevent evictions.
- Network policies must be tested rigorously to avoid blocking critical communication paths.
- Always combine diagnostic commands (
kubectl describe,top,curl) with policy reviews to isolate root causes.