Our TLS Certificate Expired in Production Because Every Alert We Had Was Watching the Wrong Thing

Our TLS Certificate Expired in Production Because Every Alert We Had Was Watching the Wrong Thing

By KP  |  TZoneLabs  |  DevOps & Cloud Engineering At 3 AM, our ingress started rejecting every HTTPS request with a certificate error. Every health check we had was green. Every pod was Running. The cluster looked perfectly healthy because none of our monitoring ever asked the one question that mattered: how many days is … Read more

Why Your PodDisruptionBudget Didn’t Save the Last EKS Upgrade

Why Your PodDisruptionBudget Didn’t Save the Last EKS Upgrade

By KP  |  TZoneLabs  |  DevOps & Cloud Engineering A PodDisruptionBudget doesn’t protect your application — it protects a number. That’s why a routine EKS managed node group upgrade recently stalled on us for 40 minutes and paged on-call at 2am, even though our minAvailable: 2 config was textbook. The PDB was doing exactly what … Read more

During a Production Failure, the Real Issue Is Often Not Where the Error Is Showing

During a Production Failure, the Real Issue Is Often Not Where the Error Is Showing

By KP  |  TZoneLabs  |  DevOps & Cloud Engineering Here is something nobody tells you when you start in DevOps: the error message is almost never the problem. It is just the messenger. A pod crashes — you blame the application. An API times out — you blame the network. A deployment fails — you … Read more

We Lost 3 Hours of Production Deployments Because of One Silent Node Provisioning Failure

We Lost 3 Hours of Production Deployments Because of One Silent Node Provisioning Failure

By KP  |  TZoneLabs  |  DevOps & Cloud Engineering We were in the middle of a production scaling event when everything went quiet — in the worst way possible. No crash. No alert. No obvious Kubernetes error. Just pods stuck in Pending, GitHub Actions deployment jobs timing out, and the entire team staring at dashboards … Read more