Setting Up Horizontal Pod Autoscaler on Kubernetes: The Settings That Actually Matter

Setting Up Horizontal Pod Autoscaler on Kubernetes: The Settings That Actually Matter

By KP  |  TZoneLabs  |  DevOps & Cloud Engineering The default Horizontal Pod Autoscaler config scales on raw CPU with no stabilization window, which is exactly why it’s common to see replica counts bounce up and down every few minutes under normal, slightly-uneven traffic. The two settings that actually decide whether HPA is useful or … Read more

We Rotated a Database Password, and Every Pod That Didn’t Restart Started Failing Auth

We Rotated a Database Password, and Every Pod That Didn’t Restart Started Failing Auth

By KP  |  TZoneLabs  |  DevOps & Cloud Engineering A scheduled quarterly rotation replaced the database credential in Secrets Manager, and within minutes most of the fleet was using it without issue. Two hours later, a subset of pods started failing every query with authentication errors — pods that hadn’t restarted in weeks, still holding … Read more

Setting Up Kubernetes RBAC: The Roles and RoleBindings That Actually Enforce Least Privilege

Setting Up Kubernetes RBAC: The Roles and RoleBindings That Actually Enforce Least Privilege

By KP  |  TZoneLabs  |  DevOps & Cloud Engineering A Role scoped to one namespace, bound to one service account, with verbs limited to exactly what that workload does, is what turns Kubernetes RBAC into an actual security boundary instead of a checkbox. Most clusters we’ve audited get this half right: the Role exists, but … Read more

A Namespace Relabel Silently Broke a NetworkPolicy Selector, and Cross-Namespace Traffic Started Timing Out

A Namespace Relabel Silently Broke a NetworkPolicy Selector, and Cross-Namespace Traffic Started Timing Out

By KP  |  TZoneLabs  |  DevOps & Cloud Engineering A payments-team service started timing out on calls to a shared internal API, intermittently, with nothing in either side’s logs pointing at why. No connection refused, no TLS error, no 5xx. Just a hang until the client’s own timeout gave up. The pods were healthy, the … Read more

A CoreDNS Cache Setting Kept Routing gRPC Traffic to Pods That No Longer Existed After Every Deploy

A CoreDNS Cache Setting Kept Routing gRPC Traffic to Pods That No Longer Existed After Every Deploy

By KP  |  TZoneLabs  |  DevOps & Cloud Engineering Every deploy of one gRPC service left about 5% of calls failing for the next 30 to 40 seconds, then it cleared up on its own. The rollout itself looked clean, every new pod passed its readiness probe on schedule. The failing calls were connecting to … Read more

Setting Up Kubernetes Cluster Autoscaler on EKS: The Settings That Actually Matter

Setting Up Kubernetes Cluster Autoscaler on EKS: The Settings That Actually Matter

By KP  |  TZoneLabs  |  DevOps & Cloud Engineering Installing Cluster Autoscaler is one Helm command. Getting it to scale up fast enough and scale down without evicting something it shouldn’t is the part that takes actual configuration. The defaults are conservative on purpose, and left untouched they either leave pending pods waiting longer than … Read more

A Mistyped Route53 Zone ID Exhausted Our ACME Rate Limit — and Killed an Unrelated Cert

A Typo’d Route53 Zone ID Quietly Exhausted Our ACME Rate Limit, and It Took Down an Unrelated Production Certificate

By KP  |  TZoneLabs  |  DevOps & Cloud Engineering Our production API certificate expired on schedule for its routine 60-day renewal, except cert-manager never got a new one issued. The cause had nothing to do with that certificate. A staging Ingress with a typo’d Route53 zone ID had been failing its own renewal silently for … Read more

Our Rolling Deploys Were Sending Live Traffic to Pods That Weren’t Ready, Because the Readiness Probe Checked the Wrong Port

Our Rolling Deploys Were Sending Live Traffic to Pods That Weren’t Ready, Because the Readiness Probe Checked the Wrong Port

By KP  |  TZoneLabs  |  DevOps & Cloud Engineering Every rolling deploy for two months sent a burst of 502 errors to a small percentage of users, for about 30 seconds each time. Support tickets blamed “the app being slow sometimes.” The actual cause: our readiness probe was checking a port the container wasn’t listening … Read more

We Burned Two Weeks of Cloud Budget Because Our Autoscaler Was Stuck in a Scale-Up, Scale-Down Loop

We Burned Two Weeks of Cloud Budget Because Our Autoscaler Was Stuck in a Scale-Up, Scale-Down Loop

By KP  |  TZoneLabs  |  DevOps & Cloud Engineering For nine days, our cluster added and removed the same three worker nodes every six minutes. Every dashboard we watched stayed green the whole time. Pods were running, nodes were Ready, CPU never crossed 60%. Nobody noticed until the AWS bill landed about 40% higher than … Read more

Our TLS Certificate Expired in Production Because Every Alert We Had Was Watching the Wrong Thing

Our TLS Certificate Expired in Production Because Every Alert We Had Was Watching the Wrong Thing

By KP  |  TZoneLabs  |  DevOps & Cloud Engineering At 3 AM, our ingress started rejecting every HTTPS request with a certificate error. Every health check we had was green. Every pod was Running. The cluster looked perfectly healthy because none of our monitoring ever asked the one question that mattered: how many days is … Read more