Setting Up Prometheus and Grafana on Kubernetes: From Install to Alerts That Actually Fire

By KP  |  TZoneLabs  |  DevOps & Cloud Engineering

The kube-prometheus-stack Helm chart gets you Prometheus, Grafana, and Alertmanager running in about ten minutes. What it doesn’t do is scrape your own application’s metrics, write alerting rules that fire on the right thing, or send those alerts anywhere a human will see them. That part is configuration, not installation, and it’s where most setups stall out at “we have dashboards” without ever getting to “we get paged before the customer notices.”

This covers the install, scraping your own services, writing alerts that don’t flap or go silent, and the sizing mistakes that turn Prometheus itself into the thing that falls over.

Installing the Stack via Helm

helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo update

helm install kube-prometheus prometheus-community/kube-prometheus-stack \
  --namespace monitoring --create-namespace \
  --set prometheus.prometheusSpec.retention=15d \
  --set prometheus.prometheusSpec.storageSpec.volumeClaimTemplate.spec.resources.requests.storage=50Gi

This one command brings up Prometheus, Grafana, Alertmanager, node-exporter, and kube-state-metrics, wired together with sensible defaults. Setting the retention and storage size explicitly at install time matters more than it looks, the chart’s defaults are conservative and most teams hit the storage ceiling before they hit the retention one.

Scraping Your Own Application’s Metrics

The chart only scrapes what it knows about out of the box. Your own services need a ServiceMonitor pointing at wherever they expose a /metrics endpoint:

apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
  name: payments-api
  namespace: monitoring
  labels:
    release: kube-prometheus
spec:
  selector:
    matchLabels:
      app: payments-api
  namespaceSelector:
    matchNames:
      - payments
  endpoints:
    - port: metrics
      interval: 30s

The release: kube-prometheus label is easy to miss and silently means nothing gets scraped. The Prometheus Operator only picks up ServiceMonitors matching the label selector configured on the Prometheus resource itself, which defaults to matching the Helm release name.

Writing Alerting Rules That Don’t Flap

apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: payments-api-alerts
  namespace: monitoring
  labels:
    release: kube-prometheus
spec:
  groups:
    - name: payments-api
      rules:
        - alert: HighErrorRate
          expr: |
            sum(rate(http_requests_total{app="payments-api",status=~"5.."}[5m]))
            /
            sum(rate(http_requests_total{app="payments-api"}[5m])) > 0.05
          for: 10m
          labels:
            severity: warning
          annotations:
            summary: "Payments API error rate above 5% for 10 minutes"

The for: 10m is what separates a real alert from pager noise. Without it, a single bad scrape or a five-minute traffic blip fires the same alert as a genuine sustained failure, and the fastest way to get an alert ignored is to have it cry wolf a few times in its first week.

Wiring Alertmanager to Actually Notify Someone

apiVersion: monitoring.coreos.com/v1alpha1
kind: AlertmanagerConfig
metadata:
  name: payments-alerts-route
  namespace: monitoring
spec:
  route:
    receiver: payments-slack
    groupBy: ["alertname"]
    groupWait: 30s
    repeatInterval: 4h
  receivers:
    - name: payments-slack
      slackConfigs:
        - apiURL:
            name: slack-webhook
            key: url
          channel: "#payments-alerts"

groupBy and repeatInterval matter as much as the route itself. Without grouping, ten pods failing the same check sends ten separate messages instead of one; without a sane repeat interval, an unresolved alert either goes silent for a day or re-pages every five minutes.

Dashboards: Import First, Build Second

Grafana’s dashboard library already has well-built dashboards for the standard exporters, node-exporter’s dashboard (ID 1860) and the Kubernetes cluster overview (ID 315) cover most infrastructure-level visibility on day one:

# In Grafana: Dashboards → Import → paste the dashboard ID
# 1860 = Node Exporter Full
# 315  = Kubernetes cluster monitoring (via Prometheus)

Save custom dashboard work for the metrics only your own services expose, business-specific counters and latencies nobody else’s dashboard could know about. Rebuilding a generic infrastructure dashboard from scratch is time spent on something already solved.

Common Mistakes

  • No resource limits on Prometheus itself. A high-cardinality metric or a sudden increase in scrape targets can push Prometheus’s memory usage past a node’s capacity, and an OOMKilled Prometheus pod is the one outage a monitoring stack isn’t supposed to cause.
  • Unbounded label cardinality. A label carrying a user ID or a raw request path turns a handful of time series into millions, and cardinality growth like that shows up as Prometheus slowly running out of memory over days, not immediately.
  • Alerts with no for duration. Fires on every transient blip, trains the team to ignore that alert specifically, right up until the blip is real.
  • Storage sized for today’s data, not six months from now. Retention and PVC size are both easy to change later, but only if someone notices the disk filling up before Prometheus starts silently dropping data.

Key Lessons

  1. Installing the stack and getting useful alerts out of it are two different projects.
    The Helm chart is the easy part; scrape configs, alert thresholds, and routing are where the real work is.
  2. A missing release label on a ServiceMonitor or PrometheusRule means it’s silently ignored.
    Check the Prometheus Operator’s own selector before assuming a scrape config or alert rule is even being read.
  3. A for duration is what makes an alert trustworthy.
    Without it, a transient blip pages the same as a sustained outage, and the alert gets tuned out either way.
  4. Label cardinality is the failure mode nobody notices until Prometheus is already struggling.
    Avoid putting anything with unbounded values, user IDs, raw paths, into a metric label.
  5. Import the standard dashboards, build only what’s actually specific to your service.
    Node and cluster-level visibility is a solved problem; don’t spend custom dashboard time re-solving it.

Summary

Component Purpose Watch For
ServiceMonitor Tells Prometheus what to scrape Missing release label = silently never scraped
PrometheusRule Defines alerting conditions No for duration turns blips into pages
AlertmanagerConfig Routes alerts to a real notification channel Missing grouping floods the channel with duplicates
Storage sizing Retention and PVC size for the TSDB Undersized storage silently drops old data first

Most of what makes a monitoring stack actually useful happens after the install finishes, in the scrape configs and alert rules nobody sees in a demo.

Read Next

If you’re running Kubernetes in production, follow along on LinkedIn for more guides like this one as they’re published.


Tags:
#Prometheus   #Grafana   #Kubernetes   #Monitoring  
#DevOps   #SRE

Leave a Comment