NineLabNineLab.ru
CasesPrices
Contacts
July 8, 2026Evgeny · Technical Director

Production monitoring: 4 metrics anyone on the team can understand


A client writes: "Payment doesn't work." You open the site — homepage loads fine. An hour later you learn the Pay button API is failing while marketing spend keeps running.

That's a typical production incident — your live site where people pay, not a dev sandbox. Below: four metrics any director or marketer can grasp, with the keywords teams actually search: production monitoring, server metrics, DevOps, Grafana.

Production monitoring dashboard: speed, errors, traffic, server load

Why "server pings" is not enough

Hosting may show 99.9% uptime while checkout returns errors, responses take 8 seconds, or the server is one spike from crashing.

Production monitoring tracks what customers experience plus headroom underneath — not 200 decorative charts.

4 metrics everyone should know

Four monitoring metrics: speed, traffic, errors, server capacity

1. Response speed

How long after a click. Watch worst cases, not only averages — if many users wait 2–3+ seconds, you lose orders.

2. Traffic and load

How many people use the site at once. A 3–5× spike in 30 minutes may be ads, viral traffic, or an attack.

3. Errors

Failed checkouts, 5xx pages, broken payments. Even 0.5% errors at scale means dozens of failures per minute.

4. Server headroom

CPU, memory, disk — like a fuel gauge. Alert at 70–80% sustained use, not at 99%.

Five questions for your IT team

  1. Who learns first if payment breaks — automatically, not from a tweet?
  2. What response time do we promise? (e.g. 99% of requests under 1s)
  3. How much downtime per month is acceptable? (99.9% ≈ 43 minutes/month)
  4. Who is on call on campaign day?
  5. Did we load-test before ads?

What should wake people at night

  • Urgent — site or payment down, call on-call
  • Important — slowdown or error spike, fix within an hour
  • Morning ticket — slow disk trend, no 3 AM call

Tools in plain terms

  • Grafana — vitals dashboard
  • Prometheus — collects server numbers (common open-source choice)
  • Telegram alerts — simple start for SMBs
  • Cloud APM (Datadog etc.) — faster setup, higher cost at scale

Bottom line

Production monitoring helps you learn before customers call. Four numbers — speed, traffic, errors, headroom — are enough for any executive. Usually part of DevOps services and pays back after one avoided outage. See downtime cost.

We can set up monitoring and test before your peak — DevOps, load testing, or performance audit.

Website monitoring FAQ

A homepage check misses a broken checkout or payment API. Monitor what customers actually do — order, login, pay.

Three alerts: site down, error spike, responses slower than 2–3 seconds. Telegram notifications are enough to start.

Not necessarily. Start with simple alerts and one clear dashboard. Add heavier tooling as traffic and team grow.

DevOps covers not only deployments but watching production: metrics, alerts, and capacity for traffic peaks.

Want to apply this in practice?

Tell us about your system — we’ll propose a work plan and the metrics worth fixing in an SLA/SLO.

All posts: DevOps & SRE

DevOps & SREAugust 27, 2026
SMB infra monitoring without DevOps: 5 signals that matter

Infrastructure monitoring without DevOps for SMBs: five signals — money-path uptime, disk, backups, slowdowns, and expiry dates. Telegram alerts, no Grafana.

Read Article
DevOps & SREJuly 27, 2026
Code from AI Agents: Checklist Before Production

AI agents write code: a pre-production control checklist for CTOs — leak, license, and silent-bug risks, review gates, and incident cost ranges in ₽.

Read Article
DevOps & SREJune 19, 2026
DevOps and CI/CD in Production: What to Set Up First

DevOps services for business: build pipeline, staging, zero-downtime deploy, monitoring and rollback — priorities for the first 4–6 weeks.

Read Article
DevOps & SREJune 19, 2026
Kubernetes in Production: A CTO Checklist Before Launching a Cluster

Production Kubernetes setup: RBAC, resources, Ingress, GitOps, monitoring, and common mistakes — a checklist before going live.

Read Article