How do you know your site is down before a customer emails support? Your load balancer can show healthy pods while the homepage returns 502. DNS can point to the wrong IP after a migration. A certificate can expire on Saturday night while
/healthinside the VPC still returns 200.
Those failures are what uptime monitoring solves. Scheduled probes from outside your network ask a simple question on repeat: can this URL, port, or hostname be reached, and does the response match what you expect?
This guide covers how to monitor website uptime with simple checks (HTTP, TCP, DNS, TLS, heartbeat), what to measure, how often to probe, how to wire alerts, and how to compare tools without overcomplicating the stack. Related: DNS monitoring, SSL certificate monitoring, monitoring and logging.
What is uptime monitoring?
Uptime monitoring is the practice of probing a service on a schedule to confirm it is available and behaving correctly. It does not rely on your app saying "I am fine" in your host dashboard, container logs, or docker ps list. It relies on external probes from outside your infrastructure: can a client reach the URL, resolve the hostname, complete TLS, and get the status code or body you expect?
Most teams start with a handful of HTTP checks on their homepage, /health endpoint, login route, and API base URL, then add TCP, DNS, ICMP, or heartbeat monitors as the stack grows. See how uptime monitoring works for the underlying loop.
This is not application performance management. APM tells you why a request was slow inside your code. Uptime monitoring tells you the service is unreachable or returning the wrong response. You need both, but the detection layer comes first, because a slow trace is useless when the load balancer returns 502. For the split between probes and logs, see monitoring and logging.
Why uptime monitoring matters
When a web application goes down, users notice before the team does. Unless you run external probes, your dashboard can show green while customers see errors. Internal health checks often skip the same path users take: CDN, DNS, TLS termination, WAF rules. Probes on the public URL close that gap in a way that /health inside the VPC never will.
The expensive part of an outage is not just the broken flow. It is the support loop that follows. Three common scenarios show what this looks like in practice.
Three examples from support inboxes
The site is down
User — Your site won't load. Just a blank white page. User — Login page too. Tried my phone, same thing. Was signing up for a trial. User — Your status page says everything is fine. That can't be right.
Engineering —
www502 from us-east and eu-west. Deployv2.4.1just landed? Engineering — Origin healthy inside VPC. External only. Rollback, who has the key? Engineering — Status page still green. Nobody updated it.
A plain HTTP check from outside your network would have caught this before support started collecting screenshots. The same approach works for any endpoint where you need a baseline alert before users open tickets: marketing pages returning 5xx, API health endpoints that stop responding after a deploy, or DNS and TLS failures that make the site unreachable entirely. See DNS monitoring and SSL certificate monitoring.
SSL certificate expired
User — Chrome says your connection is not private. I can't open the app at all. User — Same on my phone. The main website loads, the app doesn't. User — We have a customer demo in 2 hours. Any update?
Engineering — Cert on
app.example.comlapsed at midnight. Nobody got paged. Engineering — ACME failed,_acme-challenge.appstill points at old CNAME. Who edited DNS?
The app process was still running, but users never reached it because TLS was invalid. An HTTP check with an SSL expiry assertion catches renewal failures before browsers show the warning. This matters for certificate expiry on apex, www, and API subdomains, missed auto-renewal after DNS or ACME config changes, and staging certs that accidentally serve production hostnames. More detail in SSL certificate monitoring.
DNS points to the wrong place
User — Your app hasn't worked all morning from my home internet. User — Works fine when I'm on office VPN. Mobile app just says network error.
Engineering — External
digstill returns decommissioned LB IP. Cutover was yesterday. Engineering — Zone file fixed in staging. Never published to registrar. Engineering — Support queue at 40+. We don't have a runbook for this.
HTTP checks against the old IP kept passing while new users resolved to a dead address. DNS monitors verify resolution and record values independently of your app tier, catching problems after load balancer or CDN migrations, accidental TTL or registrar changes, and split-horizon DNS where external resolution differs from internal. See DNS monitoring and best API monitoring tools.
These three failure modes (hard outage, bad TLS, broken DNS) cover most of what users report first. Stack HTTP monitors on critical URLs, DNS monitors on hostnames, and SSL expiry assertions on certificates. Running probes from multiple regions is what turns "works on my machine" into an SLA you can actually defend. For the relationship between probes and SLAs, see SLO vs SLA vs SLI and best status page software.
The cost of undetected downtime
Downtime costs more than lost minutes on an availability chart. A trial user who hits a certificate error and leaves is not coming back. A customer who sees the same outage twice stops treating it as a bug and starts treating it as a reason to evaluate competitors. The support team spends the next hour confirming DNS screenshots for an incident that automated probes should have caught in seconds.
| Impact | What it costs |
|---|---|
| Trials | A homepage or signup URL that returns 5xx or fails TLS can lose a new user before they ever see the product. |
| Existing customers | Repeated incidents erode trust regardless of root cause. |
| Support cost | Every undetected incident becomes manual triage: screenshots, browser versions, call recordings, request IDs, and follow-up emails. |
| Trust | When users report an outage before your team sees it, the product feels unmanaged. |
Contracts reinforce this. If you promise 99.9% availability, the SLI has to come from external checks that run whether or not your agents are healthy. Dependency outages look like your bug: when Stripe or Cloudflare wobbles, your error rate spikes. Correlating vendor status with your own failing checks cuts MTTR. See MTTA, MTTR, MTBF, and MTTF compared for how these incident metrics relate.
Get posts like this in your inbox
Engineering insights on monitoring, incidents, and uptime. Bi-weekly, 2-minute read.
How uptime monitoring works
Every check follows the same loop: schedule, probe, assert, store, alert. A scheduler queues work per monitor and region. A probe worker issues the request. Assertions pass or fail. Results land in a time-series store. Alert rules fire when failures cross your threshold.
| Check type | What it verifies | Typical target |
|---|---|---|
| HTTP/HTTPS | Status code, headers, body, TLS | GET /health |
| TCP | Port accepts connections | db.internal:5432 |
| DNS | Resolution and record values | api.example.com |
| ICMP | Host reachable (where allowed) | Edge router |
| Heartbeat | Job ran on schedule | Cron ping URL |
Run checks from more than one region. A record that resolves in Virginia but fails in Frankfurt is a production incident for half your users. One green datapoint is not global health, which is why you need multi-region probes rather than a single ping from one datacenter.
Deep dives: DNS monitoring · SSL certificate monitoring · multi-region monitoring guide
Key metrics to track
Uptime data only helps if you read it consistently. Five numbers cover most teams before you need custom dashboards.
Availability (uptime %). Count failed checks against total checks in a rolling window, usually 30 or 90 days. Exclude planned maintenance windows so deploys do not tank your SLA report. External probes are a cleaner measure of "up for users" than internal health endpoints that stay green while the CDN fails. This is the core input for SLIs and SLOs.
Response time. Track p50 and p95, not just the average. A mean of 200ms can hide spikes that make your API feel unusable. Latency thresholds on health checks catch slow databases and cold starts before they become hard outages. Look for tools that report percentiles per region. See best API monitoring tools and response time budgets.
Error rate. Failed checks divided by total checks in the same window as availability. A monitor that flaps between pass and fail is often worse than one that stays red, because it trains on-call to ignore pages. Group by monitor and region so you see whether the problem is global or isolated to one probe location.
Time to detect. Elapsed time from the first failed external probe to the alert landing in Slack or PagerDuty. This is the part uptime monitoring controls directly. Shaving detection from five minutes to thirty seconds is often worth more than shaving fix time, because nobody starts debugging until the page fires.
Mean time to repair (MTTR). Detection is the probe's job; repair is the incident workflow. Still, track MTTR alongside probe metrics so you know whether alerts are actionable. If MTTR stays high while detection is fast, the gap is runbooks, ownership, or dependency context, not check frequency. For how these metrics relate, see MTTA, MTTR, MTBF, and MTTF compared.
Uptime probes measure whether the service works for users. Logs and traces explain why it broke. Keep that split in mind. Monitoring and logging answer different questions, and you need both after the alert fires.
What to monitor
Start with URLs and hostnames that block revenue or support tickets, then widen to infrastructure those endpoints depend on. The goal is not "every URL on the sitemap." It is the smallest set of probes that prove the service is reachable and responding correctly.
Customer-facing HTTP endpoints. Homepage, app login page, signup API, and your public /health or readiness route. Run HTTP checks with status-code assertions on marketing pages and stricter body or JSON assertions on API routes. Cover the endpoints your mobile app and partners call, not only the marketing site. See best API monitoring tools and API testing vs API monitoring.
Auth, payments, and webhooks. Monitor login and token refresh URLs, payment webhooks, and charge endpoints. Assert 200 (or expected 401 on unauthenticated health) and response time ceilings. Webhooks fail silently when TLS or DNS drifts, so a dedicated monitor on the public webhook URL catches misconfiguration before events queue on the vendor side.
DNS and TLS. Resolution for apex and www, API subdomains, and mail records if deliverability matters. Certificate expiry on every hostname users hit in the browser. These are cheap checks that prevent the "site is down but the app pod is fine" class of incidents. Deep guides: DNS monitoring and SSL certificate monitoring.
Background jobs and cron. Heartbeat monitors on backup jobs, report generators, and queue drainers. If a job stops pinging, you want a page before downstream data is stale, not when a customer notices missing invoices.
TCP and ICMP where HTTP is not enough. Database ports, message brokers, and edge routers often need a TCP connect check or ping when there is no HTTP surface. Keep these on longer intervals unless they are customer-critical paths.
Third-party dependencies. When Stripe, Cloudflare, or your auth provider wobbles, your error rate spikes even if your deploy was clean. Correlate vendor status with your own failing checks so the first responder knows whether to roll back or wait.
What not to over-monitor. Static docs, blog posts, and legal pages rarely need 30-second multi-region checks. A single HTTP monitor on status code every 5 to 15 minutes is enough. Save aggressive frequency and region coverage for homepage, API health, and payment webhooks.
Pre-deploy API tests and production monitors solve different problems: tests gate releases, monitors watch what already shipped. If your team conflates the two, read API testing vs API monitoring before you duplicate coverage or leave gaps.
Check frequency and regions
Frequency and region choice determine how fast you learn about an outage and how often you get false pages. Faster checks and more regions mean quicker detection, but also higher probe cost. The goal is matching interval to user impact.
How often to check. Payment and login paths deserve 30 to 60 second checks. Marketing pages and internal dashboards can run every 5 to 15 minutes. On DevHelm, the create-monitor flow exposes an interval slider from 10 seconds to 24 hours, with faster intervals unlocking on higher plan tiers (Free starts at 5 minutes, Starter at 1 minute, Pro at 30 seconds, Enterprise at 10 seconds). The gap matters: a 5-minute check can miss a four-minute blip entirely, while a 60-second check catches it within one or two failed probes. Compare tiers on the uptime monitoring product page.
Where to probe from. Pick regions that match real users. US East for a US-heavy SaaS, EU West for GDPR customers, AP South if you have APAC traffic. DevHelm ships four probe regions today (US East, US West, EU West, and AP South), selectable on the Schedule step of the create-monitor form.
Multi-region confirmation. One green probe is not global health. Require failures from more than one region before opening an incident, or you will page on transient network paths between a single probe and your CDN. DevHelm supports multi-region confirmation policies so a flaky link in one geography does not wake on-call at 3 a.m. See the multi-region monitoring guide for confirmation and recovery settings.
Practical defaults. Production API health: 60 seconds, three regions, multi-region confirm. Static status page: 5 minutes, one region. Certificate and DNS monitors: daily or hourly is often enough because expiry assertions matter more than sub-minute polling. New to the dashboard? The first HTTP monitor guide walks through interval, regions, and assertions on the same create flow.
Alerting that gets answered
A monitor that fails silently is worse than no monitor. Uptime monitoring only pays off when alerting turns a failed probe into a page someone actually answers.
Related guides: first HTTP monitor · first alert · testing your alerts · incident management tools comparison
DevHelm splits alerting into two layers you configure once and reuse everywhere. Alert channels are the destinations: Slack, PagerDuty, OpsGenie, Discord, Microsoft Teams, email, webhooks, and more. You create them in the dashboard or CLI, test with devhelm alert-channels test, and attach them to monitors. The full list is in the integrations overview. Notification policies are the routing rules and escalation chains that decide which channels fire for which incidents, based on severity, tags, and catch-all defaults. The alerting guide walks through the setup.
When you create a monitor, the form walks four steps: Monitor (type and target), Schedule (interval and regions), Alerting (pick channels), and Advanced (assertions, auth headers, incident policy). On the Alerting step, select existing channels or add Slack, PagerDuty, or email without leaving the form. Use Test to run a live probe against your target before you save.
Most teams start with a Slack incoming webhook because it requires no OAuth app. Paste the webhook URL, name the channel, test it, then select it on the monitor. Incident messages include severity, monitor name, affected regions, and a link back to the dashboard. Setup steps are in the Slack integration docs. For the full channel catalog (PagerDuty trigger-resolve vs Slack fire-and-forget), see the integrations overview.
A single Slack channel works for a five-person team. Larger organizations need tiered chains: Slack first, PagerDuty after N minutes, email to leads if still unacknowledged. DevHelm notification policies support delay steps and match rules, including routing by monitor tags. Patterns are documented in the alerting guide, tiered escalation, and alert routing by tag.
A failing monitor only helps if the notification reaches the right channel on the first attempt. Attach tested alert channels on each monitor, route warnings and outages to different destinations, and use maintenance windows during planned deploys so expected blips do not open incidents. Include a runbook or dashboard link in the alert payload so whoever receives the page knows what to do next. Before you depend on it in production, walk the full path from failed check to delivered notification with testing your alerts.
Status pages and dependency correlation
Internal alerts tell your team something broke. Customers still email support unless they can see the same truth your probes see. Status pages and dependency context are how you communicate outages and cut the time spent answering "is it just me?"
Status pages tied to monitor state
When an incident opens, your public status page should update from the same failed checks that triggered the alert, not from someone remembering to flip a component to red during a firefight. Map homepage, API, login, and payment monitors to status page components so degradation and recovery follow probe results automatically.
Manual status pages fail in the exact moments you need them. The on-call engineer is debugging, support is drowning in tickets, and the marketing site still shows "all systems operational." Automated pages shrink support load and set expectations while you fix the root cause. For how teams compare vendors on this feature, see best status page software.
When the outage is not yours
Not every red monitor means a bad deploy. Payment webhooks fail when Stripe has an incident. Login breaks when your identity provider wobbles. API error rates spike when Cloudflare or your CDN misroutes traffic. If you only look at your own graphs, the first hour goes to proving it is not your bug.
Track the third-party services your uptime checks depend on and surface vendor status next to your failing monitors. When a checkout URL and Stripe both go red at the same time, you skip the rollback and post a status update instead. DevHelm documents this workflow in tracking dependencies and the broader status data guide.
Wire both pieces together: probes detect the problem, alerts wake the right people, the status page tells customers, and dependency context tells your team whether to fix or wait. That full loop is what turns uptime monitoring from a dashboard widget into an incident workflow.
Assertions and advanced checks
A monitor that only checks "did we get a response?" is incomplete. You need pass/fail criteria on status codes, latency, body content, certificates, and DNS records. Assertions turn a ping into proof that the service works the way users expect.
Start with HTTP basics. Assert status code 200 (or 204/301 where appropriate) on homepage, /health, and login URLs. Add body_contains or json_path on API routes so a 200 with {"status":"error"} still fails. For API endpoints, match the JSON shape your app and partners rely on, not just that the socket connected.
Layer performance budgets. Response time assertions on the same probe catch slow endpoints before they time out. Use a warn threshold (for example 500ms) and a fail threshold (for example 2s) so on-call sees degradation early. See response time budgets for a full example.
SSL and DNS assertions. Add ssl_expiry with a 14-day floor on every public hostname. Pair DNS record assertions with HTTP checks on the same name so you know whether resolution or the origin failed. Guides: SSL certificate monitoring and DNS monitoring.
Stack assertions on one monitor. A production API health check might assert 200 status, JSON path $.status == "ok", response time under 2 seconds, and TLS valid for at least 7 days. Stacking multiple assertions on a single monitor keeps your configuration honest when the process is up but the response is wrong.
monitors: - name: API Health type: HTTP config: url: https://api.example.com/health method: GET frequencySeconds: 60 regions: [us-east, eu-west] assertions: - config: { type: status_code, expected: "200", operator: equals } severity: fail - config: { type: json_path, expression: "$.status", expected: "ok" } severity: fail - config: { type: response_time, thresholdMs: 2000 } severity: fail - config: { type: ssl_expiry, minDays: 14 } severity: fail
For assertion types per monitor (HTTP, DNS, TCP, heartbeat), see the monitors YAML reference. For API-specific patterns, the best API monitoring tools comparison covers what strong body and latency checks look like in practice.
Monitoring as code
Clicking monitors into a dashboard does not scale. When checkout URLs, alert channels, and regions live only in a UI, they drift the week after a deploy. Monitoring as code keeps monitors, channels, and policies in Git beside the services they watch, reviewed in PRs, deployed with the same discipline as application config.
A new /health route ships in a PR, and the monitor for it should land in the same PR. Staging and production stay aligned because each environment has its own devhelm.yml (or overlay), not a manual copy-paste. That distinction is what separates a reproducible monitoring setup from a bookmarked admin panel.
Define monitors and alert channels in YAML, validate offline, preview changes, then deploy:
devhelm validate devhelm.yml
devhelm plan -f devhelm.yml
devhelm deploy -f devhelm.yml --yes
The file can include tags, HTTP/DNS/TCP monitors, assertions, and Slack or PagerDuty channels with secrets referenced as ${SLACK_WEBHOOK_URL} (never committed in plain text). Walkthrough: monitoring as code guide. Full schema: monitoring as code overview.
Run devhelm deploy on merge when devhelm.yml changes, the same way you apply Terraform or Kubernetes manifests. GitHub Actions and other CI systems are documented under CI/CD pipeline and GitHub Actions integration.
If you are evaluating monitoring tools at scale, config-as-code is a hard requirement: CLI, validated YAML, plan/apply, and Terraform or API parity. The monitoring as code post explains why monitors rot in a UI and how a monitoring service should be reproducible from a repo clone.
How to evaluate uptime monitoring tools
There is no single best tool for every team. Use a rubric that matches how your product actually fails in production.
| Criterion | What to ask |
|---|---|
| Check types | HTTP, TCP, DNS, SSL expiry, heartbeat? |
| Regions | Enough probe locations for your users? |
| Assertions | Beyond status code: body, JSON, latency, TLS? |
| Alerting | Integrations you already use (Slack, PagerDuty, email)? |
| Pricing model | Per monitor, per check run, per seat? |
| Config-as-code | CLI, validated YAML, Terraform, API? |
| Status page | Included or a separate product? |
| Dependency context | Vendor outages shown next to your failing checks? |
Free ping tools work until you need multi-region confirmation, escalation chains, or monitors in Git beside your services. When you outgrow a bookmarked admin panel, compare on the criteria above rather than feature checklists alone. The website monitoring tools comparison walks through the rubric in detail. Migrating from a free-tier ping service? See DevHelm vs UptimeRobot.
How to monitor website uptime (step by step)
Uptime monitoring becomes repeatable once you treat it like any other production control. Here is the process from first probe to ongoing review.
- Pick endpoints that represent real user paths. Homepage, login, checkout or signup API, and a public
/healthroute. - Set frequency and regions. 30 to 60 seconds on revenue paths, 5 to 15 minutes on docs. Probe from regions where your customers live.
- Add assertions. Status code, JSON body shape, latency ceilings, SSL days-to-expiry, and DNS record values on the same hostnames.
- Wire alerts to on-call. Tested Slack or PagerDuty channels, escalation policies, and maintenance windows for planned deploys.
- Publish a status page. Map monitors to components so customers see the same truth your probes see.
- Store config in Git.
devhelm validate, plan, and deploy monitors with the services they watch. - Review failed checks weekly. Flapping monitors and ignored alerts are how teams learn to dismiss pages.
monitors: - name: Homepage type: HTTP config: url: https://www.example.com/ method: GET frequencySeconds: 60 regions: [us-east, eu-west, ap-south] assertions: - config: { type: status_code, expected: "200", operator: equals } severity: fail
Getting started with DevHelm
DevHelm is built for teams that need uptime monitoring alongside third-party dependency tracking. The free tier includes 50 monitors, multi-region probes, alert channels, and the CLI, with no credit card required.
Create your first HTTP monitor in the dashboard: pick a URL, set interval and regions, attach Slack or email, and add assertions on the Advanced step. The first HTTP monitor guide walks through the same flow in under ten minutes.
When Stripe, Cloudflare, or your auth provider wobbles, dependency correlation shows vendor status next to your failing checks so you know whether to roll back or wait. Learn more on the uptime monitoring product page.
Create your first monitor at app.devhelm.io. Free tier, 50 monitors, no credit card.
FAQ
What is uptime monitoring?
The practice of probing a service on a schedule from outside your infrastructure to confirm it is reachable and responding correctly. See What is uptime monitoring? for the full explanation.
How do you monitor website uptime?
Pick the URLs that represent real user paths, set check frequency and probe regions, add assertions on status codes and response time, wire alerts to Slack or PagerDuty, and store the config in Git. The full walkthrough is in How to monitor website uptime (step by step).
What should I look for in an uptime monitoring tool?
Check types, probe regions, assertion depth, alerting integrations, pricing model, config-as-code support, built-in status pages, and dependency correlation. See How to evaluate uptime monitoring tools and the website monitoring tools comparison.
Why does response time matter for uptime monitoring?
Because a service that responds but takes 10 seconds is effectively down for users. Tracking p50 and p95 response times, and setting latency assertions on HTTP monitors, catches slow databases and cold starts before they become hard outages. See Key metrics to track.
What is external monitoring and why does it matter?
External probes run outside your VPC on the same path users take through DNS, CDN, and TLS. Internal health checks often skip those layers, which is why your dashboard can show green while customers see errors. See Why uptime monitoring matters.
Get posts like this in your inbox
Engineering insights on monitoring, incidents, and uptime. Bi-weekly, 2-minute read.