Skip to content

Live

UpOwl

A monitoring SaaS that checks endpoints from multiple regions, watches cron jobs via heartbeats, and only pages you when an outage is real.

Role
Founder · sole engineer
Timeline
[n] months
Year
[2025]
Stack
Next.js, Node.js, PostgreSQL

(01)The problem

Single-region monitors raise false alarms whenever one network path hiccups, so teams learn to ignore alerts — the worst outcome for a monitoring tool.

Cron jobs fail silently. Most uptime tools only watch URLs and never notice the nightly backup that stopped running three weeks ago.

(02)My role

  • Designed and built the product end-to-end: data model, probe network, alerting pipeline, dashboard and billing.
  • Ran it on self-managed VPS infrastructure with Docker and a zero-downtime CI/CD pipeline.
  • [Anything else — design, pricing, user interviews, launch]

(03)Architecture

Lightweight probe workers in several regions run checks on a schedule and report to a central API. An alert is raised only when a quorum of regions agree, which removes nearly all false positives. Cron jobs ping a unique heartbeat URL; a late heartbeat opens an incident.

  1. 01

    Scheduler

    Distributes checks across regions with jitter

  2. 02

    Regional probes

    HTTP / TCP / keyword checks, TLS expiry

  3. 03

    Ingest API

    Writes results, rolls up time-series

  4. 04

    Quorum engine

    ≥ 2 regions must agree before an incident

  5. 05

    Notifier

    Email, Slack, webhooks with escalation

Quorum over single-probe alerts

Cross-region confirmation trades a few seconds of detection latency for alerts people actually trust.

Redis queues, Postgres truth

Hot scheduling state lives in Redis; durable incident history and roll-ups live in Postgres with time-bucketed tables.

Heartbeats for cron jobs

Inverting the check — the job calls us — makes cron monitoring work behind firewalls with zero agents to install.

(04)Tech

App

  • Next.js
  • TypeScript
  • Tailwind CSS

Backend

  • Node.js
  • PostgreSQL
  • Redis
  • BullMQ

Infra

  • Docker
  • Self-managed VPS
  • GitHub Actions
  • Caddy

(05)Outcome

[One or two sentences on traction, reliability, or what you learned.]

probe regions
5+
second check interval
30s
platform uptime
99.9%

Next case study

AI Debate Arena

Multi-agent streaming debates — Pro, Con & Judge