Live
UpOwl
A monitoring SaaS that checks endpoints from multiple regions, watches cron jobs via heartbeats, and only pages you when an outage is real.
- Role
- Founder · sole engineer
- Timeline
- [n] months
- Year
- [2025]
- Stack
- Next.js, Node.js, PostgreSQL
(01)The problem
Single-region monitors raise false alarms whenever one network path hiccups, so teams learn to ignore alerts — the worst outcome for a monitoring tool.
Cron jobs fail silently. Most uptime tools only watch URLs and never notice the nightly backup that stopped running three weeks ago.
(02)My role
- Designed and built the product end-to-end: data model, probe network, alerting pipeline, dashboard and billing.
- Ran it on self-managed VPS infrastructure with Docker and a zero-downtime CI/CD pipeline.
- [Anything else — design, pricing, user interviews, launch]
(03)Architecture
Lightweight probe workers in several regions run checks on a schedule and report to a central API. An alert is raised only when a quorum of regions agree, which removes nearly all false positives. Cron jobs ping a unique heartbeat URL; a late heartbeat opens an incident.
- 01
Scheduler
Distributes checks across regions with jitter
- 02
Regional probes
HTTP / TCP / keyword checks, TLS expiry
- 03
Ingest API
Writes results, rolls up time-series
- 04
Quorum engine
≥ 2 regions must agree before an incident
- 05
Notifier
Email, Slack, webhooks with escalation
Quorum over single-probe alerts
Cross-region confirmation trades a few seconds of detection latency for alerts people actually trust.
Redis queues, Postgres truth
Hot scheduling state lives in Redis; durable incident history and roll-ups live in Postgres with time-bucketed tables.
Heartbeats for cron jobs
Inverting the check — the job calls us — makes cron monitoring work behind firewalls with zero agents to install.
(04)Tech
App
- Next.js
- TypeScript
- Tailwind CSS
Backend
- Node.js
- PostgreSQL
- Redis
- BullMQ
Infra
- Docker
- Self-managed VPS
- GitHub Actions
- Caddy
(05)Outcome
[One or two sentences on traction, reliability, or what you learned.]
- probe regions
- 5+
- second check interval
- 30s
- platform uptime
- 99.9%
Next case study
AI Debate Arena
Multi-agent streaming debates — Pro, Con & Judge