Open source
Distributed Workflow Orchestrator
A durable execution engine: write workflows as plain code, and every step is persisted, retried and resumable — even across crashes and deploys.
- Role
- Architect & engineer
- Timeline
- [n] months
- Year
- [2025]
- Stack
- Node.js, TypeScript, PostgreSQL
(01)The problem
Long-running business processes — onboarding, billing, AI pipelines — fail halfway and leave state inconsistent.
Hosted durable-execution platforms are powerful but heavy and expensive to self-host for small teams.
(02)My role
- Designed the event-sourced execution model and step-replay semantics.
- Built the worker runtime, lease-based task claiming, retries with backoff, and timers.
- Built a dashboard to inspect runs, replay failures and cancel workflows.
(03)Architecture
Each workflow run is an append-only event history in Postgres. Workers claim tasks with leases; on resume, the workflow function is replayed against history so completed steps return their recorded results instead of running again.
- 01
Client SDK
Define workflows & steps as code
- 02
History store
Event-sourced runs in PostgreSQL
- 03
Task queue
Lease-based claiming, visibility timeouts
- 04
Workers
Deterministic replay + step execution
- 05
Timers
Durable sleeps, schedules, retries
Event sourcing for replay
Persisting every step result makes recovery a pure function of history — no bespoke checkpointing logic per workflow.
Leases, not locks
Time-bounded leases let a crashed worker's tasks be picked up automatically without a coordinator.
Postgres first
SKIP LOCKED queues and advisory locks were enough for the target scale, keeping the deployment a single database.
(04)Tech
Runtime
- Node.js
- TypeScript
Storage
- PostgreSQL
- Redis
Ops
- Docker
- OpenTelemetry
- GitHub Actions
(05)Outcome
[Throughput benchmarks, adopters, or lessons learned.]
- lost steps across crash tests
- 0
- steps / second per worker
- 1,000+
- dependency: Postgres
- 1
Next case study
Dynamic QR
QR codes whose destination changes on a schedule