Health Checks
Runix can monitor process health using HTTP, TCP, or command-based checks. When a process is marked unhealthy after consecutive failures, Runix can trigger an automatic restart.
Implementation
Located at internal/healthcheck/:
| File | Responsibility |
|---|---|
checker.go | Checker — periodic health check loop with retry logic |
http.go | HTTP check — GET request, status code validation |
tcp.go | TCP check — dial and close connection |
command.go | Command check — sh -c execution, exit code validation |
Check Types
HTTP
health_check:
type: http
url: http://localhost:8080/health
interval: 10s
timeout: 5s
retries: 3
grace_period: 5s
Sends a GET request to the URL. A response with status 200–399 is healthy. Any other status or connection error is unhealthy.
TCP
health_check:
type: tcp
tcp_endpoint: localhost:5432
interval: 5s
retries: 5
Attempts to establish a TCP connection. Successful connection = healthy. Connection refused or timeout = unhealthy.
Command
health_check:
type: command
command: "curl -sf http://localhost:8080/health || exit 1"
interval: 10s
timeout: 5s
retries: 3
Executes a shell command via sh -c. Exit code 0 = healthy. Any other exit code = unhealthy.
Configuration
| Field | Type | Default | Description |
|---|---|---|---|
type | string | (required) | http, tcp, or command |
url | string | URL for HTTP checks | |
tcp_endpoint | string | host:port for TCP checks | |
command | string | Shell command for command checks | |
interval | duration | 10s | Time between checks |
timeout | duration | 5s | Per-check timeout |
retries | int | 3 | Consecutive failures before marking unhealthy |
grace_period | duration | 0s | Delay before starting checks |
Check Lifecycle
Grace Period
The grace_period gives the process time to initialize before health checks begin. During this period, the process is assumed healthy.
health_check:
type: http
url: http://localhost:8080/health
grace_period: 10s # Wait 10s after start before checking
Retry Logic
A process is only marked unhealthy after retries consecutive failures:
health_check:
retries: 3 # Must fail 3 times in a row
Any successful check resets the failure counter to 0.
Unhealthy Callback
When a process is marked unhealthy, the onUnhealthy callback is triggered. In the supervisor, this typically initiates a restart (respecting the restart policy and max_restarts).
What's Next
- Restart Policies — How unhealthy processes get restarted
- Ready Command — Wait for process readiness
- Configuration Reference — Health check config fields