Skip to main content

Health Checks

Runix can monitor process health using HTTP, TCP, or command-based checks. When a process is marked unhealthy after consecutive failures, Runix can trigger an automatic restart.

Implementation​

Located at internal/healthcheck/:

FileResponsibility
checker.goChecker — periodic health check loop with retry logic
http.goHTTP check — GET request, status code validation
tcp.goTCP check — dial and close connection
command.goCommand check — sh -c execution, exit code validation

Check Types​

HTTP​

health_check:
type: http
url: http://localhost:8080/health
interval: 10s
timeout: 5s
retries: 3
grace_period: 5s

Sends a GET request to the URL. A response with status 200–399 is healthy. Any other status or connection error is unhealthy.

TCP​

health_check:
type: tcp
tcp_endpoint: localhost:5432
interval: 5s
retries: 5

Attempts to establish a TCP connection. Successful connection = healthy. Connection refused or timeout = unhealthy.

Command​

health_check:
type: command
command: "curl -sf http://localhost:8080/health || exit 1"
interval: 10s
timeout: 5s
retries: 3

Executes a shell command via sh -c. Exit code 0 = healthy. Any other exit code = unhealthy.

Configuration​

FieldTypeDefaultDescription
typestring(required)http, tcp, or command
urlstringURL for HTTP checks
tcp_endpointstringhost:port for TCP checks
commandstringShell command for command checks
intervalduration10sTime between checks
timeoutduration5sPer-check timeout
retriesint3Consecutive failures before marking unhealthy
grace_periodduration0sDelay before starting checks

Check Lifecycle​

Grace Period​

The grace_period gives the process time to initialize before health checks begin. During this period, the process is assumed healthy.

health_check:
type: http
url: http://localhost:8080/health
grace_period: 10s # Wait 10s after start before checking

Retry Logic​

A process is only marked unhealthy after retries consecutive failures:

health_check:
retries: 3 # Must fail 3 times in a row

Any successful check resets the failure counter to 0.

Unhealthy Callback​

When a process is marked unhealthy, the onUnhealthy callback is triggered. In the supervisor, this typically initiates a restart (respecting the restart policy and max_restarts).

What's Next​