Skip to main content

Supervisor

The supervisor is the core engine of Runix. It manages the full lifecycle of processes — from creation through execution, monitoring, restart, and shutdown.

Located at internal/supervisor/, the supervisor consists of these files:

FileResponsibility
supervisor.goSupervisor struct — the top-level orchestrator
process.goProcess struct — wraps an exec.Cmd with state tracking, hooks, and log capture
state.goState machine logic, valid transition checking
backoff.goExponential backoff calculator for restart delays
monitor.goExit monitor — goroutines that wait on process exit and trigger restart
dependencies.goTopological sort for process dependency ordering
rolling.goRolling reload logic — batched concurrent restart with rollback

Supervisor Struct​

type Supervisor struct {
processes map[string]*Process // ID → Process
config *types.RunixConfig
options Options
// ...
}

type Options struct {
LogDir string // Base log directory
Defaults types.DefaultsConfig // Default config values
}

Key Methods​

MethodDescription
New(opts Options) *SupervisorCreates a new supervisor
AddProcess(ctx, cfg) (*Process, error)Adds and starts a process
StopProcess(id, force, timeout) errorStops a process
RestartProcess(ctx, id) errorStops then starts a process
ReloadProcess(ctx, id) errorGraceful restart preserving config
RemoveProcess(id) errorStops and removes from process table
Get(target) (*Process, error)Lookup by ID, name, or unique prefix
List() []ProcessInfoReturns info for all processes
RollingReload(ctx, names, opts) errorBatched concurrent reload
Save() errorPersist process list to dump.json
Resurrect() errorRestore processes from dump.json
Shutdown()Stop all processes and clean up
LogPath(name) stringReturns stdout log path for a process
LogPathStderr(name) stringReturns stderr log path for a process

Process Lifecycle​

Exit Monitor​

Each running process has a dedicated goroutine that waits on cmd.Wait(). When the process exits:

  1. The goroutine reads the exit code
  2. It calls handleExit() which atomically transitions the state
  3. If the restart policy allows, it schedules a restart with exponential backoff
  4. The exited channel is closed exactly once — other goroutines use this to detect termination
func (p *Process) watchExit() {
err := p.cmd.Wait()
// Extract exit code
p.handleExit(exitCode)
}

The exited channel is created fresh on each Start() call and closed exactly once by handleExit(). This prevents double-close panics.

Process Group Isolation​

All child processes are started with SysProcAttr{Setpgid: true}, which places them in their own process group. This means:

  • Signals are sent to the entire group via syscall.Kill(-pid, signal)
  • Child processes spawned by the managed process are also terminated
  • Prevents zombie processes from orphaned children

Log Capture​

Each process has stdout and stderr captured to separate log files:

~/.runix/apps/<name>/stdout.log
~/.runix/apps/<name>/stderr.log

Log files are written through a PrefixWriter that prepends timestamps:

2025-01-15 10:30:00 [out] Server listening on :8080
2025-01-15 10:30:01 [err] Connection refused

Log rotation is handled by internal/logrot/rotator.go when configured.

Dependency Resolution​

When starting multiple processes (e.g., from a config file), the supervisor resolves dependencies using topological sort:

  1. Processes with depends_on: ["db", "cache"] wait for those processes to be running first
  2. The sort uses priority as a tiebreaker — higher priority processes start first
  3. Circular dependencies are detected and rejected at config validation time

See Dependencies for the full feature documentation.

Rolling Reload​

The supervisor supports rolling reload for zero-downtime updates:

type RollingReloadOptions struct {
BatchSize int // Concurrent reloads per batch
WaitReady bool // Wait for health checks between batches
ReadyTimeout string // Timeout for health check readiness
RollbackOnFailure bool // Stop and revert on first failure
}

See Rolling Reload for the full feature documentation.

Thread Safety​

The supervisor uses fine-grained locking:

  • Process state: Lock-free via atomic.Value + CompareAndSwap
  • Process map: sync.RWMutex for adding/removing processes
  • Log writers: Mutex-protected PrefixWriter
  • Metrics collector: sync.RWMutex for the metrics map

The hot path — state transitions — never acquires a mutex. This is critical because state transitions happen on every process exit and health check cycle.

What's Next​