02 · Rust · Runtime control

A model server is still a process.

llamactl puts ownership, authentication, resource state, and cleanup around local model runtimes. The current boundary is a control plane, not yet an inference server.

The supervisor

Launch is an event, not a lifecycle.

Manually starting a model backend works until two clients want the GPU, a child crashes, a profile needs different flags, or nobody remembers which process owns the listening port. llamactl is a Tokio/Axum control plane over a model registry, loader, scheduler, hardware snapshot, security policy, and process supervisor.

Process ownership is explicit

Profiles define launch configuration. The supervisor assigns a local port, starts the configured backend without a shell, records its model id, PID, port, launch time, and policy, and retains ownership until the process stops. Background supervision reaps children that exit unexpectedly and unloads eligible instances after the configured timeout.

There is an important limitation: inference traffic does not yet pass through llamactl, so “idle” currently means time since launch rather than time since the last inference request. Backend readiness is also not established by a successful spawn response. The documentation calls both limits out instead of inventing a stronger lifecycle model than the code has.

Local callers still have different authority

Requests carry a client identity and, for registered identities, a bearer token. Only SHA-256 token digests are persisted. Exact permissions such as model read, load, unload, and admin are resolved centrally and enforced by each route. Unknown callers receive only the configured unprivileged default, while a registered caller with invalid credentials is rejected rather than silently downgraded.

The control API defaults to loopback, and spawned model backends are forced onto loopback independently of the control bind address. That keeps the intended boundary local, but it is not a substitute for transport encryption or host security. The current API should not be exposed directly to an untrusted network.

State is observable, not clairvoyant

Shared state is coordinated behind asynchronous locks. Hardware sampling reports CPU, memory, and best-effort NVIDIA information; it does not yet make reliable admission or placement decisions. Health, status, and doctor routes expose what the control plane observed, including the configured profiles and owned processes.

The intended consumer boundary is still ahead

Applications such as Dungeon Dice are intended to request local runtime capability instead of owning backend downloads, subprocess supervision, ports, and lifecycle policy themselves. That is the architectural direction. The current implementation remains a working control plane: an inference proxy, readiness states, request-aware activity, resource admission, automatic tuning, and a stable versioned consumer API are planned rather than shipped.

Current state: v0.1.0 early development. Control and lifecycle supervision are real; the complete local inference boundary is not.