Local AI tooling · Rust
llamactl
A deliberately small control plane for local model runtimes: keeping process lifecycle, resource signals, health, and access control in one inspectable operational surface.
The project
Local models still need a proper operator.
llamactl currently acts as a control plane. It loads model profiles from TOML, authenticates callers, launches configured backend processes on allocated loopback ports, tracks model instances, reports health and hardware observations, and exposes status, diagnostics, and lifecycle routes over a local HTTP API.
That is useful infrastructure for consuming applications such as Dungeon Dice, but it is important not to oversell the boundary. Inference requests currently go directly to the managed backend port returned by launch; llamactl does not yet proxy inference or provide a stable inference-facing API.
- Rust process and model lifecycle supervision
- Profile-driven backend launch arguments and loopback binding
- Health plus CPU, memory, and best-effort NVIDIA observations
- Model registry with PID, port, uptime, reaping, and unload policy
- Bearer-authenticated API with per-client permissions
- Generated local configuration, policy, and model profiles
Early development at v0.1.0. Control and supervision are implemented; backend readiness probes, inference proxying, request-aware activity, deeper hardware admission, automatic tuning, packaging, and a stable consumer API remain planned.