Product
Hux — the hook in your GPU.
Hux is one Go binary that turns any GPU machine into a production-ready AI gateway
node. It's the data plane: every inference request goes through it, every metric
leaves through it, every config change lands on it. Hux is short for "hook" —
it sits next to your model server and hooks every request before the GPU sees it.
What Hux runs on your VM.
The agent itself is one systemd unit (kan-agent). On boot it brings up
a fixed set of containers under --network host:
- APISIX 3.11 — the gateway. Listens on 9080 (plain), 9443 + 443 (TLS). Handles auth, rate limits, ACLs, budgets, model routing, prometheus scraping.
- Redis 7.2 — the auth + budget cache. The agent syncs API keys + ACLs + budgets into Redis HASHes; APISIX plugins read from there at request time.
- Prometheus + DCGM-exporter + node-exporter — metrics collection. APISIX request rates, GPU utilization, GPU temperature, host CPU/RAM, container counts.
- Loki + Promtail — log aggregation. Tails APISIX, vLLM, agent logs into a fleet-local Loki instance.
- Grafana — embedded dashboards, accessible from the Kan console via iframe proxy.
Your inference engines (vLLM, SGLang, TensorRT-LLM) run as siblings, also on host
networking. Hux discovers them by scanning for OpenAI-compatible /v1/models
endpoints and registers them with Kan automatically.
What Hux does at runtime.
- Pulls config from Kan every 15 seconds. If
apisix.yaml changed (new model, key revoked, group renamed), it writes the new manifest in-place and signals APISIX to reload.
- Streams keyauth deltas into Redis. When you create or revoke a key in the console, the agent's next sync writes the change into Redis via
redis-cli --pipe within seconds.
- Heartbeats node liveness — agent version, applied config revision, container statuses — so the Kan console always knows which Huxes are healthy.
- Reconciles model deployments for any model whose lifecycle Kan owns. Discovers user-spawned containers and adopts them safely. Skips ones that are already serving on a port.
- Generates and rotates a self-signed TLS cert at boot so end users can hit HTTPS even before you upload a real cert.
- Runs commands on demand — fetch logs, restart agent, run "Discover & Adopt", uninstall — issued from the console, executed locally.
What Hux never does.
- It does not call out to Kan with prompts or completions. Ever.
- It does not phone home with model weights or per-request token contents.
- It does not require Kubernetes, Helm, or any orchestration framework you don't already have.
- It does not need outbound access to anything except
app.huxkan.com and your model registries (Hugging Face, your private registry).
System requirements.
- Linux x86_64 (Ubuntu 22.04, Debian 12, RHEL 9, Amazon Linux 2023 — all tested).
- Docker Engine 24+ (the agent uses the docker CLI for container management).
- NVIDIA driver + nvidia-container-toolkit if you want DCGM GPU telemetry.
- Outbound HTTPS to
app.huxkan.com.
- Inbound TCP on whatever ports you expose to your end users (typically 443).
Deploy Hux now →
See the architecture