Product

Hux — the hook in your GPU.

Hux is one Go binary that turns any GPU machine into a production-ready AI gateway node. It's the data plane: every inference request goes through it, every metric leaves through it, every config change lands on it. Hux is short for "hook" — it sits next to your model server and hooks every request before the GPU sees it.

What Hux runs on your VM.

The agent itself is one systemd unit (kan-agent). On boot it brings up a fixed set of containers under --network host:

Your inference engines (vLLM, SGLang, TensorRT-LLM) run as siblings, also on host networking. Hux discovers them by scanning for OpenAI-compatible /v1/models endpoints and registers them with Kan automatically.

What Hux does at runtime.

What Hux never does.

System requirements.


Deploy Hux now → See the architecture