Two halves, one gateway. Hux is the hook that lives next to your GPU
and serves all the traffic. Kan is the eye that watches every Hux
from the cloud, distributes config, and gives you a single dashboard for every fleet
you run.
Hux — the hook in your GPU.
"Hux" is hook. The agent is one Go binary that systemd runs on every GPU machine in
your fleet. On boot it brings up a stack of containers — APISIX as the gateway, Redis
for low-latency auth/budget cache, Prometheus + DCGM for metrics, Loki + Promtail for
logs, Grafana for dashboards — and stays running as the local control loop. Inference
workloads (vLLM, SGLang, TensorRT-LLM) run as siblings. Hux discovers them, registers
them with Kan, and keeps the gateway config synced.
Every request from your end users terminates at Hux's APISIX. APISIX
authenticates the API key against the local Redis cache, checks ACLs (IP / model /
route / time-window / expiry), checks per-key and per-group budgets, then proxies to
the right model upstream. The completion stream goes back to the user. Kan never sees
the prompt or the output.
Kan — the eye in the cloud.
"Kan" is eye in Tamil (கண்). It's the multi-tenant control plane that watches every
Hux you've enrolled. Four services in our cloud:
core-api — the source of truth for orgs, fleets, nodes, models,
API keys, consumer groups, budgets, and audit. Postgres-backed with full row-level
security per org. Exposes the REST API the console uses and the agent endpoints
Hux talks to.
config-sync — listens on Postgres NOTIFY for any config change
(a key revoked, a fleet renamed, a model deployed) and re-renders the affected
fleet's apisix.yaml manifest. Hux pulls it on the next 15-second
heartbeat.
quota-engine — ingests usage events Hux streams up, attributes
them to the right key/group/org, and pushes budget refreshes back down for the
gateway plugins to enforce.
console — the web app you log into. Manage everything from one
multi-tenant dashboard, even when your fleets span continents.
What crosses the wire.
Hux pulls from Kan, never the other way around. Even with that, the
only things on the wire are:
Config (Kan → Hux): the rendered apisix.yaml manifest and the keyauth-map delta. Both gzip-compressed, both signed by an org-scoped fleet token.
Usage events (Hux → Kan): request_id, timestamp, route, model name, status code, latencies, token counts, hashed client IP, error code. No prompts. No completions. No model weights. Every event is metering, not content.
How a single request flows.
End user POSTs to https://your-gateway/v1/chat/completions with their Authorization: Bearer sk_… header.
APISIX terminates TLS using the cert in your fleet's ssls: entry. SNI matches one of your operator-set public addresses (IP or hostname).
The kan-model-router plugin reads body.model, slugifies, and dispatches to the matching upstream that points at your vLLM.
kan-auth resolves the API key against the local Redis (populated via the keyauth-map sync), maps it to a consumer.
kan-acl checks IP / model / route / time-window / expiry restrictions for that consumer.
kan-budget checks per-key and per-group token/USD budgets for the current period.
APISIX proxies to vLLM. Completion streams back unchanged.
kan-quota's log phase emits a usage event, increments local Redis budget counters, and queues the event for the next batched POST to Kan's quota-engine.
Steady state is one tick every 15 s for heartbeats and config pulls (almost always
304s once everything's stable), plus a usage-event batch every few seconds when
there's traffic. The data path between your end user and your GPU never hops through
our cloud.
Why this shape.
Most "AI gateway" products force you to send prompts through their cloud — that's how
they implement auth and rate limits. Huxகn
inverts the problem: the gateway lives next to the GPU, the cloud only orchestrates.
You get a real production gateway (auth, ACLs, budgets, observability) while keeping
the regulatory and economic posture of a self-hosted system. No hidden egress, no
"trust us with your tokens", no surprise vendor data plane.