Architecture

How Huxகn is built.

Two halves, one gateway. Hux is the hook that lives next to your GPU and serves all the traffic. Kan is the eye that watches every Hux from the cloud, distributes config, and gives you a single dashboard for every fleet you run.

END USER curl / OpenAI SDK HTTPS YOUR GPU VPC · HUX APISIX auth · ACL · budget · routing REDIS keyauth · acl · budget cache vLLM · SGLang · TensorRT-LLM your models, your GPU, your VRAM Prometheus · Loki metrics · logs Grafana · DCGM dashboards · GPU telemetry all running as Docker containers under one kan-agent systemd unit HUXKAN CLOUD · KAN core-api orgs · fleets · keys · budgets · audit config-sync renders apisix.yaml on every change quota-engine ingests usage · enforces budgets console (web UI) your dashboard · multi-tenant config pulls heartbeat · usage

Hux — the hook in your GPU.

"Hux" is hook. The agent is one Go binary that systemd runs on every GPU machine in your fleet. On boot it brings up a stack of containers — APISIX as the gateway, Redis for low-latency auth/budget cache, Prometheus + DCGM for metrics, Loki + Promtail for logs, Grafana for dashboards — and stays running as the local control loop. Inference workloads (vLLM, SGLang, TensorRT-LLM) run as siblings. Hux discovers them, registers them with Kan, and keeps the gateway config synced.

Every request from your end users terminates at Hux's APISIX. APISIX authenticates the API key against the local Redis cache, checks ACLs (IP / model / route / time-window / expiry), checks per-key and per-group budgets, then proxies to the right model upstream. The completion stream goes back to the user. Kan never sees the prompt or the output.

Kan — the eye in the cloud.

"Kan" is eye in Tamil (கண்). It's the multi-tenant control plane that watches every Hux you've enrolled. Four services in our cloud:

What crosses the wire.

Hux pulls from Kan, never the other way around. Even with that, the only things on the wire are:

How a single request flows.

  1. End user POSTs to https://your-gateway/v1/chat/completions with their Authorization: Bearer sk_… header.
  2. APISIX terminates TLS using the cert in your fleet's ssls: entry. SNI matches one of your operator-set public addresses (IP or hostname).
  3. The kan-model-router plugin reads body.model, slugifies, and dispatches to the matching upstream that points at your vLLM.
  4. kan-auth resolves the API key against the local Redis (populated via the keyauth-map sync), maps it to a consumer.
  5. kan-acl checks IP / model / route / time-window / expiry restrictions for that consumer.
  6. kan-budget checks per-key and per-group token/USD budgets for the current period.
  7. APISIX proxies to vLLM. Completion streams back unchanged.
  8. kan-quota's log phase emits a usage event, increments local Redis budget counters, and queues the event for the next batched POST to Kan's quota-engine.

Steady state is one tick every 15 s for heartbeats and config pulls (almost always 304s once everything's stable), plus a usage-event batch every few seconds when there's traffic. The data path between your end user and your GPU never hops through our cloud.

Why this shape.

Most "AI gateway" products force you to send prompts through their cloud — that's how they implement auth and rate limits. Huxகn inverts the problem: the gateway lives next to the GPU, the cloud only orchestrates. You get a real production gateway (auth, ACLs, budgets, observability) while keeping the regulatory and economic posture of a self-hosted system. No hidden egress, no "trust us with your tokens", no surprise vendor data plane.


Deploy your first Hux → Read the docs