Huxகn gives you auth, rate limits, budgets, and full observability for every model running on your hardware — while every token stays inside your VPC. Hux is the hook that lives next to your GPU. Kan is the eye that watches every Hux from the cloud.
Stop re-implementing auth, quotas, and observability for each LLM you bring up. Drop a single agent on your GPU node — get a production gateway in three minutes.
Every prompt and every completion runs on the GPU you own. Kan never sees model inputs or outputs — only metered events: counts, latencies, error codes. Sovereign by default, no private deployment tier required.
Hux is one Go binary. Paste a curl into your VM and it brings up APISIX, Redis, Prometheus, Grafana, Loki, auto-discovers your vLLM containers, enrolls with Kan, and starts serving traffic. No Helm, no CRDs, no Kubernetes required.
Bind API keys to consumer groups, set RPM/RPD/token budgets per period, restrict by IP/model/route/time-window, and watch real-time usage in the Kan dashboard. Every limit enforced at the edge by APISIX, all metered events streamed back.
Whether you have one GPU on a desk or a thousand across a multi-region cluster, the path is the same.
Sign up with email, verify, and you're in the Kan console. Free 14-day trial. No credit card up front.
# From your browser:
https://app.huxkan.com/signup
A fleet is a logical pool of GPU machines that share the same gateway namespace. Name it whatever you like — "prod-us-east", "research-h100", "customer-acme" — and Kan generates a one-line installer.
SSH into the GPU VM, paste the installer. Hux brings up APISIX (gateway), Redis
(auth + budget cache), Prometheus + Grafana (metrics), Loki + Promtail (logs),
and DCGM-exporter (GPU telemetry) — all as Docker containers under a single
kan-agent systemd unit.
# Pasted from the Kan console once your fleet exists:
curl -sSL https://install.huxkan.com | FLEET_TOKEN=FLEET_ETK_… sh
Tell Kan how end users will reach this fleet (an IP, a hostname, or both — for the SNI list and the cert SAN). Then run "Discover & Adopt" — Hux scans the host for vLLM/SGLang containers and registers them as routable models. Or deploy new ones from the Models page.
Issue keys per end user or app. Group them into consumer groups and assign rate limits — requests/minute, requests/day, token budgets — and ACL rules like model allowlists, IP allowlists, time-of-day windows, expiry dates. Everything sync's into the fleet's local Redis within 15 seconds.
Your end users hit the OpenAI-compatible endpoint with their key. APISIX authenticates against the local cache, enforces ACLs and budgets, dispatches to the right model, and streams the completion back. Every request becomes a usage event you can see live in the Kan dashboard.
curl https://gateway.your-company.com/v1/chat/completions \
-H "Authorization: Bearer sk_…" \
-H "Content-Type: application/json" \
-d '{"model":"llama-3.2-1b","messages":[{"role":"user","content":"hi"}]}'
Hux is a Go agent. It sits inside your GPU VM next to vLLM/SGLang/TRT-LLM and serves as the gateway, observability stack, and config receiver. Kan is a stateless control plane in our cloud. It renders gateway configs, distributes API keys, ingests usage events, and shows you everything in a single dashboard.
This is what your Kan console looks like once a Hux is enrolled and serving traffic. Live request rates, p95 latency, hourly spend, and a streaming feed of every inference call as it lands.
Numbers shown are illustrative. Your real dashboard updates from your fleet's Prometheus + Loki — Kan never sees prompt or completion content.
Your tokens never cross our wire. Your keys never leave your fleet's Redis. Your model weights never see our servers. We see metering, you keep everything else.