Self-hosted on your GPU. Managed from our control plane.

The AI gateway
for sovereign GPU clouds.

Huxகn gives you auth, rate limits, budgets, and full observability for every model running on your hardware — while every token stays inside your VPC. Hux is the hook that lives next to your GPU. Kan is the eye that watches every Hux from the cloud.

Why Huxகn

One gateway, every model, your hardware.

Stop re-implementing auth, quotas, and observability for each LLM you bring up. Drop a single agent on your GPU node — get a production gateway in three minutes.

Tokens stay home.

Every prompt and every completion runs on the GPU you own. Kan never sees model inputs or outputs — only metered events: counts, latencies, error codes. Sovereign by default, no private deployment tier required.

One agent, three minutes.

Hux is one Go binary. Paste a curl into your VM and it brings up APISIX, Redis, Prometheus, Grafana, Loki, auto-discovers your vLLM containers, enrolls with Kan, and starts serving traffic. No Helm, no CRDs, no Kubernetes required.

Per-key everything.

Bind API keys to consumer groups, set RPM/RPD/token budgets per period, restrict by IP/model/route/time-window, and watch real-time usage in the Kan dashboard. Every limit enforced at the edge by APISIX, all metered events streamed back.

Get started

From signup to first inference in six steps.

Whether you have one GPU on a desk or a thousand across a multi-region cluster, the path is the same.

Create your Huxகn account.

Sign up with email, verify, and you're in the Kan console. Free 14-day trial. No credit card up front.

# From your browser:
https://app.huxkan.com/signup

Add a fleet.

A fleet is a logical pool of GPU machines that share the same gateway namespace. Name it whatever you like — "prod-us-east", "research-h100", "customer-acme" — and Kan generates a one-line installer.

Install Hux on your GPU VM.

SSH into the GPU VM, paste the installer. Hux brings up APISIX (gateway), Redis (auth + budget cache), Prometheus + Grafana (metrics), Loki + Promtail (logs), and DCGM-exporter (GPU telemetry) — all as Docker containers under a single kan-agent systemd unit.

# Pasted from the Kan console once your fleet exists:
curl -sSL https://install.huxkan.com | FLEET_TOKEN=FLEET_ETK_… sh

Set your public endpoint & adopt models.

Tell Kan how end users will reach this fleet (an IP, a hostname, or both — for the SNI list and the cert SAN). Then run "Discover & Adopt" — Hux scans the host for vLLM/SGLang containers and registers them as routable models. Or deploy new ones from the Models page.

Generate API keys & configure access.

Issue keys per end user or app. Group them into consumer groups and assign rate limits — requests/minute, requests/day, token budgets — and ACL rules like model allowlists, IP allowlists, time-of-day windows, expiry dates. Everything sync's into the fleet's local Redis within 15 seconds.

Start serving.

Your end users hit the OpenAI-compatible endpoint with their key. APISIX authenticates against the local cache, enforces ACLs and budgets, dispatches to the right model, and streams the completion back. Every request becomes a usage event you can see live in the Kan dashboard.

curl https://gateway.your-company.com/v1/chat/completions \
  -H "Authorization: Bearer sk_…" \
  -H "Content-Type: application/json" \
  -d '{"model":"llama-3.2-1b","messages":[{"role":"user","content":"hi"}]}'
Architecture

Hooks on every GPU. One eye watching them all.

Hux is a Go agent. It sits inside your GPU VM next to vLLM/SGLang/TRT-LLM and serves as the gateway, observability stack, and config receiver. Kan is a stateless control plane in our cloud. It renders gateway configs, distributes API keys, ingests usage events, and shows you everything in a single dashboard.

KAN — CONTROL PLANE renders configs · distributes keys · ingests usage HUX · GPU NODE APISIX · Redis · vLLM Prometheus · Grafana · Loki HUX · GPU NODE APISIX · Redis · SGLang Prometheus · Grafana · Loki HUX · GPU NODE APISIX · Redis · TRT-LLM Prometheus · Grafana · Loki
Read the full architecture →
Live preview

See your fleet, in real time.

This is what your Kan console looks like once a Hux is enrolled and serving traffic. Live request rates, p95 latency, hourly spend, and a streaming feed of every inference call as it lands.

Live overview
4 gateways · streaming
live
REQ/MIN
19,328
P95
312ms
SPEND/H
$4.82
req_01HG7XN4PWQ… gpt-4o 312ms 200
req_01HG7XN3VCQ… claude-3.5-sonnet 488ms 200
req_01HG7XN1X2T… gpt-4o-mini 204ms 200
req_01HG7XMZHQA… llama-3.2-1b 96ms 200

Numbers shown are illustrative. Your real dashboard updates from your fleet's Prometheus + Loki — Kan never sees prompt or completion content.

Sovereign by default.

Your tokens never cross our wire. Your keys never leave your fleet's Redis. Your model weights never see our servers. We see metering, you keep everything else.