Quickstart

Five minutes to your first inference call.

Sign up, drop the agent on a GPU box, hit the gateway. That's the whole tutorial.

1. Sign up.

Go to app.huxkan.com/signup. Free 14-day trial, no credit card.

2. Create a fleet.

From the console: Gateways → New fleet. Give it a name (e.g. prod-us-east). Kan generates a fleet enroll token and shows you the installer command.

3. Run the installer on your GPU VM.

SSH into the GPU VM and paste the installer:

curl -sSL https://install.huxkan.com | FLEET_TOKEN=FLEET_ETK_live_… sh

The installer puts the agent at /usr/local/bin/kan-agent, registers a systemd unit, starts it. The agent enrolls with Kan, brings up APISIX/Redis/Prometheus/Grafana/Loki containers, and starts heartbeating.

4. Set the public endpoint.

From the console: Gateways → your fleet → globe icon. Type the address(es) your end users will reach the gateway at: 10.0.1.50, gateway.acme.com, or both comma-separated. The agent uses this for the SNI list and cert SAN.

5. Adopt or deploy a model.

If you already have vLLM running on the box: Gateways → your node → Discover & Adopt. Hux scans for OpenAI-compatible servers and registers them.

Or, from Models → Deploy new, pick an HF repo, pick a node, click Deploy. Hux pulls and starts the container.

6. Issue an API key.

Keys → Generate. Pick a consumer name, optionally assign a consumer group with rate limits / budgets. Copy the sk_… key (shown once).

7. Hit the gateway.

curl https://gateway.acme.com/v1/chat/completions \
  -H "Authorization: Bearer sk_…" \
  -H "Content-Type: application/json" \
  -d '{"model":"llama-3.2-1b","messages":[{"role":"user","content":"hello"}]}'

You should see a chat completion stream back. Watch the Dashboards page in Kan — the request shows up in real-time metrics.


Sign up and start → Full API reference