Overview
The AI Router exposes two endpoints. Requests and responses follow the OpenAI format for chat completions.
https://hostingswift.com/api/v1POST /chat/completionscreate a chat completion (optionally streamed).GET /modelslist the models your key can use, with context length and rates.
Authentication
Send your key as a bearer token. Create keys in the AI Router section of your client panel. A key looks like hs-sk-..., is shown once, and only a hash is stored on our side. Keep it server-side and never ship it in browser code.
Authorization: Bearer hs-sk-xxxxxxxxxxxxxxxxxxxxxxxxEach key can be limited to specific models, given a spending budget and its own requests-per-minute limit. Revoke a key at any time.
Quickstart
Set HS_API_KEY in your environment, then send a request.
curl https://hostingswift.com/api/v1/chat/completions \
-H "Authorization: Bearer $HS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-chat",
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "Explain DNS TTLs in two sentences."}
]
}'The response is a standard chat completion object with choices and usage.
Streaming
Add "stream": true to receive server-sent events. Disable proxy buffering on your side (for curl use -N).
curl -N https://hostingswift.com/api/v1/chat/completions \
-H "Authorization: Bearer $HS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"deepseek-chat","stream":true,
"messages":[{"role":"user","content":"Write a haiku about uptime."}]}'Models
The catalogue changes over time, so list models instead of hard-coding them. Each entry includes context_length, a free flag and pricing in US cents per 1M tokens. A live table is on the AI Router page.
curl https://hostingswift.com/api/v1/models -H "Authorization: Bearer $HS_API_KEY"Request parameters
Unknown top-level fields are ignored rather than forwarded.
| Field | Type | Notes |
|---|---|---|
model | string, required | A model ID from GET /models. |
messages | array, required | 1 to 200 messages with role system, user, assistant or tool. |
stream | boolean | Stream the reply as server-sent events. |
max_tokens | integer | 1 to 8192. |
temperature | number | 0 to 2. |
top_p | number | 0 to 1. |
stop | string or array | Up to 4 stop sequences. |
presence_penalty / frequency_penalty | number | -2 to 2. |
seed | integer | Best-effort determinism where the upstream model supports it. |
response_format | object | Passed through, for example {"type":"json_object"}. |
tools / tool_choice | array / any | Passed through for models that support tool calling. |
Errors
Errors use the OpenAI error envelope. Switch on error.code, not the message text.
HTTP/1.1 429 Too Many Requests
Retry-After: 12
{"error":{"message":"Rate limit exceeded","type":"invalid_request_error","code":"rate_limited"}}| HTTP | code | Meaning |
|---|---|---|
| 400 | invalid_json, guardrail_blocked | Body is not valid JSON, or a guardrail rule rejected the prompt. |
| 401 | missing_token, invalid_api_key, key_revoked | No bearer token, an unknown key, or a revoked key. |
| 402 | insufficient_credit, budget_exceeded | Account credit is empty or the key budget is used up. |
| 403 | insufficient_scope, model_not_allowed, account_inactive | The key cannot use this endpoint or model, or the account is not active. |
| 404 | model_not_found | Unknown or disabled model ID. |
| 413 | payload_too_large | Request body is larger than 2 MB. |
| 422 | invalid_request | A field failed validation. The message names the field. |
| 429 | rate_limited | Too many requests. Wait for Retry-After seconds. |
| 502 / 503 | upstream_error, ai_disabled, provider_not_configured | The upstream provider failed, or AI is temporarily disabled. |
Rate limits and budgets
- Each key defaults to 60 requests per minute; the limit can be changed per key.
- A coarser per-IP limit applies before authentication to blunt key-guessing.
- Responses include
X-RateLimit-LimitandX-RateLimit-Remaining. A429also carriesRetry-After(seconds). - Optional per-key budgets stop spending when reached (
budget_exceeded). Account credit must be positive (insufficient_credit). - Output is capped at 8,192 tokens per request and request bodies at 2 MB.
Retry 429 and 5xx responses with exponential backoff. Do not retry 4xx errors other than 429.
Compatibility notes
- Supported today: chat completions and model listing.
- Not available: images, embeddings, audio, legacy text completions and file endpoints.
- Optional prompt guardrails can redact secrets and personal data before a prompt is sent upstream, which may alter prompt text.
- If one upstream provider fails, the gateway may retry the request through another. The
X-HS-Providerresponse header reports which one answered.