Pular para o conteúdo principal
AI Router API

API quickstart

An OpenAI-compatible chat API. Use the SDKs you already know, with https://hostingswift.com/api/v1 as the base URL.

Overview

The AI Router exposes two endpoints. Requests and responses follow the OpenAI format for chat completions.

Base URL
https://hostingswift.com/api/v1
  • POST /chat/completions create a chat completion (optionally streamed).
  • GET /models list the models your key can use, with context length and rates.

Authentication

Send your key as a bearer token. Create keys in the AI Router section of your client panel. A key looks like hs-sk-..., is shown once, and only a hash is stored on our side. Keep it server-side and never ship it in browser code.

Authorization: Bearer hs-sk-xxxxxxxxxxxxxxxxxxxxxxxx

Each key can be limited to specific models, given a spending budget and its own requests-per-minute limit. Revoke a key at any time.

Quickstart

Set HS_API_KEY in your environment, then send a request.

curl https://hostingswift.com/api/v1/chat/completions \
  -H "Authorization: Bearer $HS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-chat",
    "messages": [
      {"role": "system", "content": "You are a concise assistant."},
      {"role": "user", "content": "Explain DNS TTLs in two sentences."}
    ]
  }'

The response is a standard chat completion object with choices and usage.

Streaming

Add "stream": true to receive server-sent events. Disable proxy buffering on your side (for curl use -N).

curl -N https://hostingswift.com/api/v1/chat/completions \
  -H "Authorization: Bearer $HS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"deepseek-chat","stream":true,
       "messages":[{"role":"user","content":"Write a haiku about uptime."}]}'

Models

The catalogue changes over time, so list models instead of hard-coding them. Each entry includes context_length, a free flag and pricing in US cents per 1M tokens. A live table is on the AI Router page.

curl https://hostingswift.com/api/v1/models -H "Authorization: Bearer $HS_API_KEY"

Request parameters

Unknown top-level fields are ignored rather than forwarded.

Chat completion request parameters
FieldTypeNotes
modelstring, requiredA model ID from GET /models.
messagesarray, required1 to 200 messages with role system, user, assistant or tool.
streambooleanStream the reply as server-sent events.
max_tokensinteger1 to 8192.
temperaturenumber0 to 2.
top_pnumber0 to 1.
stopstring or arrayUp to 4 stop sequences.
presence_penalty / frequency_penaltynumber-2 to 2.
seedintegerBest-effort determinism where the upstream model supports it.
response_formatobjectPassed through, for example {"type":"json_object"}.
tools / tool_choicearray / anyPassed through for models that support tool calling.

Errors

Errors use the OpenAI error envelope. Switch on error.code, not the message text.

HTTP/1.1 429 Too Many Requests
Retry-After: 12

{"error":{"message":"Rate limit exceeded","type":"invalid_request_error","code":"rate_limited"}}
Error codes
HTTPcodeMeaning
400invalid_json, guardrail_blockedBody is not valid JSON, or a guardrail rule rejected the prompt.
401missing_token, invalid_api_key, key_revokedNo bearer token, an unknown key, or a revoked key.
402insufficient_credit, budget_exceededAccount credit is empty or the key budget is used up.
403insufficient_scope, model_not_allowed, account_inactiveThe key cannot use this endpoint or model, or the account is not active.
404model_not_foundUnknown or disabled model ID.
413payload_too_largeRequest body is larger than 2 MB.
422invalid_requestA field failed validation. The message names the field.
429rate_limitedToo many requests. Wait for Retry-After seconds.
502 / 503upstream_error, ai_disabled, provider_not_configuredThe upstream provider failed, or AI is temporarily disabled.

Rate limits and budgets

  • Each key defaults to 60 requests per minute; the limit can be changed per key.
  • A coarser per-IP limit applies before authentication to blunt key-guessing.
  • Responses include X-RateLimit-Limit and X-RateLimit-Remaining. A 429 also carries Retry-After (seconds).
  • Optional per-key budgets stop spending when reached (budget_exceeded). Account credit must be positive (insufficient_credit).
  • Output is capped at 8,192 tokens per request and request bodies at 2 MB.

Retry 429 and 5xx responses with exponential backoff. Do not retry 4xx errors other than 429.

Compatibility notes

  • Supported today: chat completions and model listing.
  • Not available: images, embeddings, audio, legacy text completions and file endpoints.
  • Optional prompt guardrails can redact secrets and personal data before a prompt is sent upstream, which may alter prompt text.
  • If one upstream provider fails, the gateway may retry the request through another. The X-HS-Provider response header reports which one answered.

Get your API key

Sign in, add credit or pick a free model, and create a key from the AI Router page of your panel.