Chat completion

POST /chat/completions

Create a chat completion with an open model. Set the OpenAI SDK base URL to https://api.munito.ai/inference/v1 and use your Munito API key. Both Authorization: Bearer and X-API-Key are accepted. Use stream for streamed output, tools for function calls, and response_format for structured output. Each model accepts its own parameters, media types, and limits, and the API answers 400 to a request outside them. Supported media can use inline data or validated remote URLs. Set store to true to retain the completion. Account limits can further restrict output.

POST
/chat/completions

Authorization

AuthorizationBearer <token>

Your Munito API key.

In: header

Header Parameters

munito-service-tier?string

priority, standard or flex. The response names the tier that is billed; see x-ratelimit-over-limit.

Default"standard"

Value in

  • "priority"
  • "standard"
  • "flex"
munito-affinity?string

An opaque token. Send it back with the next request of the conversation, so that the prompt cache can serve it.

munito-overflow?string

off, model, class or any: what may serve a chat request that its own capacity refuses. The response names the step that served it. Off by default. The request keeps its region, and you pay the price of the model that you asked for. Flex, background and free-tier requests do not overflow.

Default"off"

Value in

  • "off"
  • "model"
  • "class"
  • "any"
prefer?"respond-async"

respond-async: answer 202 at once. Then poll the location until the request finishes.

Value in

  • "respond-async"

Request Body

application/json

TypeScript Definitions

Use the request body type in TypeScript.

model*string

Model id from GET /models, e.g. qwen/qwen3.8-27b.

messages*array<>

OpenAI-style messages. Content can contain text and media parts when the selected model supports them.

max_tokens?integer

Maximum tokens to generate. Omit it to allow the output maximum of the model, which GET /models gives as max_output_length. A larger value uses that maximum.

temperature?number

Sampling temperature (0–2). Omit it to use the default of the model.

stream?boolean

Stream tokens back as server-sent events.

Defaultfalse
tools?array<>

OpenAI-style function tools the model may call (steer with tool_choice).

store?boolean

Persist the completion for later retrieval via GET /chat/completions/{id}.

Defaultfalse

Response Body

application/json

curl -X POST "https://example.com/chat/completions" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen/qwen3.8-27b",
    "messages": [
      {
        "role": "user",
        "content": "Explain intelligence sovereignty in one sentence."
      }
    ],
    "max_tokens": 512
  }'
{
  "id": "chatcmpl-7f3c1a",
  "object": "chat.completion",
  "model": "qwen/qwen3.8-27b",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Intelligence sovereignty is the ability to run, govern, and improve your own AI on infrastructure you control, with no foreign cloud in the loop."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 18,
    "completion_tokens": 34,
    "total_tokens": 52
  }
}