Create a message

POST /messages

Generate a response with the Anthropic Messages API. The Anthropic SDK uses your Munito key through X-API-Key. Set its base URL to https://api.munito.ai/inference.

  • Send the conversation in messages and the system prompt in system.
  • Set stream: true for Anthropic SSE events.
  • Use tools and tool_choice for tool calls. Reasoning models can return thinking content blocks.
  • Set thinking to turn reasoning on (enabled or adaptive) or off (disabled). The API validates budget_tokens but does not enforce it. max_tokens bounds the whole answer.

The API requires max_tokens. Values above the model's output maximum use that maximum. A response that reaches the cap returns stop_reason: "max_tokens".

Requests are stateless and billed per token. Use /messages/count_tokens to count input tokens without inference.

POST
/messages

Authorization

AuthorizationBearer <token>

Your Munito API key.

In: header

Header Parameters

munito-service-tier?string

priority, standard or flex. The response names the tier that is billed; see x-ratelimit-over-limit.

Default"standard"

Value in

  • "priority"
  • "standard"
  • "flex"
munito-affinity?string

An opaque token. Send it back with the next request of the conversation, so that the prompt cache can serve it.

munito-overflow?string

off, model, class or any: what may serve a chat request that its own capacity refuses. The response names the step that served it. Off by default. The request keeps its region, and you pay the price of the model that you asked for. Flex, background and free-tier requests do not overflow.

Default"off"

Value in

  • "off"
  • "model"
  • "class"
  • "any"

Request Body

application/json

TypeScript Definitions

Use the request body type in TypeScript.

model*string

Model id from GET /models, e.g. qwen/qwen3.8-27b.

messages*array<>

The conversation as Anthropic-style messages: content is a string or an array of content blocks.

max_tokens*integer

Maximum tokens to generate. Required. Clamped to the model's output maximum.

system?string

System prompt (top-level, Anthropic-style).

stream?boolean

Stream Anthropic SSE events (message_start … message_stop).

Defaultfalse
tools?array<>

Anthropic-style tools the model may call (steer with tool_choice).

Response Body

application/json

curl -X POST "https://example.com/messages" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen/qwen3.8-27b",
    "max_tokens": 512,
    "messages": [
      {
        "role": "user",
        "content": "Explain intelligence sovereignty in one sentence."
      }
    ]
  }'
{
  "id": "msg_7f3c1a",
  "type": "message",
  "role": "assistant",
  "model": "qwen/qwen3.8-27b",
  "content": [
    {
      "type": "text",
      "text": "Intelligence sovereignty is the ability to run, govern, and improve your own AI on infrastructure you control, with no foreign cloud in the loop."
    }
  ],
  "stop_reason": "end_turn",
  "usage": {
    "input_tokens": 18,
    "output_tokens": 34
  }
}