Create a message
POST /messages
Generate a response with the Anthropic Messages API. The Anthropic SDK uses your Munito key through X-API-Key. Set its base URL to https://api.munito.ai/inference.
- Send the conversation in
messagesand the system prompt insystem. - Set
stream: truefor Anthropic SSE events. - Use
toolsandtool_choicefor tool calls. Reasoning models can returnthinkingcontent blocks. - Set
thinkingto turn reasoning on (enabledoradaptive) or off (disabled). The API validatesbudget_tokensbut does not enforce it.max_tokensbounds the whole answer.
The API requires max_tokens. Values above the model's output maximum use that maximum. A response that reaches the cap returns stop_reason: "max_tokens".
Requests are stateless and billed per token. Use /messages/count_tokens to count input tokens without inference.
Your Munito API key.
In: header
Header Parameters
priority, standard or flex. The response names the tier that is billed; see x-ratelimit-over-limit.
"standard"Value in
- "priority"
- "standard"
- "flex"
An opaque token. Send it back with the next request of the conversation, so that the prompt cache can serve it.
off, model, class or any: what may serve a chat request that its own capacity refuses. The response names the step that served it. Off by default. The request keeps its region, and you pay the price of the model that you asked for. Flex, background and free-tier requests do not overflow.
"off"Value in
- "off"
- "model"
- "class"
- "any"
Request Body
application/json
TypeScript Definitions
Use the request body type in TypeScript.
Model id from GET /models, e.g. qwen/qwen3.8-27b.
The conversation as Anthropic-style messages: content is a string or an array of content blocks.
Maximum tokens to generate. Required. Clamped to the model's output maximum.
System prompt (top-level, Anthropic-style).
Stream Anthropic SSE events (message_start … message_stop).
falseAnthropic-style tools the model may call (steer with tool_choice).
Response Body
application/json
curl -X POST "https://example.com/messages" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen/qwen3.8-27b",
"max_tokens": 512,
"messages": [
{
"role": "user",
"content": "Explain intelligence sovereignty in one sentence."
}
]
}'{
"id": "msg_7f3c1a",
"type": "message",
"role": "assistant",
"model": "qwen/qwen3.8-27b",
"content": [
{
"type": "text",
"text": "Intelligence sovereignty is the ability to run, govern, and improve your own AI on infrastructure you control, with no foreign cloud in the loop."
}
],
"stop_reason": "end_turn",
"usage": {
"input_tokens": 18,
"output_tokens": 34
}
}